Repository navigation
macOS: Persistent SQLite TRACE target=log churn remains after rust-v0.142.0 #29532
Description
Activity
- addedbugSomething isn't workingSomething isn't workingappIssues related to the Codex desktop appIssues related to the Codex desktop appapp-serverIssues involving app server protocol or interfacesIssues involving app server protocol or interfaces
on Jun 23, 2026 github-actions commented
on Jun 23, 2026 on Jun 23, 2026 – with GitHub ActionsContributorMore actionsPotential duplicates detected. Please review them and close your issue if it is a duplicate.
- Windows Codex app continuously writes high-volume TRACE websocket logs to ~/.codex/logs_2.sqlite despite RUST_LOG=warn #29463
- Codex SQLite feedback logs can write ~640 TB/year and rapidly consume SSD endurance #28224
Powered by Codex Action
Additional macOS reproduction on the current Codex app build.
Environment:
- Platform: macOS, Apple internal SSD (
APPLE SSD AP0512Z), SMART statusVerified - Codex app:
26.616.71553, build4265 - Embedded CLI:
codex-cli 0.142.0 - App-server process:
/Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled - Active app-server PID during sample:
34793 RUST_LOG=infowas set vialaunchctl getenv RUST_LOG, but persisted SQLite rows still includedTRACEandDEBUG.- No SQLite trigger or DB mutation was applied during this sample; the checks below were read-only.
Files observed:
~/.codex/logs_2.sqlite 302,129,152 bytes (~288M) ~/.codex/logs_2.sqlite-wal 13,113,992 bytes (~13M) ~/.codex/logs_2.sqlite-shm 32,768 byteslsofshowed the active Codex app-server process holding open handles to the DB, WAL, and SHM files.60-second quiet sample:
Sample start: 2026-06-23 10:40:06 CST Sample end: 2026-06-23 10:41:07 CST max(id): 19457302 -> 19458968 (+1666) count(*): 29226 -> 29235 (+9) max(ts): 1782182405 -> 1782182465 WAL mtime: kept advancing throughout the sampleThis again suggests insert-and-prune churn: retained row count barely changed, but
max(id)advanced by 1,666 in about a minute.Rows visible in that sample window included:
TRACE log 870 rows, 473846 estimated bytes TRACE codex_app_server::message_processor 13 rows TRACE codex_api::sse::responses 11 rows TRACE hyper_util::client::legacy::pool 11 rows TRACE codex_app_server::outgoing_message 9 rows TRACE codex_mcp::connection_manager 8 rowsTotal visible
TRACErows in the sample window:940.The latest
TRACE target=logsamples were websocket/tungstenite noise, including:WouldBlock tokio-tungstenite ... Read.with_context read -> poll_read WebSocketStream.with_context Stream.poll_next Received message ... decompressing ... bytes in final frame received frame <FRAME> ...Current visible level distribution at collection time:
TRACE 15498 rows 53.03% INFO 12331 rows 42.19% DEBUG 1321 rows 4.52% WARN 71 rows ERROR 5 rowsImpact: this creates sustained writes to the persistent SQLite log sink during normal Codex app use.
RUST_LOG=infodid not stop these persistedTRACErows, so the user-level mitigation does not appear to work for this path.I have a private evidence package with raw command output, SQLite read-only queries, file metadata, hashes, and the 60-second WAL/SQLite sample. I am not attaching raw
feedback_log_bodycontents publicly because they may include local paths, conversation IDs, command fragments, or response data.Reacted by YUNG BOJANG and Runtian Lee- Platform: macOS, Apple internal SSD (
Please fix this as soon as possible! This issue is widespread and serious. Thank you.
Reacted by lock-down, soockie, stone, Runtian Lee and sukaZainan+1 Please fix this as soon as possible! This issue is widespread and serious. Thank you.
Additional current-build datapoint from macOS Desktop
26.616.71553/ build4265, bundled/globalcodex-cli 0.142.0.During active Desktop use, I temporarily removed a local SQLite mitigation trigger, recorded
max(id)=707460512, waited 15 seconds, then restored the trigger. New retained rows over that window:total_new_rows=1196 log|TRACE|934 codex_api::sse::responses|TRACE|138 codex_app_server::outgoing_message|TRACE|37 codex_mcp::connection_manager|TRACE|20 codex_config::loader::macos|DEBUG|8 codex_core::stream_events_utils|DEBUG|7 feedback_tags|INFO|7 codex_models_manager::cache|INFO|6 codex_models_manager::manager|INFO|5 codex_app_server::message_processor|TRACE|4 codex_config::loader::layer_io|DEBUG|4 codex_core::session::turn|TRACE|4 codex_core::stream_events_utils|INFO|4 codex_api::endpoint::responses_websocket|TRACE|3 codex_core::spawn|TRACE|3Rows matching the previously noisy targets (
log,codex_otel.log_only,codex_otel.trace_safe,codex_api::endpoint::responses_websocket,codex_api::sse::responses) were1075in the same 15-second window.This is another confirmation that the current Desktop app-server path is still persisting high-volume TRACE
target=logrows on the shipped 0.142.0 build. The local mitigation trigger was restored after the measurement.Additional macOS data point from another machine: I can still reproduce persistent SQLite TRACE churn after restarting Codex Desktop and after manually compacting the log DB.
Environment
- Platform: macOS arm64
- Codex Desktop:
26.616.71553, build4265 - Embedded CLI:
codex-cli 0.142.0 - PATH CLI:
codex-cli 0.142.0 - Active Desktop app-server process:
/Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled
- A separate older stdio app-server process was also present, but recent rows were dominated by the restarted Desktop app-server process.
RUST_LOGobserved in the process environment:warn~/.codex/logs_2.sqlitejournal mode observed earlier:wal
I am not attaching or pasting raw
feedback_log_bodyrows because they can contain private conversation/tool/local-path details.Before manual cleanup
Before compacting, the local log files were:
~/.codex/logs_2.sqlite 703M ~/.codex/logs_2.sqlite-wal 128M ~/.codex/logs_2.sqlite-shm 256KThe table showed high churn/free space:
rows: 143001 page_count: 179883 freelist_count: 133819 page_size: 4096A 30 second sample after restarting Codex Desktop showed continued insert-and-prune activity:
max(id) delta: about +1414 in 30s approx rate: 45-50 inserted ids/secRecent 2 minute target distribution during that sample was still dominated by TRACE rows:
TRACE log 1071 rows TRACE codex_mcp::connection_manager 55 rows TRACE codex_api::sse::responses 51 rows TRACE codex_app_server::outgoing_message 51 rows TRACE codex_api::endpoint::responses_websocket 4 rows, about 2.17 MiB estimatedAfter manual cleanup / compaction
I then retained only a small recent window of diagnostic rows and ran checkpoint/VACUUM. Integrity check returned
ok, and the files were reduced to:15:53:24 local logs_2.sqlite 1.8M logs_2.sqlite-wal 0B logs_2.sqlite-shm 256K rows: 1855However, without changing anything else, Codex Desktop quickly started growing the files again. About 80 seconds later:
15:54:44 local logs_2.sqlite 38M logs_2.sqlite-wal 37M logs_2.sqlite-shm 256K rows: 2061 max(id): 243430233Rows retained in the last 5 minutes at that point:
TRACE log 1054 rows, 0.18 MiB estimated TRACE codex_api::sse::responses 621 rows, 0.60 MiB estimated TRACE codex_mcp::connection_manager 187 rows, 0.19 MiB estimated TRACE codex_app_server::outgoing_message 48 rows DEBUG codex_core::stream_events_utils 45 rows TRACE codex_api::endpoint::responses_websocket 8 rows, 0.57 MiB estimatedExpected
After
rust-v0.142.0/ #29432 / #29457, the persistent SQLite sink should not continue high-frequency insert-and-prune churn dominated byTRACE target=logand related websocket/SSE targets during normal Desktop use.Actual
Even after restart and after compacting the database down to a tiny valid state, the active Desktop app-server began writing again immediately. The on-disk DB/WAL grew from about
1.8M + 0Bto about38M + 37Mwithin roughly 80 seconds, and the retained recent rows were still dominated by TRACE targets.This makes the current behavior look like an incomplete fix or a remaining Desktop-app persistent-log sink path that still persists noisy TRACE rows.
I’m seeing what looks like the same underlying SQLite churn pattern on macOS, though my setup differs slightly: I have not specifically disabled analytics/plugins, and the running process is
codex app-server --analytics-default-enabled.I posted the fuller diagnostic here: #29612 (comment)
Summary of my repro:
Codex Desktop: 26.616.71553 • Released Jun 22, 2026 Bundled CLI: codex-cli 0.142.0 Platform: macOS 15.3, Darwin 24.3.0, arm64 Launch method: macOS GUI / Spotlight Running process: /Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled Actual process env includes: RUST_LOG=warnOver a 5-minute idle sample, with Codex Desktop open in the background but not actively used:
new_insert_ids=2975 retained_row_delta=0 insert_ids_per_sec=9.92The DB was healthy and compacted beforehand, so this did not appear to be leftover pre-fix bloat:
quick_check=ok integrity_check=ok journal_mode=wal freelist_count=115 logs_2.sqlite≈63M logs_2.sqlite-wal≈4.1MThe exact sampled ID window showed retained rows dominated by
TRACE target=log:TRACE log 998 rows TRACE hyper_util::client::legacy::pool 2 rowsAll retained rows from that sampled window belonged to the same app-server process.
So this looks like the same post-0.142.0 behavior: file size may remain bounded, but the app-server still continuously inserts and prunes persistent SQLite log rows while idle, mostly
TRACE target=log, despiteRUST_LOG=warn.Additional macOS datapoint after cleanup/compaction, from another affected machine.
I can still reproduce the post-
0.142.0insert-and-prune churn after fully compacting the log DB. The historical bloat was cleaned up successfully, but the current app-server continues to persistTRACE target=logrows.Environment
Platform: macOS arm64 Codex Desktop: 26.616.71553 CFBundleVersion: 4265 CLI: codex-cli 0.142.0 Running process: /Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled App-server PID/start: 9920, Wed Jun 24 10:19:03 2026 AEST Process env: RUST_LOG=warn Active DB: ~/.codex/logs_2.sqliteThe active app-server process was holding open handles to
logs_2.sqlite,logs_2.sqlite-wal, andlogs_2.sqlite-shm.Before cleanup
Before manually clearing/compacting the DB, the files had grown to:
~/.codex/logs_2.sqlite 811M ~/.codex/logs_2.sqlite-wal 68M ~/.codex/logs_2.sqlite-shm 160KSQLite stats at that point showed this was mostly free pages from churn:
rows: 24495 page_size: 4096 page_count: 207595 freelist_count:195037 estimated_payload_mb: 26.66After
DELETE FROM logs, resetting the sequence, andVACUUM, the DB compacted successfully:~/.codex/logs_2.sqlite 5.3M ~/.codex/logs_2.sqlite-wal 4.0M-4.3M ~/.codex/logs_2.sqlite-shm 32KRepro after cleanup / restart
Immediately after cleanup and restart, rows began accumulating again, dominated by
TRACE target=log:rows: 1062 seq: 4027 TRACE log: 978 rowsA 20-second sample shortly after restart showed much higher active churn:
window_start=2026-06-24 10:19:48 AEST window_end=2026-06-24 10:20:08 AEST delta_seq=15923 rows_per_sec=796.1 delta_count=4Top retained rows from that ID window:
TRACE log 932 rows, 0.176 MiB TRACE codex_app_server::outgoing_message 39 rows TRACE hyper_util::client::legacy::pool 7 rows TRACE hyper_util::client::legacy::client 6 rows TRACE codex_api::sse::responses 2 rowsA later quieter 30-second sample still showed the same insert-and-prune pattern with file sizes bounded:
window_start=2026-06-24 10:23:52 AEST window_end=2026-06-24 10:24:22 AEST start_seq=67982 end_seq=68515 delta_seq=533 rows_per_sec=17.8 start_count=2049 end_count=2049 delta_count=0 start_db=5541888 end_db=5541888 delta_db=0 start_wal=4486712 end_wal=4486712 delta_wal=0Top retained rows from that sampled ID window:
TRACE log 475 rows, 0.212 MiB TRACE codex_app_server::outgoing_message 12 rows TRACE hyper_util::client::legacy::pool 10 rows TRACE hyper_util::client::legacy::client 6 rows TRACE codex_api::sse::responses 4 rows DEBUG hyper_util::client::legacy::connect::http 4 rowsI am intentionally not pasting raw
feedback_log_bodyvalues because they may contain private local paths, conversation, tool, or response data.User-visible impact
Before cleanup, while several Codex tasks were open at the same time, I observed frequent Codex restarts and UI lag. Cleaning/VACUUM reduced the historical DB size from ~811M to ~5.3M, but the high-frequency
TRACE target=loginsert/prune behavior is still present on the current Desktop app-server.So the remaining issue seems to be:
0.142.0reduces/bounds some file-size growth, but the shipped macOS Desktop app-server still persists high-frequencyTRACE target=logrows and continues SQLite churn even withRUST_LOG=warn.Additional datapoint after updating to the newer macOS app build
26.616.81150/ build4306.Environment:
- Codex app:
26.616.81150, build4306 - Embedded CLI:
codex-cli 0.142.0 - Active app-server process started after update:
/Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled - App-server PID during sample:
88553, startedWed Jun 24 09:20:21 2026 - A local SQLite mitigation trigger existed before the test, so I temporarily removed it for a controlled 30-second sample, then restored it immediately afterward.
Controlled sample with the trigger disabled:
Sample window: 2026-06-24 09:22:34 CST to 09:23:05 CST max(id): 19975524 -> 19975884 (+360) count(*): 33614 -> 33965 (+351) WAL size: 0 -> 1,030,032 bytes during the sampleRows visible in the sample window:
TRACE log 441 rows, 271850 estimated bytes TRACE codex_app_server::outgoing_message 4 rows TRACE codex_mcp::connection_manager 4 rows TRACE codex_api::endpoint::responses_websocket 1 row TRACE codex_api::sse::responses 1 row TRACE codex_core::session::turn 1 row TRACE hyper_util::client::legacy::pool 1 rowTotal
TRACErows in the 30-second window:453.The latest
TRACE target=logrows are still websocket/tungstenite noise, includingWouldBlock,tokio-tungstenite,WebSocketStream.with_context,Stream.poll_next,Received message, decompression, and<FRAME>parsing lines.So this still reproduces on app build
4306; it does not appear fully fixed by this build.- Codex app:
现在macos的codex桌面端更新到了26.622.11653,还有这个问题吗
almostlover-hao commented
on Jun 24, 2026 More actionsAdditional macOS data point from Codex Desktop showing that the issue still reproduces on my machine, and that
sqlite_homeis currently the only effective local mitigation I found.Environment
- Platform: macOS, Apple SSD (
diskutil: SMART StatusVerified, TRIM supported) - Codex Desktop app-server command:
/Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled
- Codex app version exposed in local config/env:
BROWSER_USE_CODEX_APP_VERSION=26.616.81150
- Embedded CLI:
codex-cli 0.142.0
- SQLite DBs involved:
~/.codex/logs_2.sqlite~/.codex/logs_2.sqlite-wal~/.codex/state_5.sqlite-wal
Config tried
~/.codex/config.tomlincluded:[analytics] enabled = false [otel] log_user_prompt = false exporter = "none" trace_exporter = "none" metrics_exporter = "none"
The restarted app-server process also inherited:
RUST_LOG=warnConfirmed via:
ps eww -p <app-server-pid> -o command= | tr ' ' '\n' | rg '^(RUST_LOG|CODEX|OTEL|RUST)='
Observed behavior before mitigation
After restarting Codex Desktop with the config above, a 60-second
proc_pid_rusagesample still showed:Codex-related processes total: 17.594 MB written / 60s app-server alone: 17.082 MB written / 60sThe new app-server process had already accumulated substantial writes shortly after restart:
app-server start_written_mb=167.871In an earlier active diagnostic window, app-server lifetime written bytes reached about:
app-server end_written_mb ~= 1373 MBRecent rows in
logs_2.sqlitestill included persisted TRACE records despiteRUST_LOG=warn:TRACE 2957 rows 1.65 MB DEBUG 95 rows 0.25 MB INFO 84 rows 0.07 MB WARN 13 rows 0.01 MBLargest recent targets included:
codex_api::sse::responses TRACE codex_api::endpoint::responses_websocket TRACE log TRACE codex_app_server::outgoing_message TRACE codex_mcp::connection_manager TRACEThis confirms the SQLite feedback/log sink is still persisting TRACE-level data even when the process environment has
RUST_LOG=warnand analytics/OTel exporters are disabled.Why file size alone hides the issue
The retained DB size was not enough to estimate real write volume. The file stayed bounded while
logs_2.sqlite-waland SQLite insert/prune churn continued. This matches the pattern described in this issue: current file size can look moderate while SSD writes continue due to insert, WAL write, prune, checkpoint, and rewrite activity.Local mitigation tested
I moved only Codex's SQLite-backed runtime state to a RAM disk:
sqlite_home = "/Volumes/CodexSQLiteRAM"
After restarting Codex Desktop,
lsofconfirmed these were all on the RAM disk:/Volumes/CodexSQLiteRAM/logs_2.sqlite /Volumes/CodexSQLiteRAM/logs_2.sqlite-wal /Volumes/CodexSQLiteRAM/state_5.sqlite-wal /Volumes/CodexSQLiteRAM/memories_1.sqlite-wal /Volumes/CodexSQLiteRAM/goals_1.sqlite-walA follow-up 60-second sample still showed process-level writes:
Codex-related processes total: 8.230 MB written / 60s app-server alone: 7.191 MB written / 60sBut visible internal SSD growth for the Codex SQLite files dropped to zero because the SQLite writes were on the RAM disk:
RAM_SQLITE_FILE_GROWTH_MB 0.118 SSD_OLD_SQLITE_FILE_GROWTH_MB 0.000 SSD_APP_LOG_FILE_GROWTH_MB 0.000 SSD_SESSION_FILE_GROWTH_MB 0.002A broader 30-second scan of likely Codex internal-SSD paths showed only tiny non-SQLite growth:
~/.codex 0.000 MB ~/Library/Application Support/Codex 0.007 MB ~/Library/Caches/com.openai.codex 0.000 MB ~/Library/HTTPStorages/com.openai.codex 0.000 MB ~/Library/Logs/com.openai.codex 0.000 MB TOTAL_BROAD_FILE_GROWTH_MB 0.007 MB / 30sSo
sqlite_hometo RAM disk mitigates SSD wear, but it does not fix the root cause. It only moves the high-frequency SQLite WAL writes away from the internal SSD.Expected fix
Please make the persistent SQLite log sink honor the same filtering as
RUST_LOG/ runtime log level, or provide a documented config option to disable persistent SQLite feedback logs or cap them aggressively. WithRUST_LOG=warn,[analytics].enabled=false, and all[otel]exporters set tonone, TRACE websocket/SSE/internaltarget=logrecords should not continue to be persisted tologs_2.sqliteduring normal desktop usage.Privacy note
I am not attaching the SQLite DB or raw
feedback_log_bodyrows because they may contain private local paths, tool output, and conversation/session details.- Platform: macOS, Apple SSD (
20 remaining items
Additional macOS datapoint after updating to Codex app
26.623.101652/ build4674.Environment:
- Codex app:
26.623.101652, build4674 - Embedded CLI:
codex-cli 0.142.5 - Active app-server command:
/Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled RUST_LOGwas unset in the launchd environment at the time of the check.- A local SQLite mitigation trigger existed before the test, so I temporarily removed it for a controlled 30-second sample, then restored it immediately afterward and truncated the WAL again.
- No raw
feedback_log_bodyvalues are included here because they may contain private local paths, tool output, or conversation/session details.
Notable process state after a full app restart/update:
Two codex app-server processes were running and both held open handles to logs_2.sqlite, logs_2.sqlite-wal, and logs_2.sqlite-shm. PID PPID started 94151 1 Fri Jul 3 12:07:17 2026 94919 94887 Fri Jul 3 12:07:36 2026Controlled sample with the trigger disabled:
Sample window: 2026-07-03 12:09:43 CST to 12:10:13 CST max(id): 20053926 -> 20054522 (+596) count(*): 34196 -> 34792 (+596) WAL size: 0 -> 1,833,432 bytes during the sampleRows visible in the sample window:
level target rows estimated_bytes TRACE log 577 175550 DEBUG log 8 419 TRACE codex_api::sse::responses 3 3212 TRACE hyper_util::client::legacy::pool 3 665 INFO codex_app_server_transport::transport::remote_control::websocket 1 624 INFO codex_core::stream_events_utils 1 4630 INFO feedback_tags 1 133 TRACE codex_app_server::outgoing_message 1 160 WARN codex_app_server_transport::transport::remote_control::websocket 1 935Rows by process UUID in the same sample window:
process_uuid rows trace_rows debug_rows pid:94151:1af60d6a-3dc5-4608-bc97-d376c47db0ba 315 315 0 pid:94919:1b9d9fd8-2306-4905-b74d-829cfe067eab 281 269 8After restoring the local mitigation trigger and running
PRAGMA wal_checkpoint(TRUNCATE), the local state was:trigger_present: block_trace_debug_logs logs_2.sqlite-wal size: 0A follow-up 30-second observation with the trigger enabled persisted only INFO/WARN rows, while TRACE/DEBUG rows were not persisted by the local mitigation.
This suggests
0.142.5/ build4674may have improved the large websocket request-payload trace symptom from #30771, but the broader persistent TRACE/DEBUG SQLite churn still reproduces when the local SQLite mitigation is removed. In this run, both active app-server processes contributed rows.- Codex app:
Follow-up from the same affected Mac after updating to a newer Desktop build again.
Environment:
- Platform: macOS
27.0(26A5368g),arm64 - Codex app:
26.623.101652, build4674 - Embedded CLI:
codex-cli 0.142.5 - PATH CLI:
codex-cli 0.142.5 - Active desktop app-server:
/Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled - Active desktop app-server PID during check:
98496, startedMon Jul 6 10:49:08 2026 launchctl getenv RUST_LOGreturned empty/unset during this check- A separate stdio app-server was also present:
/Users/ron/.local/bin/codex app-server --listen stdio://, PID3728 - Local SQLite mitigation triggers were still enabled:
sample_log_inserts,trim_sampled_logs
Important note: this is another protected-state check, not a raw insert-rate measurement. I did not remove the local mitigation this time. On this machine those triggers retain at most about one row every 10 seconds and keep only the newest 500 retained rows.
Current retained SQLite state after the sample:
PRAGMA quick_check: ok journal_mode: wal row_count: 500 trace_rows: 500 page_count: 32769 freelist_count: 32671 page_size: 4096 latest retained row: 2026-07-06 11:12:07 Asia/Shanghai30-second protected sample:
before: count=500 max_id=189220 trace_rows=500 max_trace_id=189220 max_ts=1783307497 latest=2026-07-06 11:11:37 after: count=500 max_id=189223 trace_rows=500 max_trace_id=189223 max_ts=1783307527 latest=2026-07-06 11:12:07File sizes at the end of the check:
~/.codex/logs_2.sqlite 134221824 bytes ~/.codex/logs_2.sqlite-wal 4185952 bytes ~/.codex/logs_2.sqlite-shm 32768 bytesTop retained targets after the sample:
TRACE log 480 TRACE hyper_util::client::legacy::pool 15 TRACE hyper_util::client::legacy::client 4 TRACE codex_app_server::outgoing_message 1The newest retained rows immediately after the sample were still from the active Desktop app-server process UUID:
189225 2026-07-06 11:12:27 TRACE log pid:98496:06d227f2-204d-4c82-a197-bd163eb74d1f 189224 2026-07-06 11:12:17 TRACE log pid:98496:06d227f2-204d-4c82-a197-bd163eb74d1f 189223 2026-07-06 11:12:07 TRACE log pid:98496:06d227f2-204d-4c82-a197-bd163eb74d1f 189222 2026-07-06 11:11:57 TRACE log pid:98496:06d227f2-204d-4c82-a197-bd163eb74d1fConclusion from this machine:
26.623.101652/ build4674/codex-cli 0.142.5still does not look fully fixed. The large websocket request-payload trace symptom may be improved, as others noted for #30771, but the broader persistent SQLite TRACE sink still appears active: retained rows remain500/500TRACE andmax_trace_idcontinues to advance under the local sampled-retention trigger.I am not pasting raw
feedback_log_bodyvalues because they may include local paths, tool output, or conversation/session details.- Platform: macOS
tangleo203 commented
on Jul 10, 2026 More actionsAdditional reproduction on ChatGPT-integrated Codex / 0.144.0-alpha.4
I can still reproduce the bridged
TRACE target=logpersistence on a newer macOS build that embeds Codex inside the ChatGPT app.Environment
- macOS, Apple Silicon
- ChatGPT app:
26.707.31428(build5059) - Embedded binary:
/Applications/ChatGPT.app/Contents/Resources/codex - Embedded CLI version:
codex-cli 0.144.0-alpha.4 - SQLite DB:
~/.codex/logs_2.sqlite
Evidence collected before installing any local workaround
The current app-server process persisted 436 rows with
target=logbetween:2026-07-10 08:27:01 +0800 2026-07-10 08:27:05 +0800Example sanitized messages included:
TRACE log: deregistering event source from poller TRACE log: shouldn't retry!At the time of inspection, the retained target counts included:
log 2333 codex_api::sse::responses 1654 hyper_util::client::legacy::pool 547No retained rows were found for
codex_otel.log_onlyorcodex_otel.trace_safe, so that part of the filtering appears effective. However, bridgedtarget=logevents are still reaching SQLite, despite #29599 being intended to reject them inside the sink.This appears to be either:
- a regression after Stop persisting bridged log events #29599 / 0.143.0, or
- a different logging path in the ChatGPT-integrated app build that bypasses the sink-level
metadata.target() == "log"guard.
Local confirmation/workaround
After the measurement, I installed a SQLite
BEFORE INSERT ... RAISE(IGNORE)trigger for the noisy targets. A test insert was ignored, and normal non-blocked log IDs continued advancing while the blocked-target maximum ID stayed unchanged. This is only a local workaround; filtering before formatting/queueing remains preferable.Raw payloads, thread IDs, local paths, and conversation contents are intentionally omitted.
Follow-up from the same affected Mac after updating to the newer ChatGPT-integrated Codex build.
Environment:
- Platform: macOS
27.0(26A5378j),arm64 - ChatGPT app:
26.707.61608, build5200 - Embedded binary:
/Applications/ChatGPT.app/Contents/Resources/codex - Embedded CLI version:
codex-cli 0.144.0-alpha.4 command -v codexnow resolves to the ChatGPT embedded binary on this machine- The previous standalone
/Applications/Codex.app/Contents/Resources/codexbinary was not present during this check - Active ChatGPT-integrated app-server:
/Applications/ChatGPT.app/Contents/Resources/codex -c features.code_mode_host=true app-server --analytics-default-enabled - Active ChatGPT-integrated app-server PID during check:
92069, startedMon Jul 13 10:03:30 2026 - A separate stdio app-server was also present:
~/.local/bin/codex app-server --listen stdio://, PID89531 launchctl getenv RUST_LOGreturned empty/unset during this check- Local SQLite mitigation triggers were still enabled:
sample_log_inserts,trim_sampled_logs
Important note: this is another protected-state check, not a raw insert-rate measurement. I did not remove the local mitigation this time. On this machine those triggers retain at most about one row every 10 seconds and keep only the newest 500 retained rows.
Current retained SQLite state after the sample:
PRAGMA quick_check: ok journal_mode: wal row_count: 500 trace_rows: 451 debug_rows: 33 info_rows: 10 warn_rows: 6 error_rows: 0 page_count: 32769 freelist_count: 32649 page_size: 4096 latest retained row: 2026-07-13 10:33:48 Asia/Shanghai30-second protected sample:
before: count=500 max_id=221476 trace_rows=452 max_trace_id=221475 debug_rows=32 info_rows=10 warn_rows=6 error_rows=0 latest=2026-07-13 10:33:18 after: count=500 max_id=221479 trace_rows=451 max_trace_id=221479 debug_rows=33 info_rows=10 warn_rows=6 error_rows=0 latest=2026-07-13 10:33:48File sizes at the end of the check:
~/.codex/logs_2.sqlite 134221824 bytes ~/.codex/logs_2.sqlite-wal 44388912 bytes ~/.codex/logs_2.sqlite-shm 32768 bytesTop retained targets after the sample:
TRACE log 333 TRACE hyper_util::client::legacy::pool 27 TRACE codex_app_server::outgoing_message 25 TRACE hyper_util::client::legacy::client 25 TRACE codex_api::sse::responses 24 TRACE codex_app_server::message_processor 17 DEBUG codex_http_client::default_client 15 DEBUG codex_core::stream_events_utils 10 DEBUG opentelemetry-otlp 8 INFO codex_http_client::custom_ca 6 WARN codex_app_server_transport::transport::remote_control::websocket 5The newest retained rows immediately after the sample included new TRACE rows from both the ChatGPT-integrated app-server and the stdio app-server:
221479 2026-07-13 10:33:48 TRACE codex_app_server::message_processor pid:89531:7334f365-0926-407f-aac9-abe8fc58abf9 221478 2026-07-13 10:33:38 TRACE codex_api::sse::responses pid:92069:71bd96dc-8afa-4d14-85b4-d909fc2345a0 221477 2026-07-13 10:33:28 DEBUG codex_http_client::default_client pid:92069:71bd96dc-8afa-4d14-85b4-d909fc2345a0 221476 2026-07-13 10:33:18 DEBUG codex_core::stream_events_utils pid:92069:71bd96dc-8afa-4d14-85b4-d909fc2345a0 221475 2026-07-13 10:33:08 TRACE codex_app_server::outgoing_message pid:92069:71bd96dc-8afa-4d14-85b4-d909fc2345a0Conclusion from this machine: the old
500/500 TRACEprotected-state symptom has changed in the ChatGPT-integrated0.144.0-alpha.4build: retained rows now include DEBUG/INFO/WARN as well. However, the broader issue still does not look fully fixed here.max_trace_idcontinued to advance during the protected sample, andTRACE target=logremains the largest retained target. This matches the concern in the newer ChatGPT-integrated reports: part of the older filtering appears improved, but bridged TRACE/log-style events still reach the persistent SQLite DB.I am not pasting raw
feedback_log_bodyvalues because they may include local paths, tool output, or conversation/session details.- Platform: macOS
Our series of fixes for this problem are in the latest shipping desktop app, CLI, and IDE extension.
Not quite yet.
0.144.1does not have #31789, #31790, #31791, and #31792.0.145.0-alpha.4includes them, so0.145.0likely will too.Also, please see #32496. A different root cause is causing about 1.3 TB/year in unnecessary writes, according to the numbers in the issue. (Much less than was on this issue.) Codex is continually rewriting
models_cache.jsonandscope_v3.json.Originally posted by @darlingm in #28224 (comment)
I’ve completed an independent reference implementation and validation effort for the persistent SQLite TRACE/DEBUG issue described here.
Repository:
https://github.com/Freyliu0516/Codex-Log-Doctor
Proposed upstream change
The upstream patch is deliberately small and separate from the user-side diagnostic tool:
- change the persistent SQLite sink’s default level from
TRACEtoWARN; - keep the existing target-specific exclusions;
- filter low-level events before they are formatted, queued, or sent to SQLite;
- add a regression test that emits 100,000 TRACE events and 100,000 DEBUG events;
- verify that SQLite retains only the expected WARN event.
Reference patch:
https://github.com/Freyliu0516/Codex-Log-Doctor/tree/main/upstream
The independent tool does not patch or replace an installed Codex binary. The
upstream/directory exists only as a reviewable reference implementation.Validation
The patched Codex source was tested against the pinned upstream base:
-
codex-stateprimary tests: 151 passed; -
binary tests: 3 passed;
-
doctests: 1 passed;
-
persistence regression:
- 100,000 TRACE iterations;
- 100,000 DEBUG iterations;
- 0 unexpected rows persisted;
- the expected WARN row was retained.
The separate Codex Log Doctor repository also passed:
- formatting;
- workspace-wide Clippy;
- all workspace tests;
- release builds on Linux, macOS, and Windows;
- supply-chain checks;
- source-install smoke tests;
- upstream patch drift checks.
The repository additionally contains 18 deterministic component and failure-mode scenarios covering SQLite/WAL contention, queue saturation, backup and rollback, interrupted operations, unsupported schemas, and report redaction. These are supporting tests, not a claim that all 18 scenarios exercise the complete Codex app-server pipeline.
Public CI evidence is linked from the merged initial implementation PR:
Freyliu0516/Codex-Log-Doctor#1
User-side mitigation
The tool can install a fingerprint-gated SQLite
BEFORE INSERTtrigger that rejects TRACE rows. It requires:- an exact supported schema fingerprint;
- no detected Codex process using the database;
- exclusive SQLite access;
- sufficient disk space;
- a verified SQLite backup;
- explicit
--apply; - rollback metadata and post-operation integrity checks.
This is intended only as a reversible workaround. It reduces retained TRACE rows,
MAX(id)growth, WAL writes, indexing, and subsequent pruning, but it does not eliminate upstream event construction, formatting, queueing, transactions, or attempted INSERT execution.Early filtering in Codex remains the preferred fix.
Request
Would the Codex team be open to reviewing the persistence-filter approach?
If changing the persistent SQLite default from TRACE to WARN is consistent with the intended feedback and diagnostics behavior, I would appreciate guidance on whether the team would invite a focused upstream PR limited to:
codex-rs/state/src/log_db.rs;- the SQLite sink regression test;
- any required documentation or configuration updates.
I’m also happy to adjust the proposed target-specific policy if there are operational events that the feedback pipeline must continue retaining below WARN.
- change the persistent SQLite sink’s default level from
I can still reproduce significant persistent SQLite TRACE churn on the current ChatGPT-integrated Codex desktop build, and I think this deserves clarification after #28224 was closed.
Environment
- Platform: macOS / Apple Silicon
- App:
/Applications/ChatGPT.app - ChatGPT.app version:
26.707.72221 - Build:
5307 - Bundled Codex CLI:
codex-cli 0.144.2 - Primary DB:
~/.codex/logs_2.sqlite - Also observed project-scoped
CODEX_HOMEDBs such asclient-hub-codex-home/logs_2.sqlite
Controlled test
I had been using a local SQLite stopgap trigger:
CREATE TRIGGER IF NOT EXISTS block_log_inserts BEFORE INSERT ON logs BEGIN SELECT RAISE(IGNORE); END;
With the trigger enabled, short samples show:
id_delta=0count_delta=0- WAL stable
I then temporarily removed the trigger for a controlled test and restored it immediately afterward.
Result: within roughly a one-minute active Codex window, the global
~/.codex/logs_2.sqliteaccumulated about 2.6k rows and the WAL grew to about 4.15 MiB.Dominant persisted targets were still TRACE/DEBUG-style internal plumbing:
TRACE codex_api::sse::responsesTRACE codex_mcp::connection_managerTRACE codex_app_server::outgoing_messageTRACE hyper_util::client::legacy::*DEBUG codex_core::stream_events_utils
After restoring the trigger, another short sample returned to:
id_delta=0count_delta=0- WAL stable
Why I am posting
#28224 was closed as fixed, but the current shipping desktop app on my machine still bundles
codex-cli 0.144.2. From the discussion around #28224, it sounds like the more relevant logging fixes (#31789, #31790, #31791, #31792) are expected in the0.145.0-alpha.4+line, not necessarily in this bundled desktop version.So I would like to clarify:
- Is this residual persistent TRACE/DEBUG SQLite churn considered expected behavior?
- Which desktop app / bundled CLI version should include the full fix for Stop persisting RMCP service traces #31789-Summarize streamed response item logs #31792?
- Is there an official threshold for acceptable idle/active Codex local disk writes?
- Should users keep local mitigations like
block_log_insertsuntil the bundled desktop CLI reaches the fixed line? - Can OpenAI add an I/O regression test for idle and active Codex usage, so this class of issue does not regress silently?
I would also encourage other affected users to share comparable measurements with:
- app version/build
- bundled CLI version
- platform
- sample duration
MAX(id)delta- retained row delta
- WAL byte delta
- top persisted log targets
Fresh measurements on codex-cli 0.144.3 (Linux/ext4), plus a minimal tested patch and an A/B benchmark. TL;DR: most of the remaining churn is flush policy, not log content — batching the sink by bytes/severity instead of every 2 s cuts its physical writes ~5x on a steady-state workload, down to SQLite's WAL+checkpoint double-write floor, with zero change to what
/feedbacksees.Measurements: the writes are flush overhead, not content
Four resumed TUI sessions left idle; write rates from
/proc/<pid>/iodeltas, content attribution from queryinglogs_2.sqlite(read-only) for rows inserted inside the same instrumented window:metric value physical writes in a 59 s idle window (2 busiest processes) 28.2 MB log rows retained in that same window 630 rows / 774 KB physical-to-retained write amplification ~36x combined idle write rate, 4 sessions ~353 KB/s ≈ 30 GB/day one idle session running 9.3 days 123 GB lifetime writes The DB itself stays at ~25 MB — retention caps work. The wear comes from the durability policy:
LOG_FLUSH_INTERVAL = 2smeans ~43,000 transactions/day/process, each paying WAL page images + (at retention-cap steady state, which every long-lived session reaches) the windowed prune DELETE inside the same transaction.Row-level attribution also shows why target muting (#29457, #29599, #31789, #31791) can't converge: after accounting for those mutes, idle traffic is dominated by codex's own targets —
codex_app_server::outgoing_message(TRACE),codex_api::sse::responses(TRACE),codex_core::stream_events_utils(DEBUG),feedback_tags(INFO). Muting those defeats/feedback; keeping them keeps the churn. The reasoning that closed #23275 ("bounding individual call sites is brittle and misses future tracing events") applies equally to per-target filters.Patch: flush by bytes/severity/60 s instead of every 2 s
Branch: https://github.com/friday-james/codex-1/tree/fix/log-db-write-churn — 3 files, +127/−9, mostly tests.
- Flush the sink buffer on any of: ~1 MiB of estimated pending content, a 60 s tick, or a WARN/ERROR entry (flushed immediately, so crash-relevant diagnostics still reach disk right away).
/feedbackalready callslog_db.flush().awaitbefore reading (feedback_processor.rs), so feedback content is byte-identical, even mid-session.LogEntry::estimated_bytes()is factored out and shared by the sink buffer and the insert path, so the write-side threshold and partition budgets agree.- Acknowledged trade-off: on SIGKILL, up to 60 s of TRACE/DEBUG tail can be lost (WARN/ERROR are not affected, and any explicit flush path is not affected). Today's policy loses up to 2 s.
A/B benchmark (each phase an isolated process)
Harness: https://github.com/friday-james/codex-1/tree/bench/log-db-write-churn (
state/examples/write_churn_bench.rs). It replays the measured idle profile — 5.3 rows/s, ~1.2 KiB/row, thread partition pre-filled past its retention cap (steady state of a long-lived session) — for 3 minutes per policy, measuring/proc/self/iowrite_bytes. Two runs, results identical to ±0.2%:old policy (2 s / 128 rows): 14.5 MB in 180 s -> 7.02 GB/day new policy (60 s / 1 MiB / sev): 2.9 MB in 180 s -> 1.42 GB/day (~4.9x less)The new policy's 2.9 MB against 1.17 MB of raw content is SQLite's WAL-then-checkpoint double write plus prune/B-tree overhead — i.e. the patch removes nearly all removable churn; going lower requires persisting less content or relaxing durability guarantees. Note the benchmark is single-process and therefore conservative: live sessions share one
logs_2.sqliteacross processes, where checkpoint contention pushes the old policy 2–3x higher than the bench (and in the limit produces the unbounded-WAL growth reported in #28997).Happy to open this as a PR per the invitation policy in contributing.md, or the approach is free to take as-is. Two further options if there's interest: prune hysteresis (once a partition crosses its cap, delete down to ~75% so steady-state inserts don't re-run the windowed DELETE every batch), and honoring an env override for the sink filter, which several reports expected
RUST_LOGto do (#29463, #31111).Request for clarification: current shipping build still lacks the broader SQLite fix
Following up with the current shipping macOS build:
- ChatGPT/Codex app:
26.707.91948, build5440 - Bundled CLI:
codex-cli 0.144.5 - PATH CLI:
codex-cli 0.144.3 - npm stable:
0.144.5
The
rust-v0.144.5release contains only the dangerous-command detection fix in #33455. Its persistent log filter still uses:Targets::new().with_default(LevelFilter::TRACE)
Source: https://github.com/openai/codex/blob/rust-v0.144.5/codex-rs/state/src/log_db.rs
The
0.144.4...0.144.5release comparison also does not include #31789, #31790, #31791, or #31792. This appears inconsistent with the statement in #28224 that the series of fixes was already in the latest shipping desktop app, CLI, and IDE extension.My local SQLite database is currently protected by this reversible stopgap:
CREATE TRIGGER block_log_inserts BEFORE INSERT ON logs BEGIN SELECT RAISE(IGNORE); END;
Protected-state verification on the current build shows
rows=0,seq=0, andmax(id)=0. I did not remove the trigger for another raw sample because the shipping0.144.5source contains no relevant persistent-log change. I am not posting any rawfeedback_log_bodyvalues because they may contain private paths, command fragments, tool output, or conversation context.Could the maintainers please clarify:
- Is
LevelFilter::TRACEstill the intended production default for the persistent SQLite feedback log? - Which exact stable CLI version and desktop build will contain the complete fix, including the relevant target/callsite reductions and any write-flush policy change?
- Is there a planned ETA or release line for that fix?
- Until that version ships, should affected users continue using a
BEFORE INSERT ... RAISE(IGNORE)trigger? - If the trigger is not recommended, what is the officially supported mitigation that prevents excessive writes without losing essential
/feedbackdiagnostics? - Can a regression test or documented acceptable idle/active write threshold be added so users can verify the fix?
A version-specific answer would help users know when it is safe to remove the trigger and restore normal persistent diagnostic logging.
Reacted by CoreJa- ChatGPT/Codex app:
Follow-up after updating this Mac to the current ChatGPT-integrated Codex build.
Environment now:
- Platform: macOS
27.0(26A5378n), Apple Silicon /arm64 - App:
/Applications/ChatGPT.app - ChatGPT app:
26.715.31925, build5551 codexresolves to:/Applications/ChatGPT.app/Contents/Resources/codex- CLI:
codex-cli 0.145.0-alpha.18 - Old standalone
/Applications/Codex.appis not present on this machine RUST_LOGis unset
I kept the local SQLite mitigation in place for this check. The active triggers are still:
sample_log_inserts trim_sampled_logsSo this is a protected-state verification only; I did not remove the mitigation for another raw write-churn sample, and I am not posting raw
feedback_log_bodyvalues publicly.Current protected-state sample from
~/.codex/logs_2.sqlite:quick_check: ok journal_mode: wal rows: 500 TRACE rows: 104 DEBUG rows: 128 INFO rows: 259 WARN rows: 9 ERROR rows: 0 sqlite size: 134,221,824 bytes wal size: 4,194,192 bytes page_count: 32769 freelist_count: 32650 page_size: 409630-second retained-row sample with the mitigation still enabled:
max(id): 240689 -> 240692 (+3) rows: 500 -> 500 TRACE rows: 105 -> 104 max TRACE id: 240689 -> 240689 sqlite size: unchanged at 134,221,824 bytes wal size: unchanged at 4,194,192 bytesNewest retained rows in that sample were
DEBUG codex_http_client::default_clientandINFO codex_http_client::custom_ca; the newest retained TRACE row stayed at id240689(TRACE codex_app_server::message_processor). The retained sample still includes TRACE entries fromcodex_app_server::message_processor,codex_api::sse::responses, andcodex_app_server::outgoing_message, but I cannot infer the unmitigated raw insert rate from this protected sample.Given #4988716077 and #4989141674, this updated app/CLI datapoint may help separate two questions:
- whether the 0.145.x ChatGPT-integrated build has the target/callsite reductions expected from the 0.145 alpha line;
- whether the broader SQLite sink write/flush policy issue remains separate from the target filtering work.
A version-specific maintainer answer would still be useful: is
26.715.31925/ build5551/codex-cli 0.145.0-alpha.18expected to be safe enough for affected users to remove local SQLite mitigations, or should users keep them until a stable desktop/CLI release with an explicit persistent-log fix ships?- Platform: macOS
Hey, I had a look at this.
The SQLite sink already has a guard that drops bridged events with the
logtarget before they reach the queue. The missing piece was a test that proves the guard still works when the outer tracing filter is permissive, which is the path that allowed this noise through in the first place.I've added that regression test. It sends a bridged
target="log"event alongside a normal Codex state event, flushes the sink, and checks that only the normal event was written to SQLite.Branch: https://github.com/TheJhyeFactor/codex/tree/fix/sqlite-log-bridge-regression-29532
I ran
CARGO_NET_OFFLINE=true cargo test -p codex-statelocally on macOS. All 153 unit tests, 3 binary tests, and 1 doctest passed.I couldn't open the PR because GitHub returned a
CreatePullRequestpermission error for my account, so I've linked the tested branch here for review.Follow-up after another ChatGPT-integrated Codex update on the same Mac.
Environment now:
- ChatGPT app:
26.715.72359, build5718 - Embedded CLI:
codex-cli 0.145.0-alpha.30 codexstill resolves through/Applications/ChatGPT.app/Contents/Resources/codex- The old standalone
/Applications/Codex.appis still not present RUST_LOGis unset
I kept the local SQLite mitigation enabled for this check, so this is still only a protected-state verification. I did not remove the mitigation for an unbounded raw write-churn sample, and I am not posting raw
feedback_log_bodyvalues or detailed local SQLite metadata publicly.Qualitatively, in the protected sample I no longer see bridged
TRACE target=logamong the newest/top retained entries, which seems consistent with the regression-test discussion in #5018554274. However, normal Codex TRACE targets are still present in the retained sample, and this protected check cannot answer the broader SQLite sink write/flush-policy question raised in #4988716077.Could maintainers clarify two version-specific points for affected users?
- Is
26.715.72359/ build5718/codex-cli 0.145.0-alpha.30expected to contain the relevant macOS persistent-log target filtering fix? - Separately, is there any shipped or planned fix for the SQLite sink write/flush-policy churn, or should users keep local SQLite mitigations until a stable release explicitly mentions that broader fix?
Reacted by Jhye O'Meley- ChatGPT app:
Analysis (community)
Thanks for tracking this — attempting a residual-focused analysis aligned with docs/contributing.md (invitation-only: analysis in-thread, no unsolicited PR).
- Issue under review: macOS: Persistent SQLite TRACE target=log churn remains after rust-v0.142.0 #29532 — macOS: Persistent SQLite TRACE target=log churn remains after rust-v0.142.0
- Theme / radar cluster: residual TRACE write after claimed fix (
logs_2.sqlite SSD) - Reporter posture: static + local evidence where available; not claiming a maintainer invite.
Observed (from issue + public residual signal)
Summary After upgrading to
rust-v0.142.0, I can still reproduce persistent SQLite log churn on macOS in~/.codex/logs_2.sqlite. This looks like a partial fix: -#29432appears to have helped:codex_api::endpoint::responses_websocketis much lower than before. - However,#29457does not appear to fully solve the persistent SQLite churn in this app build/runtime. The release note says noisy persistent log targets were filtered, and the PR says bridgedtarget=logevents should be excLocal residual measurement (2026-07-24, this reporter)
On a host running codex-cli 0.146.0-alpha.6 (newer than 0.142–0.145 residual reports):
Signal-14 INTERIM evidence (from qwen_s14 mid-flight)
captured: 2026-07-24T15:10:05Z
/Users/bing/.codex/logs_2.sqlitesize=1778548736 mtime=Fri Jul 24 23:09:38 2026/Users/bing/.codex/logs_2.sqlite-walsize=30405632 mtime=Fri Jul 24 23:10:04 2026- codex version:
codex-cli 0.146.0-alpha.6
level counts (ro)
TRACE|94794 DEBUG|21528 INFO|19618 WARN|796 ERROR|370rows/max_ts: 137106|2026-07-24 23:10:05
This is not a claim that every install still burns TB/year — it is a version-pinned observation that TRACE still lands in the SQLite sink after the #28224-era mitigations on a current alpha, consistent with residual issues still open.
Root-cause hypothesis (to be confirmed on
main)- Feedback/log pipeline still admits TRACE (or default level too verbose) into
logs_2.sqlite. - WAL/checkpoint policy does not cap growth under idle or streaming load.
- Partial cut of TRACE paths may leave high-volume targets writing.
High-level fix outline (not a PR)
- Default sink level: never store TRACE in SQLite unless explicitly enabled.
- Honor
RUST_LOG/ app log level end-to-end into the SQLite writer. - Hard caps: max DB size / WAL size / rotation with user-visible warning.
- Fail-first tests: write-rate budget under idle + streaming fixtures.
Tests
- Unit: writer rejects TRACE when configured level is INFO.
- Integration: after N seconds idle with default config, row growth rate < bound.
Questions for maintainers
- Is the intended long-term policy documented for this area (especially sandbox gitdir / log TRACE defaults)?
- Preferred residual anchor issue if several open threads describe the same mechanism?
- If the high-level outline above is roughly aligned, would an invited PR with failing-first tests be welcome later — or is an internal fix preferred?
Happy to refine with more measurements or code pointers. No unsolicited PR from me.
Analysis (community)
Residual analysis for #29532 (TRACE churn remains after rust-v0.142+ claims) — invitation-only under docs/contributing.md. No unsolicited PR.
Distinct from #17320 (default TRACE policy) and #28997 (WAL budget): this thread is about fix claims vs still-shipping residual on current builds.
Deterministic evidence (code + runtime)
Code (
5dd992acd3):codex-rs/state/src/log_db.rsdefault_filter():.with_default(LevelFilter::TRACE)(L55)- only a short denylist of noisy targets (
hyper_util→WARN,log→OFF, a fewcodex_otel*/codex_api::responses_websocket_timingOFF,rmcp::service→INFO)
- High-fanout product targets are not demoted: SSE responses, app-server outgoing messages, TUI markdown stream still admit TRACE into the SQLite sink via
LogDbLayer::on_event. - Unit tests in
codex-rs/state/src/log_db_filter_tests.rsstill document TRACE retention for non-filtered targets.
Runtime (
codex-cli 0.146.0-alpha.6):logs_2.sqlite≈ 1778548736 bytes; TRACE rows 94907 of ≈137150 total.- Top TRACE targets on this host:
codex_api::sse::responses→ 37252codex_app_server::outgoing_message→ 19558codex_tui::markdown_stream→ 15085
- That distribution shows residual TRACE is product-path, not only third-party crates already WARN-filtered.
Root-cause hypothesis
Partial target denylists / “we fixed TRACE spam” claims did not change the default LevelFilter::TRACE for the SQLite sink, nor demote the highest-volume first-party streaming targets. Users on 0.146.x still observe multi-GB DBs and TRACE-dominant histograms.
High-level fix outline (not a PR)
- Ship default sink level INFO (or DEBUG) with explicit TRACE opt-in.
- Demote first-party high-fanout targets (
codex_api::sse::responses,codex_tui::markdown_stream,codex_app_server::outgoing_message) even if global default stays DEBUG. - Release note matrix: version → default sink level → measured TRACE share under a 5-minute stream repro.
- Fail-first: assert default filter does not retain generic TRACE for those three targets.
Questions for maintainers
- Which release was supposed to close the residual claimed on macOS: Persistent SQLite TRACE target=log churn remains after rust-v0.142.0 #29532, and against which default?
- Preferred anchor among macOS: Persistent SQLite TRACE target=log churn remains after rust-v0.142.0 #29532 / Excessive SQLite WAL writes during streaming due to TRACE logs ignoring RUST_LOG #17320 / logs_2.sqlite-wal grows without bound into tens of GB #28997 for the “still TRACE by default on 0.146.x” proof?
- Would an invited PR for default-level + first-party demotions with tests be useful?
No unsolicited PR from me.
Reacted by Jhye O'MeleyAdditional current-stable reproduction and bounded mitigation result (2026-08-15):
- Environment: macOS 26.6.1 (arm64), official stable
codex-cli 0.147.0, persistent app-server. - A new sanitized macOS Disk Writes diagnostic for that exact 0.147.0 process reports 8,589.94 MB over 25,456 seconds (337.44 KB/s versus a 99.42 KB/s threshold), with representative stacks dominated by SQLite
pwrite, WAL checkpoint, andfsync. A second current-version report records 2,147.49 MB over 44,782 seconds (47.95 KB/s). - Source inspection at
rust-v0.147.0still shows the dedicated log DB using its own default TRACE filter. The current documented configuration has no setting that independently raises or disables this SQLite sink;RUST_LOGcontrols a different layer. - Updating alone is therefore not a discriminator here: the writer and latest stable release are both 0.147.0.
As a local stop-loss only, after stopping and independently verifying zero Codex handles on the DB/WAL/SHM, I installed one persistent
BEFORE INSERTtrigger on the dedicatedlogstable. It usesRAISE(IGNORE)only whenupper(NEW.level) IN ('TRACE', 'DEBUG'); INFO/WARN/ERROR and all existing rows remain intact. No database was deleted, truncated, vacuumed, checkpointed, relocated, or read for content. Rollback is the exact namedDROP TRIGGERoperation.An isolated truth-table/rollback test and live schema plus
quick_checkacceptance passed. After restart through the official managed daemon, a representative 30-second active sample wrote 217,088 bytes (7.24 KB/s), below the smaller report's 24.86 KB/s threshold and materially below both red rates. WAL metadata still changes because higher-severity diagnostics remain enabled.This is not presented as an upstream fix or maintainer-endorsed workaround. It may disappear after a DB migration/recreation. A supported persisted log-level control (or an INFO-or-higher sink default), with migration behavior documented, would remove the need for this schema-level containment.
- Environment: macOS 26.6.1 (arm64), official stable
Summary
After upgrading to
rust-v0.142.0, I can still reproduce persistent SQLite log churn on macOS in~/.codex/logs_2.sqlite.This looks like a partial fix:
#29432appears to have helped:codex_api::endpoint::responses_websocketis much lower than before.#29457does not appear to fully solve the persistent SQLite churn in this app build/runtime. The release note says noisy persistent log targets were filtered, and the PR says bridgedtarget=logevents should be excluded from the SQLite sink, butTRACE target=logis still the dominant persisted target here.I also posted an earlier data point in the closed umbrella issue
#28224here:#28224 (comment)
Opening a new issue because
#28224is closed and this is post-rust-v0.142.0macOS data.Environment
26.616.71553, build4265codex-cli 0.142.0codex-cli 0.142.0/Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled7763, startedTue Jun 23 10:17:10 2026RUST_LOG: empty~/.codex/config.toml:[analytics] enabled = false0logs_2.sqlitejournal mode:walNo local
block_log_insertstrigger was present during this sample.Files observed
At the start of the sample:
The active app-server process was holding the DB/WAL/SHM files open.
Read-only 60 second sample
Sample window:
SQLite metadata before/after:
So the retained row count stayed flat, but
sqlite_sequence/max_idadvanced by1,470in about 60 seconds. This looks like continued insert-and-prune churn.File mtimes also moved during the sample:
The WAL size stayed stable at about 4.4 MB, so file size alone hides the ongoing write activity.
Recent persisted targets
Rows in the last 60 seconds:
TRACE targets in the last 5 minutes:
Expected behavior
After
rust-v0.142.0, Codex should not continuously persist high-frequencyTRACE target=logrows intologs_2.sqliteby default, especially after#29457.Actual behavior
Codex app-server continued to insert high-frequency
TRACE target=logrows into the persistent SQLite log sink. The retained row count stayed flat, butmax(id)/sqlite_sequencekept advancing and the DB/WAL mtimes kept moving, indicating ongoing insert-and-prune churn.Privacy note
I am intentionally not pasting raw
feedback_log_bodycontents because they may contain private conversation, local path, tool, or response data.