Repository navigation
Codex SQLite feedback logs can write ~640 TB/year and rapidly consume SSD endurance #28224
Description
Activity
- addedbugSomething isn't workingSomething isn't workingCLIIssues related to the Codex CLIIssues related to the Codex CLI
on Jun 14, 2026 github-actions commented
on Jun 14, 2026 on Jun 14, 2026 – with GitHub ActionsContributorMore actionsPotential duplicates detected. Please review them and close your issue if it is a duplicate.
- Codex Desktop app-server leaks child processes and writes excessive logs after crash/restart #26869
- Desktop launch can fail when logs_2.sqlite grows large: app-server SQLite pool times out during startup #27741
- Cold start scales badly with large ~/.codex state: 419 MB logs DB + 625 MB sessions cause 1–5 min startup on WSL2 #28166
Powered by Codex Action
cc @charley-oai @jif-oai since #12969 introduced the SQLite feedback log sink at TRACE level.
I pushed a minimal branch here:
https://github.com/1996fanrui/codex/tree/codex/reduce-sqlite-feedback-log-writesThe proposed patch does not disable feedback logs. It only changes the SQLite feedback log sink from global TRACE to this default filter:
warn,codex_api=info,codex_app_server=info,codex_core=info,codex_mcp=info,codex_state=info,codex_tui=infoIn my local sample, the retained SQLite log content was mostly TRACE plus
codex_otelINFO mirror rows, so this should drop the high-volume low-signal writes while keeping INFO+ logs for core Codex targets.I tried to open an upstream PR, but GitHub returned:
GraphQL: 1996fanrui does not have the correct permissions to execute CreatePullRequestOther possible fixes if the default-filter approach is not the preferred direction:
- add a config escape hatch, for example
sqlite_logs_enabled = falseor a feedback-log-specific filter setting; - keep SQLite feedback logs opt-in or sampled instead of always-on at TRACE;
- truncate or summarize large websocket/SSE/raw event payloads before storing them;
- enforce a global write/size budget in addition to per-thread retention.
Reacted by Ali Sherief, Maurycy Pawłowski-Wieroński, Ville Seppänen, Ev0lv3nta, f1refa11, Wesley Piapia, Eric McDonald, Rob and Ignat Remizov- add a config escape hatch, for example
This is a very serious bug, as it can not only cause login failures on Linux boxes if the system is rebooted while the disk is full, but Codex in
/goalmode will actively delete files and folders on your disk in a vain attempt to gain disk space. So the severity of this issue needs to be increased as data loss is possible when using Codex in this way.Reacted by Julio M, Maurycy Pawłowski-Wieroński, Bodya, kkocdko, Jonathan Froehlich, Jacques Germishuys, Emilie Rollandin, Marco Raddatz, Brian Chew, singhn-ymmd and 71 moreWTF this is ridiculous. PLEASE have more respect for your users.
Reacted by Hongbo Miao, imi, Bin Xie, Isaac, stqfdyr, twiceYuan, taosimple2, Vonfry, Yongpeng Lin, Arendelle and 250 moreI prepared a PR-ready branch against current
main, but upstream PR creation is still blocked for my account with:GraphQL: Nicolas0315 does not have the correct permissions to execute `CreatePullRequest`Branch:
https://github.com/Nicolas0315/codex/tree/codex/log-db-default-filterProposed PR title:
[codex] Limit persisted SQLite log verbosityProposed PR body, following the recent Codex PR format:
Why
logs_2.sqliteis currently attached with a globalTRACEfilter in both the app server and TUI. On my Windows Codex app install, that persisted high-volume websocket internals and OTel mirror records into the diagnostic SQLite log DB:.codex/logs_2.sqlite: 729,575,424 bytes.codex/logs_2.sqlite-wal: 39,255,392 bytes- retained rows: 269,141
TRACE: 123,026 rows / 269,065,577 estimated bytes- top persisted targets included
codex_api::endpoint::responses_websocket,codex_otel.log_only,codex_otel.trace_safe, andlog
This matches the write-amplification report here and the Desktop launch failure mode in #27741. The change keeps diagnostic logs enabled, but makes the persisted SQLite sink less noisy by default.
What changed
- Add a shared
log_db::default_filter()for the persistent SQLite log sink. - Persist
INFO+by default instead ofTRACE. - Raise known high-volume mirror/dependency targets to
WARN+:codex_otel.log_onlycodex_otel.trace_safehyper_util::client::legacy::*logopentelemetry_sdk
- Use the shared filter from both app-server and TUI log DB setup.
- Update the SQLite sink filter test to cover retained Codex
INFOlogs and dropped low-value OTel/mirror logs.
Validation
git diff --checkjust test -p codex-state sqlite_sink_default_filter_drops_low_value_logsjust test -p codex-tui sqlite_sink_default_filter_drops_low_value_logscompiled the package, then exited with no matching tests.just fmtcompleted Rust/Python formatting, then failed on Windows in the Bazel buildifier step with[WinError 2].- A broader
just test -p codex-app-serverrun compiled and executed tests but timed out after reporting831 passed, 4 failed, 6 skipped; the visible failures appeared unrelated to this log filtering change, including a skills fixture warning count mismatch.
Reacted by Hongbo Miao, Ville Seppänen, Carter Stach, Thorsten Sommer, 1gaojer, ChetanPatteparapu1, Sabit Zahin, Majimay, Jirafa, Robert Rollins and 2 moreReacted by yumi and AslanHere I have made two scripts which can serve as temporary mitigations for the WAL issue until OpenAI fixes it.
The first script trims the WAL while Codex processes are running:
#!/usr/bin/env bash # trim-codex-wal.sh # Trims the logs_2.sqlite and WAL files by truncating the checkpoint database set -u CODEX_DIR="${CODEX_HOME:-$HOME/.codex}" DB="$CODEX_DIR/logs_2.sqlite" WAL="$CODEX_DIR/logs_2.sqlite-wal" SHM="$CODEX_DIR/logs_2.sqlite-shm" bytes() { if [ -e "$1" ]; then stat -c '%s' "$1" 2>/dev/null || stat -f '%z' "$1" 2>/dev/null || echo 0 else echo 0 fi } human() { if command -v numfmt >/dev/null 2>&1; then numfmt --to=iec --suffix=B "$1" 2>/dev/null || echo "$1 bytes" else echo "$1 bytes" fi } show_sizes() { echo echo "Codex SQLite files:" for f in "$DB" "$WAL" "$SHM"; do printf " %-45s %12s\n" "$f" "$(human "$(bytes "$f")")" done } echo "[LIVE RUN] Codex WAL truncate only" echo "Codex dir: $CODEX_DIR" echo "DB: $DB" echo "WAL: $WAL" echo "SHM: $SHM" if [ ! -f "$DB" ]; then echo echo "ERROR: Codex DB not found:" echo " $DB" exit 1 fi if ! command -v sqlite3 >/dev/null 2>&1; then echo echo "ERROR: sqlite3 is not installed or not in PATH." exit 1 fi show_sizes BEFORE_WAL="$(bytes "$WAL")" echo echo "Running SQLite WAL checkpoint/truncate..." echo "Command:" echo " PRAGMA wal_checkpoint(TRUNCATE);" RESULT="$( sqlite3 "$DB" <<'SQL' 2>&1 PRAGMA busy_timeout = 5000; PRAGMA wal_checkpoint(TRUNCATE); SQL )" RC=$? AFTER_WAL="$(bytes "$WAL")" echo echo "SQLite output:" if [ -n "$RESULT" ]; then printf '%s\n' "$RESULT" else echo "(no output)" fi show_sizes echo echo "Result:" echo " WAL before: $(human "$BEFORE_WAL")" echo " WAL after: $(human "$AFTER_WAL")" if [ "$RC" -ne 0 ]; then echo echo "Failed: sqlite3 returned exit code $RC." echo "Likely causes: DB is busy, Codex is actively writing, or disk is critically full." exit "$RC" fi if [ "$AFTER_WAL" -lt "$BEFORE_WAL" ]; then echo echo "Success: Codex WAL was truncated." elif [ "$AFTER_WAL" -eq "$BEFORE_WAL" ]; then echo echo "No size change." echo "Possible reasons:" echo " - WAL was already small/checkpointed." echo " - Codex or another process currently has an active transaction." echo " - Codex is still writing and recreated WAL content immediately." else echo echo "WAL grew during the run, probably because Codex was actively writing." fi echo echo "Done."To run as a cron job:
# Run `crontab -e` in your shell and then add the following line at the end */15 * * * * /bin/bash "$HOME/trim-codex-wal.sh" > "$HOME/.codex-wal-clean.log" 2>&1Note that the main table is trimmed, but there might be other tables which still accumulate data over a much slower period of time, requiring the script below.
The second script deletes the log and WAL files and kills all running Codex processes in order to immediately free disk space:
#!/usr/bin/env bash # fix-codex-wal.sh # Deletes the logs-2.sqlite* files and sends SIGTERM followed by SIGKILL to all codex-related processes # in order to close the open file handles which will allow disk space to be freed set -u CODEX_DIR="${CODEX_HOME:-$HOME/.codex}" DB="$CODEX_DIR/logs_2.sqlite" WAL="$CODEX_DIR/logs_2.sqlite-wal" SHM="$CODEX_DIR/logs_2.sqlite-shm" PIDS="" add_pid() { case "${1:-}" in ''|*[!0-9]*) return ;; "$$") return ;; esac case " $PIDS " in *" $1 "*) ;; *) PIDS="$PIDS $1" ;; esac } bytes() { if [ -e "$1" ]; then stat -c '%s' "$1" 2>/dev/null || stat -f '%z' "$1" 2>/dev/null || echo 0 else echo 0 fi } human() { if command -v numfmt >/dev/null 2>&1; then numfmt --to=iec --suffix=B "$1" 2>/dev/null || echo "$1 bytes" else echo "$1 bytes" fi } show_sizes() { echo echo "Codex SQLite files:" for f in "$DB" "$WAL" "$SHM"; do printf " %-45s %12s\n" "$f" "$(human "$(bytes "$f")")" done } show_processes() { if [ -z "$(printf '%s' "$PIDS" | tr -d ' ')" ]; then echo "None." return fi printf "%-8s %-8s %-6s %-10s %-12s %s\n" "PID" "PPID" "STAT" "TTY" "ELAPSED" "CMD" for pid in $PIDS; do ps -p "$pid" -o pid=,ppid=,stat=,tty=,etime=,cmd= 2>/dev/null done } echo "[LIVE RUN] Codex WAL/SHM cleanup" echo "Codex dir: $CODEX_DIR" echo "DB: $DB" echo "WAL: $WAL" echo "SHM: $SHM" show_sizes echo echo "[1/5] Finding Codex-related processes..." for pid in $(pgrep -u "$USER" -x codex 2>/dev/null); do add_pid "$pid" done for pid in $(pgrep -u "$USER" -x node 2>/dev/null); do cmd="$(ps -p "$pid" -o cmd= 2>/dev/null || true)" case "$cmd" in *codex*|*.codex*|*openai*codex*) add_pid "$pid" ;; esac done echo echo "[2/5] Finding processes holding Codex WAL/SHM files..." if command -v lsof >/dev/null 2>&1; then for f in "$WAL" "$SHM"; do if [ -e "$f" ]; then for pid in $(lsof -t -- "$f" 2>/dev/null); do add_pid "$pid" done fi done for pid in $( lsof -nP +L1 2>/dev/null | awk -v dir="$CODEX_DIR" ' index($0, dir "/logs_2.sqlite-wal") && /deleted/ { print $2 } index($0, dir "/logs_2.sqlite-shm") && /deleted/ { print $2 } ' ); do add_pid "$pid" done else echo "lsof not found; cannot check open/deleted WAL handles." fi echo echo "[3/5] Processes to terminate:" show_processes if [ -n "$(printf '%s' "$PIDS" | tr -d ' ')" ]; then echo echo "Sending SIGTERM..." for pid in $PIDS; do echo "TERM $pid: $(ps -p "$pid" -o cmd= 2>/dev/null || echo unknown)" kill "$pid" 2>/dev/null || true done sleep 3 STILL="" for pid in $PIDS; do if kill -0 "$pid" 2>/dev/null; then STILL="$STILL $pid" fi done if [ -n "$(printf '%s' "$STILL" | tr -d ' ')" ]; then echo echo "Some processes survived SIGTERM; sending SIGKILL:" printf "%-8s %-8s %-6s %-10s %-12s %s\n" "PID" "PPID" "STAT" "TTY" "ELAPSED" "CMD" for pid in $STILL; do ps -p "$pid" -o pid=,ppid=,stat=,tty=,etime=,cmd= 2>/dev/null || true done for pid in $STILL; do echo "KILL $pid: $(ps -p "$pid" -o cmd= 2>/dev/null || echo unknown)" kill -9 "$pid" 2>/dev/null || true done else echo "All targeted processes exited after SIGTERM." fi fi echo echo "[4/5] Deleting Codex WAL/SHM files..." if [ -e "$WAL" ]; then echo "Deleting: $WAL" rm -f -- "$WAL" else echo "Not found: $WAL" fi if [ -e "$SHM" ]; then echo "Deleting: $SHM" rm -f -- "$SHM" else echo "Not found: $SHM" fi show_sizes echo echo "[5/5] Final cleanup check..." if command -v lsof >/dev/null 2>&1; then echo "Remaining deleted Codex files still held open, if any:" lsof -nP +L1 2>/dev/null | grep -F "$CODEX_DIR" || echo "None found." else echo "lsof not found; skipping deleted-handle check." fi echo echo "Disk usage:" df -h "$HOME" 2>/dev/null || true du -xsh "$CODEX_DIR" 2>/dev/null || true echo echo "Done."Reacted by Hongbo Miao, cwegener, BG4JEC, David Paluy, 求余, ifBars, ayamkv, Linus, tukuyomi032, Statusnone and 4 moreplz fix
Reacted by Lê Tấn Lộc, Maliki Kadafa Tsaqif, Jubilee and SavvaReacted by Ed Johnson-WilliamsThe log DB should have an explicit byte budget and retention policy. Durable usage/diagnostic events can stay queryable, while websocket/SSE TRACE payloads need sampling, level caps, and rotation independent of the state DB. A startup receipt with effective log level, DB byte cap, WAL checkpoint policy, and dropped-row count would make write pressure visible before SSD wear becomes the signal.
Generated with ax.
Additional native Windows check against current
main:- Upstream source inspected:
63f009e9dad2e70454b7ed6434d8aa28dfb52b51 - Windows 11 Pro 10.0.26200
- Codex Desktop active during the sample
Current source still attaches the persistent SQLite log layer with a global TRACE filter in both app-server and TUI:
codex-rs/app-server/src/lib.rslog_db::start.with_filter(Targets::new().with_default(Level::TRACE))
codex-rs/tui/src/lib.rslog_db::start.with_filter(Targets::new().with_default(Level::TRACE))
I also checked the current log sink implementation:
codex-rs/state/src/log_db.rs- events are formatted into
feedback_log_body, queued, then inserted by a background task. - the sink drops low-level
opentelemetry_sdkTRACE/DEBUG, which helps one specific noisy target.
- events are formatted into
codex-rs/state/src/runtime/logs.rs- each batch inserts rows, then prunes inside the same transaction.
That means DB size or retained row count can stay stable while insert/prune churn continues.
Local aggregate-only measurement from
.codex\logs_2.sqliteover 15 seconds, without reading log bodies:- retained rows before: 56,766
- retained rows after: 56,766
- retained row delta: 0
MAX(id)before: 191,131,613MAX(id)after: 191,131,862MAX(id)delta: 249 rows inserted in 15s- retained estimated bytes delta: +135,974 bytes
Top retained target+level groups in this Windows sample were still dominated by the same categories already reported here:
codex_api::endpoint::responses_websocketTRACEcodex_otel.log_onlyINFOcodex_otel.trace_safeINFOcodex_core::stream_events_utilsINFOcodex_client::transportTRACErmcp::serviceTRACEcodex_api::sse::responsesTRACElogTRACE
This reinforces the earlier point: visible DB growth is not a reliable proxy for write pressure. Even with a stable retained row count, the append-and-prune path is still doing work. The low-risk patch shape still looks like filtering before formatting/queueing/inserting, not only pruning after insert.
I also rebuilt a narrow local candidate patch against the current
origin/mainhead:- base:
f774455(code-mode: linearize cell terminal state (#29286)) - adds
codex_state::log_db::default_filter() - uses the shared filter from app-server and TUI instead of
Targets::new().with_default(Level::TRACE) - keeps global
WARN+ - keeps
INFO+for core Codex targets (codex_api,codex_app_server,codex_core,codex_mcp,codex_state,codex_tui) - raises known high-volume mirror/dependency targets to
WARN+(codex_otel.log_only,codex_otel.trace_safe,hyper_util,log,opentelemetry_sdk)
Local validation:
cargo test -p codex-state sqlite_sink_default_filter_drops_low_value_logs-> 1 passedcargo test -p codex-state log_db-> 7 passedcargo check -p codex-app-server -p codex-tui-> passedcargo fmt --check-> passed; stable rustfmt emitted repo-config warnings thatimports_granularityis nightly-onlygit diff --check-> passed
Per
docs/contributing.md, I am not opening a PR unless invited, but this latest local patch is PR-shaped if maintainers want this direction.- Upstream source inspected:
I hope the developers see this issue.
Reacted by yanghaku, Merrkry, Janos Dobronszki, ooeyuna, Maurycy Pawłowski-Wieroński, nitish-anekantvad, George Liu (eva2000), michael-heinrich, ilyes, Nikita and 27 moreShame on such low software quality with such high memory prices and SSD cost...
Reacted by Lê Tấn Lộc, Miyuru, Hikmet Altıntaş, Amirhossein Fakari, Muhammad Yaseen, Marco Raddatz, Simon Kocurek, Erkan Okyay, Asuka Minato, santiago-afonso and 11 moreReacted by I-Would-Like-To-Report-A-Bug-Please, Anirudh, Asuka Minato and milkii_ii151 remaining items
Load more actionsWhich version has it fixed? Is it in the desktop app?
Reacted by Glenn 'devalias' Grant and Alexander GusevOur series of fixes for this problem are in the latest shipping desktop app, CLI, and IDE extension.
Reacted by Ken WarnerReacted by waittting, Glenn 'devalias' Grant, Yiyang Lei and 小羽Our series of fixes for this problem are in the latest shipping desktop app, CLI, and IDE extension.
Not quite yet. 0.144.1 does not have #31789, #31790, #31791, and #31792. 0.145.0-alpha.4 includes them, so 0.145.0 likely will too.
Also, please see #32496. A different root cause is causing about 1.3 TB/year in unnecessary writes, according to the numbers in the issue. (Much less than was on this issue.) Codex is continually rewriting
models_cache.jsonandscope_v3.json.Would it be possible for OpenAI to add a regression test that ensures: (1) Codex during idle doesn't make more than a very minimal amount of I/O; and (2) Codex during usage doesn't absurdly increase its I/O? It might be best to split 2, into one version that fakes the model to test the harness in isolation, and one end-to-end test that makes sure new model/harness combinations don't come up with something that bumps it significantly. The latter could of course trip from random model behavior. If OpenAI logged and compared the values, they would not only see spikes but also worrisome trends in I/O usage that could be investigated.
While at it, testing CPU and GPU usage similarly would be a big help. There are several issues about crazy high usage, often tracked down to animations being unnecessarily repeatedly updated a crazy number of times a second.
Reacted by waittting, Joker_, Glenn 'devalias' Grant, Merrkry, sbbu, Ronald Gamez and Alexander Gusev@etraut-openai Codex has been tracking this issue for me. I asked its opinion (GPT 5.6 Sol High), and it doesn't think this is fully solved:
I retested the SQLite logging behavior on macOS. The upstream fixes have substantially reduced the original churn, but the current desktop build still persists noisy TRACE/DEBUG activity at a much higher rate than when the local filter is enabled.
Scenario App/engine Mitigation Duration MAX(id) change Rate Net retained rows Earlier affected build (July 1) codex-cli 0.142.5 Off 96.9s +3,212 33.1 IDs/s +57 Current desktop build (July 12) App 26.707.51957, bundled codex-cli 0.144.0-alpha.4 Off 41.45s +305 7.36 IDs/s +9 Same current build App 26.707.51957, bundled codex-cli 0.144.0-alpha.4 On 41.89s +19 0.45 IDs/s +1 The mitigation is a local SQLite trigger that rejects TRACE/DEBUG records and known noisy targets. It does not block INFO, WARN, or ERROR logging.
This shows two things:
- The upstream fixes improved the unmitigated rate substantially: 33.1 -> 7.36 IDs/second.
- On the current build, enabling the filter still reduces churn by about 16x: 7.36 -> 0.45 IDs/second.
During the current-build test without mitigation, 305 IDs were allocated while only 9 additional rows remained. That is consistent with continuing insert-and-prune activity.
The newly retained records were dominated by:
- codex_app_server::outgoing_message / TRACE: 133
- codex_api::sse::responses / TRACE: 35
- codex_mcp::connection_manager / TRACE: 30
- hyper_util TRACE/DEBUG: 67
No log bodies were read during these measurements.
Is this residual TRACE/DEBUG churn considered expected after the fixes? If not, is desktop app 26.707.51957 missing part of the fix, and which desktop or bundled CLI version should contain it?
Reacted by rodalpho, waittting, Joker_, Glenn 'devalias' Grant, YewFence, singhn-ymmd, agentHits and eimsIs this residual TRACE/DEBUG churn considered expected after the fixes? If not, is desktop app 26.707.51957 missing part of the fix, and which desktop or bundled CLI version should contain it?
@gustavopch It seems like you're still using
0.144.0-alpha.4in those tests; whereas an earlier comment from @darlingm suggests that there are more fixes that will land in0.145.0-alpha.4+, that sound like they should address some if not all of those:Not quite yet. 0.144.1 does not have #31789, #31790, #31791, and #31792. 0.145.0-alpha.4 includes them, so 0.145.0 likely will too.
Reacted by Gustavo and singhn-ymmdWe should be getting partial refunds or a non-expiring limit reset as compensation for this.
--@mtsitzerSo are we getting some refunds/limit resets or ANY compensation for killing our SSDs? For weeks or nah?
Reacted by singhn-ymmd, Yiyang Lei and Alexander Gusev@0xdevalias Ah, thanks for pointing it out. I just updated from
26.707.51957to26.707.61608and the bundled Codex CLI is still0.144.0-alpha.4. I guess we'll have to wait a few more days.Reacted by Glenn 'devalias' GrantSo has this issue been fixed in the latest version of the Codex desktop app?
Reacted by Alexander Gusev, an_dev, Elias, Ognjen Vasic, Amit Acharya, kyubuns, Vinicius Machado Rodrigues, mmcintyre123, a3502796980, scotej and 4 more- added a commit that references this issue
on Jul 26, 2026 - added a commit that references this issue
on Jul 28, 2026 - added a commit that references this issue
on Jul 28, 2026 - added a commit that references this issue
on Aug 5, 2026 - added a commit that references this issue
on Aug 10, 2026 - added a commit that references this issue
on Aug 19, 2026 - added a commit that references this issue
on Aug 24, 2026
Update at Jun 23, 2026: the following 3 PRs are merged, it could avoid 85% logs(feedback from my codex), so let me close this issue.
Thanks @jif-oai for the fix.
Simple Workaround from @beskay #28224 (comment)
Issue
Codex is continuously writing a large amount of data to the local SQLite feedback log database:
~/.codex/logs_2.sqlite~/.codex/logs_2.sqlite-wal~/.codex/logs_2.sqlite-shmOn my machine, after about 21 days of uptime, the main SSD has written about 37 TB. Process/file-level checks show Codex SQLite logs are the main continuous writer.
That extrapolates to roughly 640 TB/year. On a 1 TB SSD, that is about 640 full-drive writes per year. Some consumer SSDs are rated around 600 TBW, so this could consume roughly a full drive's warranted write endurance in less than a year.
Evidence1
A later snapshot makes the write amplification easier to see:
logs_2.sqlitefile sizeSo the database currently retains only ~0.5M rows, while the SQLite AUTOINCREMENT counter has already advanced past 5.5B ids.
That is roughly a 10,000x gap between retained rows and historical inserted row ids. Even using the current ~1.2 GiB database size as a rough baseline, this points to 10TB+ scale historical log churn, before accounting for WAL, indexes, pruning, checkpoints, page rewrites, and filesystem/device-level write amplification.
Evidence2
Current retained rows in
logs_2.sqlite:Level distribution:
Largest target+level pairs:
codex_api::endpoint::responses_websocketcodex_otel.log_onlycodex_otel.trace_safelogcodex_client::transportcodex_core::stream_events_utilscodex_api::sse::responsesThe top sources are mostly global TRACE logs, mirrored telemetry logs, and raw websocket/SSE payload logging.
TRACEalone is about 70.7% of retained bytes.codex_otel.log_only+codex_otel.trace_safeadd another 25.3%. Filtering these categories should remove roughly 96% of retained log bytes in this sample without fully disabling feedback logs.Sanitized examples from the most frequent TRACE source:
target=logThese are high-frequency retained samples. Raw websocket/SSE payload bodies are intentionally not included because they may contain private conversation content.
Sanitized examples from frequent INFO sources
The dominant INFO sources are mostly repeated OpenTelemetry mirror events. IDs are redacted.
Write amplification
The retained DB size hides the real write volume. In a 15-second sample:
About 36,211 rows were inserted in 15 seconds, while retained row count stayed flat. This suggests continuous insert-and-prune write amplification: rows are inserted, indexed, written to WAL, then pruned.
Likely cause
The SQLite feedback log sink is installed with a global TRACE default:
This persists all targets at TRACE level by default, including dependency/internal logs and large raw protocol payloads.
Proposed fix
Keep feedback logs enabled, but narrow what is persisted by default:
target=log,hyper_util, tokio-tungstenite internals, inotify spam, and low-level OpenTelemetry SDK logs.codex_otel.log_only/codex_otel.trace_safeevents unless they are explicitly useful for feedback debugging.An optional escape hatch such as
sqlite_logs_enabled = falsewould still be useful, but the main fix should be better default filtering.Related issues and discussions
logs_2.sqlite-walgrows indefinitely and remains allocated after deletion because stale/suspended Codex TUI processes keep the deleted WAL open #22444codexprocesses. #20563