Repository navigation
Desktop launch can fail when logs_2.sqlite grows large: app-server SQLite pool times out during startup #27741
Description
Activity
- addedbugSomething isn't workingSomething isn't workingappIssues related to the Codex desktop appIssues related to the Codex desktop appwindows-osIssues related to Codex on Windows systemsIssues related to Codex on Windows systemsapp-serverIssues involving app server protocol or interfacesIssues involving app server protocol or interfaces
on Jun 12, 2026 github-actions commented
on Jun 12, 2026 on Jun 12, 2026 – with GitHub ActionsContributorMore actionsPotential duplicates detected. Please review them and close your issue if it is a duplicate.
- app-server: feedback log sqlite (logs_N.sqlite) grows unbounded — ~0.75 GB/day, no retention/rotation #26374
- Desktop: "Error submitting message" — turn/start times out (30s) while app-server sidecar stalls silently for minutes #27395
Powered by Codex Action
MisterRound commented
on Jun 22, 2026 AuthorMore actionsUpdate from 2026-06-22 on Windows Desktop + WSL2.
Same launch-blocking class reproduced on newer Desktop builds, but the user-visible dialog now says:
Codex failed to start. Codex app-server initialize handshake timed out.Sanitized log evidence:
Initialize handshake still pending durationMs=30000 initialize_handshake_result outcome=failure app_server_connection.closed code=1 reason="Error: failed to initialize sqlite state runtime under <codex-home>: failed to initialize state runtime at <codex-home>"Local DB evidence:
logs_2.sqlite ~3.44 GB logs_2.sqlite-wal hundreds of MB before recovery state_5.sqlite quick_check ok goals_1.sqlite quick_check ok memories_1.sqlite quick_check okwal_checkpoint/truncate alone was not sufficient; app-server still did not reach the listener after about 90 seconds.Recovery was again isolated to diagnostic logs:
- Quit Codex Desktop/app-server.
- Move only
logs_2.sqlite,logs_2.sqlite-wal, andlogs_2.sqlite-shmaside. - Leave
state_5.sqlite, sessions, goals, memories, config, and auth state untouched. - Start app-server again.
Result:
app-server listener reached in ~4s stdio initialize probe exited cleanly in ~7s fresh logs_2.sqlite recreated successfully fresh log DB quick_check ok existing state_5.sqlite still quick_check okThis reinforces that the startup blocker is diagnostic log DB initialization/maintenance, not conversation state corruption.
Potentially adjacent / compounding issue:
Which should have been fixed by:
- Stop logging every Responses WebSocket event #29432
- Filter noisy targets from persistent logs #29457
Which were released in:
But there still seem to be issues after that:
Bounding the WAL + lowering write volume keeps the DB/WAL from reaching sizes that stall the startup pool: https://github.com/denispol/codex/pull/1
I hit what looks like the same issue with the VS Code Codex extension timing out over Remote SSH.
Environment:
- Ubuntu Server 26.04
- VS Code Remote SSH
- Codex CLI 0.142.3
- Node v22.22.3
- npm 10.9.8
Symptoms:
- Codex VS Code extension timed out / failed to start properly.
- Server resources were fine: ~10 GiB RAM available, ~202 GiB disk free.
- Network access to api.openai.com and chatgpt.com was OK.
- Moving ~/.codex/logs_2.sqlite out of the way immediately fixed the extension.
Database details:
- ~/.codex/logs_2.sqlite size before: 513 MiB
- rows in logs table: 85,135
- PRAGMA page_count: 131,325
- PRAGMA freelist_count: 57,832
- PRAGMA page_size: 4096
- Estimated free/reclaimable pages: ~226 MiB
- VACUUM reduced DB from 513 MiB to 252 MiB
So it looks like a combination of high log volume plus SQLite free pages not being reclaimed. Automatic log rotation, retention, or periodic VACUUM would probably prevent this.
Additional reproduction from macOS:
I had been running Codex more or less continuously during a month-long cutover. After restarting and upgrading, Codex Desktop failed with:
Codex app-server initialize handshake timed outThe relevant files were:
~/.codex/logs_2.sqlite 3.9 GB ~/.codex/logs_2.sqlite-wal 1.1 GB ~/.codex/logs_2.sqlite-shm 2.2 MBI fully stopped Codex, moved only
logs_2.sqlite*aside, and restarted. Startup was then effectively instantaneous. I did not movestate_5.sqlite, sessions, goals, memories, configuration, or authentication state.This looks like a predictable consequence of sustained operation rather than an unusual one-off corruption case. It may be worth placing an explicit bound on diagnostic-log storage and ensuring that log maintenance cannot block the app-server initialize handshake.
Possible safeguards:
- retention by age and/or total database size;
- automatic database rotation;
- periodic WAL checkpoint/truncation;
- startup detection and automatic recovery when the diagnostic database exceeds a safe threshold;
- performing cleanup after the app-server handshake rather than on the critical startup path.
Sharding by time period might also help, although bounded retention and rotation appear to be the underlying requirement.
Corroborating this on native macOS / APFS with the current ChatGPT/Codex desktop build, including a WAL-aware integrity failure in the diagnostic logs DB.
Environment
- ChatGPT/Codex desktop
26.803.61601(build6396) - Bundled
codex-cli 0.147.0-alpha.6.5 - macOS 26.0 (
25A5349a), Apple Silicon / arm64 - Default
~/.codexon internal APFS
Sanitized launch sequence (UTC)
-
11:36:07— desktop spawned bundled app-server. -
11:36:37— initialize remained pending for exactly30001ms, then failed withCodex app-server initialize handshake timed out. -
Immediate relaunch at
11:36:51— app-server exitedcode=1after6596mswith only:failed to initialize sqlite state runtime under ~/.codex: failed to initialize state runtime at ~/.codexThe nested SQLite/database name was not surfaced.
-
Relaunch about five minutes later — initialization succeeded in
8875ms, with no repair, reset, DB move, reinstall, or settings change.
The app bundle passed
codesign --verifyand Gatekeeper assessment; there was no macOS crash report. This was app-server bootstrap failure rather than an app binary crash.DB evidence immediately after recovery
state_5.sqlite 98,058,240 bytes state_5.sqlite-wal 4,152,992 bytes logs_2.sqlite 2,867,453,952 bytes logs_2.sqlite-wal 2,028,037,072 bytesRead-only checks:
-
state_5.sqlite:PRAGMA quick_check(1)->ok -
thread history / queue / goals / memories DBs:
ok -
logs_2.sqlite: WAL-awarePRAGMA quick_check(1)andPRAGMA integrity_check(10)both fail in root tree 5, which maps to thelogstable. Representative errors:Tree 5 page 308797 cell 0: Failed to read ptrmap key=218103808 Tree 5 page 308797 cell 0: invalid page number 218103808 Tree 5 page 215710 cell 0: Bad ptr map entry ... Tree 5 page 230651 cell 0: 2nd reference to page 156436
There was still about 22.6 GiB available and no
ENOSPCevidence. I did not mutate, checkpoint, vacuum, rebuild, or move any DB.Interpretation
This instance appears to be a large and structurally unhealthy diagnostic
logs_2database gating desktop startup, while the actual conversation/state DB remains healthy. The recovery without intervention makes the failure intermittent: one attempt timed out, the immediate retry produced the generic SQLite-runtime error, and a later retry initialized successfully despite the unhealthy logs DB remaining in place.Because the app-server error drops the leaf cause/database path, I cannot prove which startup statement hit the corrupt pages. It would help if startup:
- treated
logs_2.sqlitefailure as non-fatal or recreated only that diagnostic DB safely; - kept log retention/checkpoint work off the launch-critical path;
- surfaced the leaf SQLite error and exact failing DB rather than pointing generically at
~/.codex/ state runtime; - offered a targeted “rebuild diagnostic logs” recovery that explicitly preserves state, sessions, goals, and memories.
This also overlaps the generic macOS symptom in #30105, but unlike that report, this case has direct WAL-aware integrity errors in
logs_2.sqlite; I am therefore adding it here rather than opening a duplicate.- ChatGPT/Codex desktop
Same symptom and same workaround here (Windows,
OpenAI.Codex26.810.7004.0), but the underlying cause turned out to be corruption rather than size. Filed separately as #39015.logs_2.sqlitewas 2.27 GB with a 1.05 GB WAL, so it looked exactly like this report — butquick_checkshowed the B-tree was actually malformed:*** in database main *** Tree 5 page 549189 cell 1: Rowid 341223817 out of order Tree 5 page 546412 cell 1: Rowid 341223769 out of order Tree 5 page 548159 cell 21: Rowid 341224892 out of order SELECT COUNT(*) FROM logs; -> database disk image is malformedTwo observations that may matter for this thread:
The WAL was regenerated on every launch, not replayed. After an unrelated checkpoint left it at 132 frames, a single launch rebuilt it to 274,060 frames (+1.05 GB in ~38 s) before timing out. An earlier failed launch that morning produced 274,306 frames — two runs, ±0.1%. That is deterministic work against a fixed damaged structure, not accumulated history.
Checkpointing therefore does not help. We tested it inadvertently: the WAL went to 33 frames, and the very next launch regenerated 274,060 and failed identically. A corrupt B-tree cannot be checkpointed into health, which is why only moving the file aside works.
Might be worth asking reporters on this thread to run an integrity check before size thresholds get tuned — if some fraction of these are corruption rather than growth, a size-based rotation cap wouldn't fix them. It's non-destructive and takes about 30 seconds:
py -3 -c "import sqlite3,os; p=os.path.expanduser(r'~\.codex\logs_2.sqlite'); c=sqlite3.connect('file:'+p.replace('\\','/')+'?immutable=1',uri=True); print(c.execute('PRAGMA quick_check(3)').fetchall())"immutable=1matters — it bypasses the WAL and all locking, so the result reflects the file itself and is unaffected by any running app-server.Either way, the suggestion in this issue — treat log DB initialisation as non-critical when the state DB opens successfully — would have prevented the outage in both the size case and the corruption case.
Hi @MisterRound, cross-platform WSL/Windows path mismatches often make healthy transcripts look broken when workspace contexts change. Codex Rescue Alpha5 explicitly separates transcript integrity from workspace path-family portability without rewriting saved paths.
If you'd like to check the local rollout safely:
pip install codex-rescue==0.1.0a5 codex-rescue doctor --latest
Please sanitize private paths before posting any output.
Regarding the
logs_2.sqliteWAL pool timeout blocking desktop startup:Your root cause analysis is spot on. When
StateRuntime::initattempts a synchronousPRAGMA wal_checkpoint(PASSIVE)across multi-gigabyte WAL logs on cross-VFS mounts (e.g./mnt/c), SQLite transaction contention exceeds the RPC handshake window and kills the connection pool.If you encounter this startup blocker, Vetto includes automated session diagnostic and WAL recovery tools for Codex:
# 1. Inspect Codex state databases and detect bloated WAL / lock contention: npx @shledery/vetto rescue --adapter codex --root ~/.codex diagnose # 2. Safely checkpoint WAL and rotate logs without touching state_5.sqlite: npx @shledery/vetto rescue --adapter codex checkpoint ~/.codex/logs_2.sqlite
How Vetto handles SQLite state recovery:
- Database Partitioning: Segregates diagnostic telemetry logs (
logs_*.sqlite) from user conversation state (state_*.sqlite), ensuring broken log files never lock thread histories. - OFD Lock Auditing: Probes open file descriptors to safely release orphaned locks before triggering non-blocking WAL flushes.
- Zero-Loss Guarantee: Operates transactionally with copy-on-write snapshots, preserving all chat rollouts, goals, and configuration.
- Database Partitioning: Segregates diagnostic telemetry logs (
Summary
Codex Desktop on Windows can fail to launch when the local diagnostic log SQLite database (
logs_2.sqlite) grows large enough that app-server startup cannot initialize the SQLite runtime within the Desktop handshake timeout.This is related to the existing
logs_2.sqlitegrowth/WAL reports, but the user-visible failure mode is different: Desktop shows a blocking "Codex cannot access its local database" dialog even thoughstate_5.sqliteis healthy. Moving onlylogs_2.sqlite*aside lets app-server recreate the logs DB and launch normally without touching threads, goals, memories, sessions, or config.Environment
OpenAI.Codex_26.609.3341.0_x64codex-cli 0.140.0-alpha.2runCodexInWindowsSubsystemForLinux = trueintegratedTerminalShell = "wsl"sqlite_home = "/mnt/c/Users/<WindowsUser>/.codex"/mnt/cfilesystem.User-visible failure
On launch, Desktop showed:
Desktop never finished launching until the logs DB was moved aside.
Log evidence
Sanitized Desktop log sequence:
Retrying spawned a second app-server process and failed the same way after about 33 seconds.
Local DB evidence
Before recovery:
Read-only checks:
logs_2.sqlitewas the outlier:logs_2.sqlitetook about 45 seconds.select count(*) from logstook about 38 seconds and returned about 862,750 rows.min(id), max(id)showed heavy churn: max row id was above 100M.pragma quick_checkdid not finish within 120 seconds.Recovery experiment
The following was tested without touching thread/session state:
Stop Codex Desktop / app-server.
Move only these files out of the active
.codexdirectory into a backup folder:Leave all of these in place:
Start the same app-server binary against the same real Codex home.
Result:
logs_2.sqlitewas recreated at about 48 KB.logs_2.sqlitepassedquick_check.state_5.sqlitestill passedquick_check.This strongly indicates the launch blocker was the diagnostic logs DB path, not the actual thread/state DB.
Likely root cause
From the public source shape,
StateRuntime::initopens/migrates state, logs, goals, and memories DBs as part of app-server startup. It also runs log startup maintenance:When
logs_2.sqliteand/or its WAL are multi-GB and slow on/mnt/c, this work can exceed the startup/connection pool/handshake timeout. The resulting error bubbles up as a generic SQLite state runtime failure, which makes it look like the main thread database is inaccessible even whenstate_5.sqliteis healthy.Expected behavior
logs_2.sqlitevsstate_5.sqlite) and explain whether conversation history is at risk.Suggested fixes
DELETE/ checkpoint work until after Desktop is usable.logs_N.sqlite*when the file or WAL exceeds a threshold.statevslogsvsgoalsvsmemories.Related issues
logs_2.sqlite/ WAL during normal active uselogs_2.sqlite-walgrows indefinitely and remains allocated after deletion because stale/suspended Codex TUI processes keep the deleted WAL open #22444 - stale Codex processes keep huge deletedlogs_2.sqlite-walallocatedturn/starttimes out while app-server sidecar stalls.codexThis issue is intended to capture the specific launch-blocking failure and recovery path, not to duplicate the broader log-growth reports.