Skip to content

Desktop launch can fail when logs_2.sqlite grows large: app-server SQLite pool times out during startup #27741

Description

@MisterRound

Summary

Codex Desktop on Windows can fail to launch when the local diagnostic log SQLite database (logs_2.sqlite) grows large enough that app-server startup cannot initialize the SQLite runtime within the Desktop handshake timeout.

This is related to the existing logs_2.sqlite growth/WAL reports, but the user-visible failure mode is different: Desktop shows a blocking "Codex cannot access its local database" dialog even though state_5.sqlite is healthy. Moving only logs_2.sqlite* aside lets app-server recreate the logs DB and launch normally without touching threads, goals, memories, sessions, or config.

Environment

  • Product: Codex Desktop for Windows
  • Windows Store package observed in logs: OpenAI.Codex_26.609.3341.0_x64
  • Bundled WSL CLI observed: codex-cli 0.140.0-alpha.2
  • Mode: Desktop app-server running in WSL
  • Config shape:
    • runCodexInWindowsSubsystemForLinux = true
    • integratedTerminalShell = "wsl"
    • sqlite_home = "/mnt/c/Users/<WindowsUser>/.codex"
  • SQLite home is on the Windows-backed /mnt/c filesystem.

User-visible failure

On launch, Desktop showed:

Codex cannot access its local database.

Database path: Path unavailable in app-server startup error

Error: (code=1, signal=null).
Most recent error: Error: failed to initialize sqlite state runtime under
/mnt/c/Users/<WindowsUser>/.codex: failed to initialize state runtime at
/mnt/c/Users/<WindowsUser>/.codex: pool timed out while waiting for an open connection

Desktop never finished launching until the logs DB was moved aside.

Log evidence

Sanitized Desktop log sequence:

Launching app ... package OpenAI.Codex_26.609.3341.0_x64
[wsl] eligible distro list ... Ubuntu Running 2
stdio_transport_spawned executablePath=C:\Users\<WindowsUser>\.codex\bin\wsl\<hash>\codex spawnCommand=wsl.exe
Initialize handshake still pending durationMs=10020
app_server_connection.closed code=1 reason="Error: failed to initialize sqlite state runtime under /mnt/c/Users/<WindowsUser>/.codex: failed to initialize state runtime at /mnt/c/Users/<WindowsUser>/.codex: pool timed out while waiting for an open connection"
initialize_handshake_result durationMs=33522 outcome=failure

Retrying spawned a second app-server process and failed the same way after about 33 seconds.

Local DB evidence

Before recovery:

state_5.sqlite          ~35.6 MB
logs_2.sqlite           ~4.5 GB
logs_2.sqlite-wal       ~1.08 GB
logs_2.sqlite-shm       ~1.47 MB
goals_1.sqlite          small
memories_1.sqlite       small

Read-only checks:

state_5.sqlite quick_check: ok
goals_1.sqlite quick_check: ok
memories_1.sqlite quick_check: ok

logs_2.sqlite was the outlier:

  • A simple metadata query against logs_2.sqlite took about 45 seconds.
  • select count(*) from logs took about 38 seconds and returned about 862,750 rows.
  • min(id), max(id) showed heavy churn: max row id was above 100M.
  • WAL-aware pragma quick_check did not finish within 120 seconds.
  • An immutable read-only check reported freelist / pointer-map errors, but because immutable mode ignores WAL frames this should be treated as corruption-suspicious rather than definitive corruption.

Recovery experiment

The following was tested without touching thread/session state:

  1. Stop Codex Desktop / app-server.

  2. Move only these files out of the active .codex directory into a backup folder:

    logs_2.sqlite
    logs_2.sqlite-wal
    logs_2.sqlite-shm
    
  3. Leave all of these in place:

    state_5.sqlite
    goals_1.sqlite*
    memories_1.sqlite*
    sessions/**
    archived_sessions/**
    session_index.jsonl
    config.toml
    auth files
    plugin/cache/config files
    
  4. Start the same app-server binary against the same real Codex home.

Result:

  • app-server startup succeeded.
  • A fresh logs_2.sqlite was recreated at about 48 KB.
  • Fresh logs_2.sqlite passed quick_check.
  • Existing state_5.sqlite still passed quick_check.
  • Thread/session files were not moved or edited.

This strongly indicates the launch blocker was the diagnostic logs DB path, not the actual thread/state DB.

Likely root cause

From the public source shape, StateRuntime::init opens/migrates state, logs, goals, and memories DBs as part of app-server startup. It also runs log startup maintenance:

delete logs older than retention window
PRAGMA wal_checkpoint(PASSIVE)

When logs_2.sqlite and/or its WAL are multi-GB and slow on /mnt/c, this work can exceed the startup/connection pool/handshake timeout. The resulting error bubbles up as a generic SQLite state runtime failure, which makes it look like the main thread database is inaccessible even when state_5.sqlite is healthy.

Expected behavior

  • A large, slow, or damaged diagnostic log DB should not prevent Desktop from launching.
  • Startup should not synchronously block on log retention/checkpoint work that can take tens of seconds or minutes.
  • If the logs DB is unhealthy or too slow, Codex should rotate/regenerate it automatically or offer a safe "clear diagnostic logs" recovery path.
  • The error dialog should identify the specific DB that failed (logs_2.sqlite vs state_5.sqlite) and explain whether conversation history is at risk.
  • Log maintenance should be bounded, backgrounded, and non-fatal where possible.

Suggested fixes

  • Treat log DB initialization/maintenance as non-critical for app launch when state DB opens successfully.
  • Add a bounded timeout around log startup maintenance, and defer heavy DELETE / checkpoint work until after Desktop is usable.
  • Add automatic rotation of logs_N.sqlite* when the file or WAL exceeds a threshold.
  • Configure WAL limits/checkpointing/retention so the diagnostic DB cannot grow to multi-GB under normal Desktop/goal/app-server usage.
  • Surface a safe recovery button/action: "Move diagnostic logs aside and relaunch".
  • Improve diagnostics so the dialog reports the actual failing DB path and classifies state vs logs vs goals vs memories.

Related issues

This issue is intended to capture the specific launch-blocking failure and recovery path, not to duplicate the broader log-growth reports.

Activity

  1. added
    bugSomething isn't working
    appIssues related to the Codex desktop app
    windows-osIssues related to Codex on Windows systems
    app-serverIssues involving app server protocol or interfaces
    on Jun 12, 2026
  2. MisterRound commented on Jun 22, 2026

    @MisterRound
    Author

    Update from 2026-06-22 on Windows Desktop + WSL2.

    Same launch-blocking class reproduced on newer Desktop builds, but the user-visible dialog now says:

    Codex failed to start.
    Codex app-server initialize handshake timed out.
    

    Sanitized log evidence:

    Initialize handshake still pending durationMs=30000
    initialize_handshake_result outcome=failure
    app_server_connection.closed code=1 reason="Error: failed to initialize sqlite state runtime under <codex-home>: failed to initialize state runtime at <codex-home>"
    

    Local DB evidence:

    logs_2.sqlite      ~3.44 GB
    logs_2.sqlite-wal  hundreds of MB before recovery
    state_5.sqlite     quick_check ok
    goals_1.sqlite     quick_check ok
    memories_1.sqlite  quick_check ok
    

    wal_checkpoint/truncate alone was not sufficient; app-server still did not reach the listener after about 90 seconds.

    Recovery was again isolated to diagnostic logs:

    1. Quit Codex Desktop/app-server.
    2. Move only logs_2.sqlite, logs_2.sqlite-wal, and logs_2.sqlite-shm aside.
    3. Leave state_5.sqlite, sessions, goals, memories, config, and auth state untouched.
    4. Start app-server again.

    Result:

    app-server listener reached in ~4s
    stdio initialize probe exited cleanly in ~7s
    fresh logs_2.sqlite recreated successfully
    fresh log DB quick_check ok
    existing state_5.sqlite still quick_check ok
    

    This reinforces that the startup blocker is diagnostic log DB initialization/maintenance, not conversation state corruption.

  3. denispol commented on Jun 25, 2026

    @denispol

    Bounding the WAL + lowering write volume keeps the DB/WAL from reaching sizes that stall the startup pool: https://github.com/denispol/codex/pull/1

  4. hareidx commented on Jun 28, 2026

    @hareidx

    I hit what looks like the same issue with the VS Code Codex extension timing out over Remote SSH.

    Environment:

    • Ubuntu Server 26.04
    • VS Code Remote SSH
    • Codex CLI 0.142.3
    • Node v22.22.3
    • npm 10.9.8

    Symptoms:

    • Codex VS Code extension timed out / failed to start properly.
    • Server resources were fine: ~10 GiB RAM available, ~202 GiB disk free.
    • Network access to api.openai.com and chatgpt.com was OK.
    • Moving ~/.codex/logs_2.sqlite out of the way immediately fixed the extension.

    Database details:

    • ~/.codex/logs_2.sqlite size before: 513 MiB
    • rows in logs table: 85,135
    • PRAGMA page_count: 131,325
    • PRAGMA freelist_count: 57,832
    • PRAGMA page_size: 4096
    • Estimated free/reclaimable pages: ~226 MiB
    • VACUUM reduced DB from 513 MiB to 252 MiB

    So it looks like a combination of high log volume plus SQLite free pages not being reclaimed. Automatic log rotation, retention, or periodic VACUUM would probably prevent this.

  5. bozo32 commented on Jul 23, 2026

    @bozo32

    Additional reproduction from macOS:

    I had been running Codex more or less continuously during a month-long cutover. After restarting and upgrading, Codex Desktop failed with:

    Codex app-server initialize handshake timed out

    The relevant files were:

    ~/.codex/logs_2.sqlite       3.9 GB
    ~/.codex/logs_2.sqlite-wal   1.1 GB
    ~/.codex/logs_2.sqlite-shm   2.2 MB
    

    I fully stopped Codex, moved only logs_2.sqlite* aside, and restarted. Startup was then effectively instantaneous. I did not move state_5.sqlite, sessions, goals, memories, configuration, or authentication state.

    This looks like a predictable consequence of sustained operation rather than an unusual one-off corruption case. It may be worth placing an explicit bound on diagnostic-log storage and ensuring that log maintenance cannot block the app-server initialize handshake.

    Possible safeguards:

    • retention by age and/or total database size;
    • automatic database rotation;
    • periodic WAL checkpoint/truncation;
    • startup detection and automatic recovery when the diagnostic database exceeds a safe threshold;
    • performing cleanup after the app-server handshake rather than on the critical startup path.

    Sharding by time period might also help, although bounded retention and rotation appear to be the underlying requirement.

  6. paulostergaard commented on Aug 12, 2026

    @paulostergaard

    Corroborating this on native macOS / APFS with the current ChatGPT/Codex desktop build, including a WAL-aware integrity failure in the diagnostic logs DB.

    Environment

    • ChatGPT/Codex desktop 26.803.61601 (build 6396)
    • Bundled codex-cli 0.147.0-alpha.6.5
    • macOS 26.0 (25A5349a), Apple Silicon / arm64
    • Default ~/.codex on internal APFS

    Sanitized launch sequence (UTC)

    1. 11:36:07 — desktop spawned bundled app-server.

    2. 11:36:37 — initialize remained pending for exactly 30001ms, then failed with Codex app-server initialize handshake timed out.

    3. Immediate relaunch at 11:36:51 — app-server exited code=1 after 6596ms with only:

      failed to initialize sqlite state runtime under ~/.codex:
      failed to initialize state runtime at ~/.codex
      

      The nested SQLite/database name was not surfaced.

    4. Relaunch about five minutes later — initialization succeeded in 8875ms, with no repair, reset, DB move, reinstall, or settings change.

    The app bundle passed codesign --verify and Gatekeeper assessment; there was no macOS crash report. This was app-server bootstrap failure rather than an app binary crash.

    DB evidence immediately after recovery

    state_5.sqlite          98,058,240 bytes
    state_5.sqlite-wal       4,152,992 bytes
    logs_2.sqlite        2,867,453,952 bytes
    logs_2.sqlite-wal    2,028,037,072 bytes
    

    Read-only checks:

    • state_5.sqlite: PRAGMA quick_check(1) -> ok

    • thread history / queue / goals / memories DBs: ok

    • logs_2.sqlite: WAL-aware PRAGMA quick_check(1) and PRAGMA integrity_check(10) both fail in root tree 5, which maps to the logs table. Representative errors:

      Tree 5 page 308797 cell 0: Failed to read ptrmap key=218103808
      Tree 5 page 308797 cell 0: invalid page number 218103808
      Tree 5 page 215710 cell 0: Bad ptr map entry ...
      Tree 5 page 230651 cell 0: 2nd reference to page 156436
      

    There was still about 22.6 GiB available and no ENOSPC evidence. I did not mutate, checkpoint, vacuum, rebuild, or move any DB.

    Interpretation

    This instance appears to be a large and structurally unhealthy diagnostic logs_2 database gating desktop startup, while the actual conversation/state DB remains healthy. The recovery without intervention makes the failure intermittent: one attempt timed out, the immediate retry produced the generic SQLite-runtime error, and a later retry initialized successfully despite the unhealthy logs DB remaining in place.

    Because the app-server error drops the leaf cause/database path, I cannot prove which startup statement hit the corrupt pages. It would help if startup:

    1. treated logs_2.sqlite failure as non-fatal or recreated only that diagnostic DB safely;
    2. kept log retention/checkpoint work off the launch-critical path;
    3. surfaced the leaf SQLite error and exact failing DB rather than pointing generically at ~/.codex / state runtime;
    4. offered a targeted “rebuild diagnostic logs” recovery that explicitly preserves state, sessions, goals, and memories.

    This also overlaps the generic macOS symptom in #30105, but unlike that report, this case has direct WAL-aware integrity errors in logs_2.sqlite; I am therefore adding it here rather than opening a duplicate.

  7. AS-1-Dev commented on Aug 17, 2026

    @AS-1-Dev

    Same symptom and same workaround here (Windows, OpenAI.Codex 26.810.7004.0), but the underlying cause turned out to be corruption rather than size. Filed separately as #39015.

    logs_2.sqlite was 2.27 GB with a 1.05 GB WAL, so it looked exactly like this report — but quick_check showed the B-tree was actually malformed:

    *** in database main ***
    Tree 5 page 549189 cell 1:  Rowid 341223817 out of order
    Tree 5 page 546412 cell 1:  Rowid 341223769 out of order
    Tree 5 page 548159 cell 21: Rowid 341224892 out of order
    
    SELECT COUNT(*) FROM logs;  ->  database disk image is malformed
    

    Two observations that may matter for this thread:

    The WAL was regenerated on every launch, not replayed. After an unrelated checkpoint left it at 132 frames, a single launch rebuilt it to 274,060 frames (+1.05 GB in ~38 s) before timing out. An earlier failed launch that morning produced 274,306 frames — two runs, ±0.1%. That is deterministic work against a fixed damaged structure, not accumulated history.

    Checkpointing therefore does not help. We tested it inadvertently: the WAL went to 33 frames, and the very next launch regenerated 274,060 and failed identically. A corrupt B-tree cannot be checkpointed into health, which is why only moving the file aside works.

    Might be worth asking reporters on this thread to run an integrity check before size thresholds get tuned — if some fraction of these are corruption rather than growth, a size-based rotation cap wouldn't fix them. It's non-destructive and takes about 30 seconds:

    py -3 -c "import sqlite3,os; p=os.path.expanduser(r'~\.codex\logs_2.sqlite'); c=sqlite3.connect('file:'+p.replace('\\','/')+'?immutable=1',uri=True); print(c.execute('PRAGMA quick_check(3)').fetchall())"
    

    immutable=1 matters — it bypasses the WAL and all locking, so the result reflects the file itself and is unaffected by any running app-server.

    Either way, the suggestion in this issue — treat log DB initialisation as non-critical when the state DB opens successfully — would have prevented the outage in both the size case and the corruption case.

  8. boombx403-byte commented on Aug 18, 2026

    @boombx403-byte

    Hi @MisterRound, cross-platform WSL/Windows path mismatches often make healthy transcripts look broken when workspace contexts change. Codex Rescue Alpha5 explicitly separates transcript integrity from workspace path-family portability without rewriting saved paths.

    If you'd like to check the local rollout safely:

    pip install codex-rescue==0.1.0a5
    codex-rescue doctor --latest

    Please sanitize private paths before posting any output.

  9. shleder commented on Aug 28, 2026

    @shleder

    Regarding the logs_2.sqlite WAL pool timeout blocking desktop startup:

    Your root cause analysis is spot on. When StateRuntime::init attempts a synchronous PRAGMA wal_checkpoint(PASSIVE) across multi-gigabyte WAL logs on cross-VFS mounts (e.g. /mnt/c), SQLite transaction contention exceeds the RPC handshake window and kills the connection pool.

    If you encounter this startup blocker, Vetto includes automated session diagnostic and WAL recovery tools for Codex:

    # 1. Inspect Codex state databases and detect bloated WAL / lock contention:
    npx @shledery/vetto rescue --adapter codex --root ~/.codex diagnose
    
    # 2. Safely checkpoint WAL and rotate logs without touching state_5.sqlite:
    npx @shledery/vetto rescue --adapter codex checkpoint ~/.codex/logs_2.sqlite

    How Vetto handles SQLite state recovery:

    • Database Partitioning: Segregates diagnostic telemetry logs (logs_*.sqlite) from user conversation state (state_*.sqlite), ensuring broken log files never lock thread histories.
    • OFD Lock Auditing: Probes open file descriptors to safely release orphaned locks before triggering non-blocking WAL flushes.
    • Zero-Loss Guarantee: Operates transactionally with copy-on-write snapshots, preserving all chat rollouts, goals, and configuration.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    appIssues related to the Codex desktop appapp-serverIssues involving app server protocol or interfacesbugSomething isn't workingperformancewindows-osIssues related to Codex on Windows systems

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions