Skip to content

Codex SQLite feedback logs can write ~640 TB/year and rapidly consume SSD endurance #28224

Description

@1996fanrui

Update at Jun 23, 2026: the following 3 PRs are merged, it could avoid 85% logs(feedback from my codex), so let me close this issue.

Thanks @jif-oai for the fix.

Simple Workaround from @beskay #28224 (comment)

  1. Quit codex
  2. Run this command:
sqlite3 ~/.codex/logs_2.sqlite "CREATE TRIGGER IF NOT EXISTS
  block_log_inserts BEFORE INSERT ON logs BEGIN SELECT RAISE(IGNORE);
  END;"

Following is the original issue

Issue

Codex is continuously writing a large amount of data to the local SQLite feedback log database:

  • ~/.codex/logs_2.sqlite
  • ~/.codex/logs_2.sqlite-wal
  • ~/.codex/logs_2.sqlite-shm

On my machine, after about 21 days of uptime, the main SSD has written about 37 TB. Process/file-level checks show Codex SQLite logs are the main continuous writer.

That extrapolates to roughly 640 TB/year. On a 1 TB SSD, that is about 640 full-drive writes per year. Some consumer SSDs are rated around 600 TBW, so this could consume roughly a full drive's warranted write endurance in less than a year.

Evidence1

A later snapshot makes the write amplification easier to see:

metric value
current logs_2.sqlite file size 1.2 GiB
current retained rows 506,149
total allocated row ids 5,543,677,486

So the database currently retains only ~0.5M rows, while the SQLite AUTOINCREMENT counter has already advanced past 5.5B ids.

That is roughly a 10,000x gap between retained rows and historical inserted row ids. Even using the current ~1.2 GiB database size as a rough baseline, this points to 10TB+ scale historical log churn, before accounting for WAL, indexes, pruning, checkpoints, page rewrites, and filesystem/device-level write amplification.

Evidence2

Current retained rows in logs_2.sqlite:

metric value
retained rows 681,774
estimated retained log content 1,035.6 MiB

Level distribution:

level estimated MiB byte %
TRACE 732.5 70.7%
INFO 266.5 25.7%
DEBUG 30.6 3.0%
WARN 5.9 0.6%

Largest target+level pairs:

target level estimated MiB
codex_api::endpoint::responses_websocket TRACE 527.4
codex_otel.log_only INFO 141.2
codex_otel.trace_safe INFO 121.2
log TRACE 97.4
codex_client::transport TRACE 60.1
codex_core::stream_events_utils DEBUG 27.5
codex_api::sse::responses TRACE 19.1

The top sources are mostly global TRACE logs, mirrored telemetry logs, and raw websocket/SSE payload logging. TRACE alone is about 70.7% of retained bytes. codex_otel.log_only + codex_otel.trace_safe add another 25.3%. Filtering these categories should remove roughly 96% of retained log bytes in this sample without fully disabling feedback logs.

Sanitized examples from the most frequent TRACE source: target=log

These are high-frequency retained samples. Raw websocket/SSE payload bodies are intentionally not included because they may contain private conversation content.

128,764x TRACE log: inotify event: ... mask: OPEN, name: Some("ld.so.cache")
 37,982x TRACE log: inotify event: ... mask: OPEN, name: Some("locale.alias")
 23,843x TRACE log: inotify event: ... mask: OPEN, name: Some("passwd")
  3,639x TRACE log: <tokio-tungstenite checkout>/src/compat.rs:131 AllowStd.with_context
  3,505x TRACE log: <tokio-tungstenite checkout>/src/lib.rs:245 WebSocketStream.with_context
  3,362x TRACE log: <tokio-tungstenite checkout>/src/compat.rs:154 Read.read
  3,356x TRACE log: <tokio-tungstenite checkout>/src/compat.rs:157 Read.with_context read -> poll_read
  3,230x TRACE log: <tokio-tungstenite checkout>/src/lib.rs:294 Stream.poll_next
  3,227x TRACE log: <tokio-tungstenite checkout>/src/lib.rs:304 Stream.with_context poll_next -> read()
  3,213x TRACE log: inotify event: ... mask: OPEN, name: Some("nsswitch.conf")
  2,001x TRACE log: WouldBlock
  1,217x TRACE log: Masked: false
  1,169x TRACE log: Opcode: Data(Text)
  1,169x TRACE log: First: 11000001
Sanitized examples from frequent INFO sources

The dominant INFO sources are mostly repeated OpenTelemetry mirror events. IDs are redacted.

843x INFO codex_client::custom_ca:
  using system root certificates because no CA override environment variable was selected ...

334x INFO codex_otel.trace_safe:
  session_loop{thread_id=<redacted>}:submission_dispatch{otel.name="op.dispatch.user_input" submission.id=<redacted> codex.op="user_input"}:turn{otel.name="session_task.turn" thread.id=<redacted> ...}

333x INFO codex_otel.log_only:
  session_loop{thread_id=<redacted>}:submission_dispatch{otel.name="op.dispatch.user_input" submission.id=<redacted> codex.op="user_input"}:turn{otel.name="session_task.turn" thread.id=<redacted> ...}

332x INFO codex_otel.log_only:
  session_loop{thread_id=<redacted>}:submission_dispatch{otel.name="op.dispatch.user_input_with_turn_context" submission.id=<redacted> codex.op="user_input_with_turn_context"}:turn{otel.name="session_task.turn" thread.id=<redacted> ...}

332x INFO codex_otel.trace_safe:
  session_loop{thread_id=<redacted>}:submission_dispatch{otel.name="op.dispatch.user_input_with_turn_context" submission.id=<redacted> codex.op="user_input_with_turn_context"}:turn{otel.name="session_task.turn" thread.id=<redacted> ...}

Write amplification

The retained DB size hides the real write volume. In a 15-second sample:

metric before after
retained rows 681,774 681,774
max row id 5,003,347,015 5,003,383,226

About 36,211 rows were inserted in 15 seconds, while retained row count stayed flat. This suggests continuous insert-and-prune write amplification: rows are inserted, indexed, written to WAL, then pruned.

Likely cause

The SQLite feedback log sink is installed with a global TRACE default:

Targets::new().with_default(Level::TRACE)

This persists all targets at TRACE level by default, including dependency/internal logs and large raw protocol payloads.

Proposed fix

Keep feedback logs enabled, but narrow what is persisted by default:

  1. Do not use global TRACE for the SQLite feedback log sink.
  2. Drop or raise thresholds for low-value dependency noise, especially target=log, hyper_util, tokio-tungstenite internals, inotify spam, and low-level OpenTelemetry SDK logs.
  3. Avoid persisting full raw websocket/SSE payloads by default. Store summaries instead: event kind, duration, success/error, token usage, and payload byte length.
  4. Avoid persisting mirrored codex_otel.log_only / codex_otel.trace_safe events unless they are explicitly useful for feedback debugging.
  5. Add a global logs DB size/write cap. Per-thread caps are not enough when many threads/processes exist.

An optional escape hatch such as sqlite_logs_enabled = false would still be useful, but the main fix should be better default filtering.

Related issues and discussions

Activity

  1. 1996fanrui commented on Jun 14, 2026

    @1996fanrui
    Author

    cc @charley-oai @jif-oai since #12969 introduced the SQLite feedback log sink at TRACE level.

    I pushed a minimal branch here:
    https://github.com/1996fanrui/codex/tree/codex/reduce-sqlite-feedback-log-writes

    The proposed patch does not disable feedback logs. It only changes the SQLite feedback log sink from global TRACE to this default filter:

    warn,codex_api=info,codex_app_server=info,codex_core=info,codex_mcp=info,codex_state=info,codex_tui=info
    

    In my local sample, the retained SQLite log content was mostly TRACE plus codex_otel INFO mirror rows, so this should drop the high-volume low-signal writes while keeping INFO+ logs for core Codex targets.

    I tried to open an upstream PR, but GitHub returned:

    GraphQL: 1996fanrui does not have the correct permissions to execute CreatePullRequest
    

    Other possible fixes if the default-filter approach is not the preferred direction:

    • add a config escape hatch, for example sqlite_logs_enabled = false or a feedback-log-specific filter setting;
    • keep SQLite feedback logs opt-in or sampled instead of always-on at TRACE;
    • truncate or summarize large websocket/SSE/raw event payloads before storing them;
    • enforce a global write/size budget in addition to per-thread retention.
  2. ZenulAbidin commented on Jun 17, 2026

    @ZenulAbidin

    This is a very serious bug, as it can not only cause login failures on Linux boxes if the system is rebooted while the disk is full, but Codex in /goal mode will actively delete files and folders on your disk in a vain attempt to gain disk space. So the severity of this issue needs to be increased as data loss is possible when using Codex in this way.

  3. jordanade commented on Jun 17, 2026

    @jordanade

    WTF this is ridiculous. PLEASE have more respect for your users.

  4. Nicolas0315 commented on Jun 17, 2026

    @Nicolas0315

    I prepared a PR-ready branch against current main, but upstream PR creation is still blocked for my account with:

    GraphQL: Nicolas0315 does not have the correct permissions to execute `CreatePullRequest`
    

    Branch:
    https://github.com/Nicolas0315/codex/tree/codex/log-db-default-filter

    Compare:
    https://github.com/openai/codex/compare/main...Nicolas0315:codex/log-db-default-filter?expand=1

    Proposed PR title:

    [codex] Limit persisted SQLite log verbosity
    

    Proposed PR body, following the recent Codex PR format:

    Why

    logs_2.sqlite is currently attached with a global TRACE filter in both the app server and TUI. On my Windows Codex app install, that persisted high-volume websocket internals and OTel mirror records into the diagnostic SQLite log DB:

    • .codex/logs_2.sqlite: 729,575,424 bytes
    • .codex/logs_2.sqlite-wal: 39,255,392 bytes
    • retained rows: 269,141
    • TRACE: 123,026 rows / 269,065,577 estimated bytes
    • top persisted targets included codex_api::endpoint::responses_websocket, codex_otel.log_only, codex_otel.trace_safe, and log

    This matches the write-amplification report here and the Desktop launch failure mode in #27741. The change keeps diagnostic logs enabled, but makes the persisted SQLite sink less noisy by default.

    What changed

    • Add a shared log_db::default_filter() for the persistent SQLite log sink.
    • Persist INFO+ by default instead of TRACE.
    • Raise known high-volume mirror/dependency targets to WARN+:
      • codex_otel.log_only
      • codex_otel.trace_safe
      • hyper_util::client::legacy::*
      • log
      • opentelemetry_sdk
    • Use the shared filter from both app-server and TUI log DB setup.
    • Update the SQLite sink filter test to cover retained Codex INFO logs and dropped low-value OTel/mirror logs.

    Validation

    • git diff --check
    • just test -p codex-state sqlite_sink_default_filter_drops_low_value_logs
    • just test -p codex-tui sqlite_sink_default_filter_drops_low_value_logs compiled the package, then exited with no matching tests.
    • just fmt completed Rust/Python formatting, then failed on Windows in the Bazel buildifier step with [WinError 2].
    • A broader just test -p codex-app-server run compiled and executed tests but timed out after reporting 831 passed, 4 failed, 6 skipped; the visible failures appeared unrelated to this log filtering change, including a skills fixture warning count mismatch.
  5. ZenulAbidin commented on Jun 18, 2026

    @ZenulAbidin

    Here I have made two scripts which can serve as temporary mitigations for the WAL issue until OpenAI fixes it.

    The first script trims the WAL while Codex processes are running:

    #!/usr/bin/env bash
    # trim-codex-wal.sh
    # Trims the logs_2.sqlite and WAL files by truncating the checkpoint database
    set -u
    
    CODEX_DIR="${CODEX_HOME:-$HOME/.codex}"
    DB="$CODEX_DIR/logs_2.sqlite"
    WAL="$CODEX_DIR/logs_2.sqlite-wal"
    SHM="$CODEX_DIR/logs_2.sqlite-shm"
    
    bytes() {
      if [ -e "$1" ]; then
        stat -c '%s' "$1" 2>/dev/null || stat -f '%z' "$1" 2>/dev/null || echo 0
      else
        echo 0
      fi
    }
    
    human() {
      if command -v numfmt >/dev/null 2>&1; then
        numfmt --to=iec --suffix=B "$1" 2>/dev/null || echo "$1 bytes"
      else
        echo "$1 bytes"
      fi
    }
    
    show_sizes() {
      echo
      echo "Codex SQLite files:"
      for f in "$DB" "$WAL" "$SHM"; do
        printf "  %-45s %12s\n" "$f" "$(human "$(bytes "$f")")"
      done
    }
    
    echo "[LIVE RUN] Codex WAL truncate only"
    echo "Codex dir: $CODEX_DIR"
    echo "DB:        $DB"
    echo "WAL:       $WAL"
    echo "SHM:       $SHM"
    
    if [ ! -f "$DB" ]; then
      echo
      echo "ERROR: Codex DB not found:"
      echo "  $DB"
      exit 1
    fi
    
    if ! command -v sqlite3 >/dev/null 2>&1; then
      echo
      echo "ERROR: sqlite3 is not installed or not in PATH."
      exit 1
    fi
    
    show_sizes
    
    BEFORE_WAL="$(bytes "$WAL")"
    
    echo
    echo "Running SQLite WAL checkpoint/truncate..."
    echo "Command:"
    echo "  PRAGMA wal_checkpoint(TRUNCATE);"
    
    RESULT="$(
      sqlite3 "$DB" <<'SQL' 2>&1
    PRAGMA busy_timeout = 5000;
    PRAGMA wal_checkpoint(TRUNCATE);
    SQL
    )"
    RC=$?
    
    AFTER_WAL="$(bytes "$WAL")"
    
    echo
    echo "SQLite output:"
    if [ -n "$RESULT" ]; then
      printf '%s\n' "$RESULT"
    else
      echo "(no output)"
    fi
    
    show_sizes
    
    echo
    echo "Result:"
    echo "  WAL before: $(human "$BEFORE_WAL")"
    echo "  WAL after:  $(human "$AFTER_WAL")"
    
    if [ "$RC" -ne 0 ]; then
      echo
      echo "Failed: sqlite3 returned exit code $RC."
      echo "Likely causes: DB is busy, Codex is actively writing, or disk is critically full."
      exit "$RC"
    fi
    
    if [ "$AFTER_WAL" -lt "$BEFORE_WAL" ]; then
      echo
      echo "Success: Codex WAL was truncated."
    elif [ "$AFTER_WAL" -eq "$BEFORE_WAL" ]; then
      echo
      echo "No size change."
      echo "Possible reasons:"
      echo "  - WAL was already small/checkpointed."
      echo "  - Codex or another process currently has an active transaction."
      echo "  - Codex is still writing and recreated WAL content immediately."
    else
      echo
      echo "WAL grew during the run, probably because Codex was actively writing."
    fi
    
    echo
    echo "Done."
    

    To run as a cron job:

    # Run `crontab -e` in your shell and then add the following line at the end
    */15 * * * * /bin/bash "$HOME/trim-codex-wal.sh" > "$HOME/.codex-wal-clean.log" 2>&1
    

    Note that the main table is trimmed, but there might be other tables which still accumulate data over a much slower period of time, requiring the script below.

    The second script deletes the log and WAL files and kills all running Codex processes in order to immediately free disk space:

    #!/usr/bin/env bash
    # fix-codex-wal.sh
    # Deletes the logs-2.sqlite* files and sends SIGTERM followed by SIGKILL to all codex-related processes
    # in order to close the open file handles which will allow disk space to be freed
    set -u
    
    CODEX_DIR="${CODEX_HOME:-$HOME/.codex}"
    DB="$CODEX_DIR/logs_2.sqlite"
    WAL="$CODEX_DIR/logs_2.sqlite-wal"
    SHM="$CODEX_DIR/logs_2.sqlite-shm"
    
    PIDS=""
    
    add_pid() {
      case "${1:-}" in
        ''|*[!0-9]*) return ;;
        "$$") return ;;
      esac
    
      case " $PIDS " in
        *" $1 "*) ;;
        *) PIDS="$PIDS $1" ;;
      esac
    }
    
    bytes() {
      if [ -e "$1" ]; then
        stat -c '%s' "$1" 2>/dev/null || stat -f '%z' "$1" 2>/dev/null || echo 0
      else
        echo 0
      fi
    }
    
    human() {
      if command -v numfmt >/dev/null 2>&1; then
        numfmt --to=iec --suffix=B "$1" 2>/dev/null || echo "$1 bytes"
      else
        echo "$1 bytes"
      fi
    }
    
    show_sizes() {
      echo
      echo "Codex SQLite files:"
      for f in "$DB" "$WAL" "$SHM"; do
        printf "  %-45s %12s\n" "$f" "$(human "$(bytes "$f")")"
      done
    }
    
    show_processes() {
      if [ -z "$(printf '%s' "$PIDS" | tr -d ' ')" ]; then
        echo "None."
        return
      fi
    
      printf "%-8s %-8s %-6s %-10s %-12s %s\n" "PID" "PPID" "STAT" "TTY" "ELAPSED" "CMD"
    
      for pid in $PIDS; do
        ps -p "$pid" -o pid=,ppid=,stat=,tty=,etime=,cmd= 2>/dev/null
      done
    }
    
    echo "[LIVE RUN] Codex WAL/SHM cleanup"
    echo "Codex dir: $CODEX_DIR"
    echo "DB:        $DB"
    echo "WAL:       $WAL"
    echo "SHM:       $SHM"
    
    show_sizes
    
    echo
    echo "[1/5] Finding Codex-related processes..."
    
    for pid in $(pgrep -u "$USER" -x codex 2>/dev/null); do
      add_pid "$pid"
    done
    
    for pid in $(pgrep -u "$USER" -x node 2>/dev/null); do
      cmd="$(ps -p "$pid" -o cmd= 2>/dev/null || true)"
    
      case "$cmd" in
        *codex*|*.codex*|*openai*codex*)
          add_pid "$pid"
          ;;
      esac
    done
    
    echo
    echo "[2/5] Finding processes holding Codex WAL/SHM files..."
    
    if command -v lsof >/dev/null 2>&1; then
      for f in "$WAL" "$SHM"; do
        if [ -e "$f" ]; then
          for pid in $(lsof -t -- "$f" 2>/dev/null); do
            add_pid "$pid"
          done
        fi
      done
    
      for pid in $(
        lsof -nP +L1 2>/dev/null | awk -v dir="$CODEX_DIR" '
          index($0, dir "/logs_2.sqlite-wal") && /deleted/ { print $2 }
          index($0, dir "/logs_2.sqlite-shm") && /deleted/ { print $2 }
        '
      ); do
        add_pid "$pid"
      done
    else
      echo "lsof not found; cannot check open/deleted WAL handles."
    fi
    
    echo
    echo "[3/5] Processes to terminate:"
    show_processes
    
    if [ -n "$(printf '%s' "$PIDS" | tr -d ' ')" ]; then
      echo
      echo "Sending SIGTERM..."
    
      for pid in $PIDS; do
        echo "TERM $pid: $(ps -p "$pid" -o cmd= 2>/dev/null || echo unknown)"
        kill "$pid" 2>/dev/null || true
      done
    
      sleep 3
    
      STILL=""
    
      for pid in $PIDS; do
        if kill -0 "$pid" 2>/dev/null; then
          STILL="$STILL $pid"
        fi
      done
    
      if [ -n "$(printf '%s' "$STILL" | tr -d ' ')" ]; then
        echo
        echo "Some processes survived SIGTERM; sending SIGKILL:"
    
        printf "%-8s %-8s %-6s %-10s %-12s %s\n" "PID" "PPID" "STAT" "TTY" "ELAPSED" "CMD"
        for pid in $STILL; do
          ps -p "$pid" -o pid=,ppid=,stat=,tty=,etime=,cmd= 2>/dev/null || true
        done
    
        for pid in $STILL; do
          echo "KILL $pid: $(ps -p "$pid" -o cmd= 2>/dev/null || echo unknown)"
          kill -9 "$pid" 2>/dev/null || true
        done
      else
        echo "All targeted processes exited after SIGTERM."
      fi
    fi
    
    echo
    echo "[4/5] Deleting Codex WAL/SHM files..."
    
    if [ -e "$WAL" ]; then
      echo "Deleting: $WAL"
      rm -f -- "$WAL"
    else
      echo "Not found: $WAL"
    fi
    
    if [ -e "$SHM" ]; then
      echo "Deleting: $SHM"
      rm -f -- "$SHM"
    else
      echo "Not found: $SHM"
    fi
    
    show_sizes
    
    echo
    echo "[5/5] Final cleanup check..."
    
    if command -v lsof >/dev/null 2>&1; then
      echo "Remaining deleted Codex files still held open, if any:"
      lsof -nP +L1 2>/dev/null | grep -F "$CODEX_DIR" || echo "None found."
    else
      echo "lsof not found; skipping deleted-handle check."
    fi
    
    echo
    echo "Disk usage:"
    df -h "$HOME" 2>/dev/null || true
    du -xsh "$CODEX_DIR" 2>/dev/null || true
    
    echo
    echo "Done."
    
  6. Gerry9000 commented on Jun 19, 2026

    @Gerry9000

    plz fix

  7. Necmttn commented on Jun 21, 2026

    @Necmttn

    The log DB should have an explicit byte budget and retention policy. Durable usage/diagnostic events can stay queryable, while websocket/SSE TRACE payloads need sampling, level caps, and rotation independent of the state DB. A startup receipt with effective log level, DB byte cap, WAL checkpoint policy, and dropped-row count would make write pressure visible before SSD wear becomes the signal.


    Generated with ax.

  8. Nicolas0315 commented on Jun 21, 2026

    @Nicolas0315

    Additional native Windows check against current main:

    • Upstream source inspected: 63f009e9dad2e70454b7ed6434d8aa28dfb52b51
    • Windows 11 Pro 10.0.26200
    • Codex Desktop active during the sample

    Current source still attaches the persistent SQLite log layer with a global TRACE filter in both app-server and TUI:

    • codex-rs/app-server/src/lib.rs
      • log_db::start
      • .with_filter(Targets::new().with_default(Level::TRACE))
    • codex-rs/tui/src/lib.rs
      • log_db::start
      • .with_filter(Targets::new().with_default(Level::TRACE))

    I also checked the current log sink implementation:

    • codex-rs/state/src/log_db.rs
      • events are formatted into feedback_log_body, queued, then inserted by a background task.
      • the sink drops low-level opentelemetry_sdk TRACE/DEBUG, which helps one specific noisy target.
    • codex-rs/state/src/runtime/logs.rs
      • each batch inserts rows, then prunes inside the same transaction.

    That means DB size or retained row count can stay stable while insert/prune churn continues.

    Local aggregate-only measurement from .codex\logs_2.sqlite over 15 seconds, without reading log bodies:

    • retained rows before: 56,766
    • retained rows after: 56,766
    • retained row delta: 0
    • MAX(id) before: 191,131,613
    • MAX(id) after: 191,131,862
    • MAX(id) delta: 249 rows inserted in 15s
    • retained estimated bytes delta: +135,974 bytes

    Top retained target+level groups in this Windows sample were still dominated by the same categories already reported here:

    • codex_api::endpoint::responses_websocket TRACE
    • codex_otel.log_only INFO
    • codex_otel.trace_safe INFO
    • codex_core::stream_events_utils INFO
    • codex_client::transport TRACE
    • rmcp::service TRACE
    • codex_api::sse::responses TRACE
    • log TRACE

    This reinforces the earlier point: visible DB growth is not a reliable proxy for write pressure. Even with a stable retained row count, the append-and-prune path is still doing work. The low-risk patch shape still looks like filtering before formatting/queueing/inserting, not only pruning after insert.

    I also rebuilt a narrow local candidate patch against the current origin/main head:

    • base: f774455 (code-mode: linearize cell terminal state (#29286))
    • adds codex_state::log_db::default_filter()
    • uses the shared filter from app-server and TUI instead of Targets::new().with_default(Level::TRACE)
    • keeps global WARN+
    • keeps INFO+ for core Codex targets (codex_api, codex_app_server, codex_core, codex_mcp, codex_state, codex_tui)
    • raises known high-volume mirror/dependency targets to WARN+ (codex_otel.log_only, codex_otel.trace_safe, hyper_util, log, opentelemetry_sdk)

    Local validation:

    • cargo test -p codex-state sqlite_sink_default_filter_drops_low_value_logs -> 1 passed
    • cargo test -p codex-state log_db -> 7 passed
    • cargo check -p codex-app-server -p codex-tui -> passed
    • cargo fmt --check -> passed; stable rustfmt emitted repo-config warnings that imports_granularity is nightly-only
    • git diff --check -> passed

    Per docs/contributing.md, I am not opening a PR unless invited, but this latest local patch is PR-shaped if maintainers want this direction.

  9. cooltzy commented on Jun 22, 2026

    @cooltzy

    I hope the developers see this issue.

  10. kernbug commented on Jun 22, 2026

    @kernbug

    Shame on such low software quality with such high memory prices and SSD cost...

  11. 151 remaining items

  12. rodalpho commented on Jul 12, 2026

    @rodalpho

    Which version has it fixed? Is it in the desktop app?

  13. etraut-openai commented on Jul 12, 2026

    @etraut-openai
    Contributor

    Our series of fixes for this problem are in the latest shipping desktop app, CLI, and IDE extension.

  14. darlingm commented on Jul 12, 2026

    @darlingm
    Contributor

    Our series of fixes for this problem are in the latest shipping desktop app, CLI, and IDE extension.

    Not quite yet. 0.144.1 does not have #31789, #31790, #31791, and #31792. 0.145.0-alpha.4 includes them, so 0.145.0 likely will too.

    Also, please see #32496. A different root cause is causing about 1.3 TB/year in unnecessary writes, according to the numbers in the issue. (Much less than was on this issue.) Codex is continually rewriting models_cache.json and scope_v3.json.

    Would it be possible for OpenAI to add a regression test that ensures: (1) Codex during idle doesn't make more than a very minimal amount of I/O; and (2) Codex during usage doesn't absurdly increase its I/O? It might be best to split 2, into one version that fakes the model to test the harness in isolation, and one end-to-end test that makes sure new model/harness combinations don't come up with something that bumps it significantly. The latter could of course trip from random model behavior. If OpenAI logged and compared the values, they would not only see spikes but also worrisome trends in I/O usage that could be investigated.

    While at it, testing CPU and GPU usage similarly would be a big help. There are several issues about crazy high usage, often tracked down to animations being unnecessarily repeatedly updated a crazy number of times a second.

  15. gustavopch commented on Jul 12, 2026

    @gustavopch

    @etraut-openai Codex has been tracking this issue for me. I asked its opinion (GPT 5.6 Sol High), and it doesn't think this is fully solved:


    I retested the SQLite logging behavior on macOS. The upstream fixes have substantially reduced the original churn, but the current desktop build still persists noisy TRACE/DEBUG activity at a much higher rate than when the local filter is enabled.

    Scenario App/engine Mitigation Duration MAX(id) change Rate Net retained rows
    Earlier affected build (July 1) codex-cli 0.142.5 Off 96.9s +3,212 33.1 IDs/s +57
    Current desktop build (July 12) App 26.707.51957, bundled codex-cli 0.144.0-alpha.4 Off 41.45s +305 7.36 IDs/s +9
    Same current build App 26.707.51957, bundled codex-cli 0.144.0-alpha.4 On 41.89s +19 0.45 IDs/s +1

    The mitigation is a local SQLite trigger that rejects TRACE/DEBUG records and known noisy targets. It does not block INFO, WARN, or ERROR logging.

    This shows two things:

    1. The upstream fixes improved the unmitigated rate substantially: 33.1 -> 7.36 IDs/second.
    2. On the current build, enabling the filter still reduces churn by about 16x: 7.36 -> 0.45 IDs/second.

    During the current-build test without mitigation, 305 IDs were allocated while only 9 additional rows remained. That is consistent with continuing insert-and-prune activity.

    The newly retained records were dominated by:

    • codex_app_server::outgoing_message / TRACE: 133
    • codex_api::sse::responses / TRACE: 35
    • codex_mcp::connection_manager / TRACE: 30
    • hyper_util TRACE/DEBUG: 67

    No log bodies were read during these measurements.

    Is this residual TRACE/DEBUG churn considered expected after the fixes? If not, is desktop app 26.707.51957 missing part of the fix, and which desktop or bundled CLI version should contain it?

  16. 0xdevalias commented on Jul 13, 2026

    @0xdevalias

    Is this residual TRACE/DEBUG churn considered expected after the fixes? If not, is desktop app 26.707.51957 missing part of the fix, and which desktop or bundled CLI version should contain it?

    @gustavopch It seems like you're still using 0.144.0-alpha.4 in those tests; whereas an earlier comment from @darlingm suggests that there are more fixes that will land in 0.145.0-alpha.4+, that sound like they should address some if not all of those:

    Not quite yet. 0.144.1 does not have #31789, #31790, #31791, and #31792. 0.145.0-alpha.4 includes them, so 0.145.0 likely will too.

  17. daryll-swer commented on Jul 13, 2026

    @daryll-swer

    We should be getting partial refunds or a non-expiring limit reset as compensation for this.
    --@mtsitzer

    So are we getting some refunds/limit resets or ANY compensation for killing our SSDs? For weeks or nah?

  18. gustavopch commented on Jul 13, 2026

    @gustavopch

    @0xdevalias Ah, thanks for pointing it out. I just updated from 26.707.51957 to 26.707.61608 and the bundled Codex CLI is still 0.144.0-alpha.4. I guess we'll have to wait a few more days.

  19. HGT158 commented on Jul 20, 2026

    @HGT158

    So has this issue been fixed in the latest version of the Codex desktop app?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    CLIIssues related to the Codex CLIbugSomething isn't workingperformance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions