Skip to content

fix(server): thread list reads no longer freeze the server on large databases - #14703

Open
saphid wants to merge 2 commits into
pingdotgg:mainfrom
saphid:perf/v2-sqlite-reader-worker
Open

saphid wants to merge 2 commits into
pingdotgg:mainfrom
saphid:perf/v2-sqlite-reader-worker

Conversation

@saphid

@saphid saphid commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Problem

On a large database, reading the thread list freezes the whole server for seconds. getShellSnapshot runs on every client connect and once a minute from the pull request discovery pass. It holds the only SQLite connection, on the main thread, for the whole read, so websocket keepalives, other requests, and writes all wait. On a clone of a 33 GB database (about 3,400 threads), each snapshot blocked the event loop for 1.4 to 2.2 s. Clients miss keepalives and reconnect, and each reconnect reads the snapshot again. Details in #14701.

Change

  • nodeSqliteClient gets a readerWorker option and a readOnly marker. A transaction wrapped in readOnly runs on a second, read-only connection in a worker thread, so it blocks neither the event loop nor the writer. Everything unmarked keeps using the existing connection.
  • Callers opt in, and only where a snapshot that keeps moving while it is read is acceptable:
    • the websocket connect snapshot, the archived-threads snapshot (ws.ts) and the HTTP snapshot (http.ts). Each still reads projects, threads and the application sequence in one transaction.
    • the once-a-minute pull request discovery pass (ThreadPullRequestService). It passes the snapshot's own sequence to its existing stale-write guard.
  • getShellSnapshot itself is not marked. Checks that act on the result, such as checkpoint restore safety and project deletion, read on the writer exactly as on main.
  • The server turns the reader on for its database file (Sqlite.ts).

How the reader behaves:

  • Startup: the worker starts on first use, from inline source, so it also works from bundles and the single-executable server build. It opens the file read-only (which no SQL statement can undo, unlike PRAGMA query_only) and checks the database is in WAL mode, then reports ready.
  • Fallback: the reader turns off for the life of the process, with one warning in the log, and every read uses the writer as before, if:
    • the worker cannot start, cannot open the file, finds a non-WAL database, or has not reported within readerStartupTimeout (10 s by default);
    • the database is in memory (tests).
    • readOnly code inside an open write transaction also stays on the writer.
  • Cancellation: a request cancelled while the worker is still starting leaves at once. Nothing is locked yet at that point.
  • Writes and BEGIN: a write on the reader fails with "attempt to write a readonly database". Main now opens writer transactions with BEGIN IMMEDIATE; the reader rewrites any BEGIN to a plain deferred BEGIN, so marked reads never take the write lock.
  • Crashes: a transaction stays on the worker it began on. If that worker exits, the rest of its statements fail with a normal SqlError instead of continuing on a new worker. COMMIT and ROLLBACK then succeed as no-ops, which is safe because the reader never writes. The next read starts a new worker.

Scope and approval

Fixes #14701, triaged as a bug by @juliusmarminge on 2026-10-02 (triage comment). Of the three fixes suggested there, this is "don't hold the only connection for the whole read", and it takes the discovery pass off the main thread. The other two (store the turn-item counts on the thread row; give the discovery pass a narrower query) would make the read itself cheaper. They are separate changes and are not attempted here. The change is limited to the SQLite adapter, the four marked callers, and the server's layer config.

Verification

A/B on a clone of a real 33 GB statev2.sqlite (APFS clone of a consistent backup; 3,347 threads; real getShellSnapshot through ProjectionStore.layer and makeSqlitePersistenceLive). Base and candidate ran in alternating rounds, 12 snapshots each. The main-thread gap was recorded with a 2 ms setInterval. In parallel, a probe ran SELECT 1 on the writer every 25 ms.

Base (de95adc) Candidate
Longest main-thread gap per snapshot 1,411 to 2,240 ms 13 to 17 ms
Snapshot duration 1,502 to 2,375 ms 1,464 to 1,918 ms
Writer probes completed during a snapshot 1 (it waits for the snapshot) 62 to 80

The A/B above ran before the reader switched from PRAGMA query_only to a read-only open. One rerun of the candidate after that switch, with the same setup, gave a longest gap of 12 to 21 ms over 6 snapshots. monitorEventLoopDelay under-reported the base stall in this setup (max about 30 ms). The timer gap and the probe log agree with each other and with the 1.4 to 2.6 s stalls seen in live server traces.

Stress test on GCP. A synthetic 36 GB statev2.sqlite shaped like ours: 3,441 threads, 253k turn items, and a heavy tail of 8 threads with over 2,000 items. Each server ran on Linux with a fake Codex provider, and a load driver supplied websocket clients, writers, and reconnect storms. Base and candidate alternated on the same VM, 3 repetitions each. Values are medians [min, max].

Shape Worst main-thread pause, 10x write load 100-client reconnect storm: command p99 Writes during that storm
2 vCPU, 8 GB, network disk base 1,847 [680, 2,058] ms, candidate 620 [533, 764] ms base 60 s (timeouts), candidate 420 ms base 0, candidate 20.3 events/s
8 vCPU, 32 GB, local NVMe base 344 ms, candidate 260 ms base 60 s (timeouts), candidate 78 ms base 0, candidate 3.9 events/s
8 vCPU, 64 GB, database cached base 322 ms, candidate 251 ms base 60 s (timeouts), candidate 71 ms base 0, candidate 3.8 events/s
  • Only the 8 GB shape reproduces stalls over 1 s on base. On the larger shapes the database stays cached, so the query is fast there. On the 8 GB shape the load generator was itself CPU-starved, so its ping timings are approximate. The storm and pause measurements come from the server.
  • The snapshot is identical between base and candidate on every shape (canonical hash).
  • Failure cases, all passing:
    • Killing the reader worker mid-query six times: each in-flight read failed cleanly within 10 s, the next read started a new worker, and the server stayed up.
    • SIGTERM and kill -9 during a read: clients recovered in about 2.5 s.
    • Database integrity after disk-full and kill -9: PRAGMA quick_check returned ok.
    • I/O contention and CPU starvation: the candidate was equal to or better than base.
  • Costs:
    • Total server CPU roughly doubles (for example 0.18 to 0.40 core-seconds per second on the NVMe shape), from the second connection and copying rows back to the main thread.
    • Peak memory on the 8 GB shape rose by about 420 MB under steady load. During storms the candidate used less memory than base.

The A/B and stress numbers were measured at 8330f48362a, where getShellSnapshot was marked inside the store. Those runs exercised the snapshot routes, which are marked the same way now. The discovery pass was not measured on the reader.

Focused tests on the current head

  • vp test run packages/shared/src/nodeSqliteClient.test.ts: 21 passed, 13 of them added by this PR.
    • With readerWorker: false, three of the core tests fail: a writer commit during an open read deadlocks, a write on the reader is accepted, and a crashed reader reports no error.
    • The two fallback tests (unreadable database path, non-WAL database) fail on the version without the startup check.
    • The three startup tests fail on the version without a deadline, an interruptible wait and a single turn-off: a worker that never reports, cancellation during startup, and two concurrent first reads that must log one warning. Two of them hang until the 60 s test timeout.
    • A race test fails without the final recheck: two staggered first reads, where the worker reports ready after one read's deadline has retired it. The second read must not receive the terminating worker.
  • vp test run on ThreadPullRequestService.test.ts, Sqlite.test.ts, ProjectionStore.test.ts, PullRequestSyncReactor.test.ts, SessionStore.test.ts, ProjectSettingsUpgrade.integration.test.ts, OrchestratorReplayRecovery.integration.test.ts and OrchestratorReplayRestartBackgroundNote.integration.test.ts: 81 passed. Several use a real database file, so they go through the reader.
  • tsc --noEmit for packages/shared and apps/server: no errors. vp lint, vp fmt --check and knip --exports for the touched workspaces: clean.

Runtimes: an inline-source worker that opens node:sqlite and reads rows works under the desktop app's Electron 44.4.2 (ELECTRON_RUN_AS_NODE) and in a Node 26.7 single-executable build.

What this can make worse

These are known costs or exposures, not defects found in testing:

  • WAL growth: writes now continue during a snapshot, so the WAL grows while a read is open. Back-to-back reads on the single reader may leave little room for a checkpoint to reset it. journal_size_limit (32 MB) trims the file only after a reset, so it is not a cap. The stress runs peaked at 49 MB during reconnect storms, against 4 MB on base.
  • CPU and memory:
    • Total server CPU roughly doubles, because rows are produced in the worker and then copied to the main thread.
    • Peak memory on the 8 GB shape rose by about 420 MB under steady load. During storms the candidate used less memory than base.
    • Coalescing identical concurrent snapshots, or an opt-out, would reduce this. Neither is in this PR.
  • Queueing and cancellation: marked reads queue on one reader. A read from a client that has disconnected keeps running until it finishes. Both behaviours are the same as before, except that the queue is now separate from the writer.
  • Remaining pauses: the candidate still pauses for 200 to 800 ms under load. Those come from other work: the shell read itself takes about 1 ms on the main thread with this change, against about 180 ms without it. Other reads that scan every thread still use the writer, for example the settlement sweep (getSettlementCandidates) and the pull request sync scan (getThreadsWithPullRequests). Each can be marked separately once measured.

Not checked:

  • a unit test that drives the websocket or HTTP snapshot routes (the stress runs exercise them end to end, and the nested-transaction test covers their pattern);
  • real providers (the stress runs use a fake Codex provider);
  • Windows;
  • mobile and web clients (the server is the only side that changed);
  • a stress rerun after the per-caller marking and the startup check.

Reviewed by GPT-6 Astra (cross-vendor): approved the original change after two rounds; a risk review of the merged branch led to the startup check and the per-caller marking; re-reviews of those led to the startup deadline, the interruptible wait and the single turn-off, and then to the late-ready recheck. The recheck has a test that fails without it but was not reviewed again. Macroscope found that SQL could turn query_only back off; the reader opens read-only. The 2026-10-05 rebase was independently reviewed by GPT-6.1 Sol with no actionable findings.

Claude Opus 5.5 in T3 Code (Claude Code harness).

🤖 Generated with Claude Code

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Oct 2, 2026
Comment thread packages/shared/src/nodeSqliteClient.ts Outdated
@macroscopeapp

macroscopeapp Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR adds a substantial worker-thread SQLite execution path and enables it by default for the production server, affecting transaction routing, concurrency, resource usage, and snapshot behavior. It also adds a line-level lint suppression, so the change requires human review.

No code changes detected at 3f31329. Prior analysis still applies.

You can add or adjust custom eligibility rules. Learn more.

@juliusmarminge juliusmarminge added the macroscope-review Opt PRs made by unvouched contributors in for Macroscope review. Vouched contributors auto-reviews label Oct 2, 2026 — with ChatGPT Codex Connector
@juliusmarminge
juliusmarminge deleted the branch pingdotgg:main October 2, 2026 19:23
@juliusmarminge juliusmarminge added the triage:keep-open Keeps this PR open despite not necessarily passing the contribution guide fully label Oct 2, 2026
@juliusmarminge juliusmarminge reopened this Oct 2, 2026
@juliusmarminge
juliusmarminge changed the base branch from t3code/codex-turn-mapping to main October 2, 2026 20:35
Comment thread packages/contracts/src/orchestrationV2.ts
@juliusmarminge
juliusmarminge force-pushed the perf/v2-sqlite-reader-worker branch from 8330f48 to f24db69 Compare October 2, 2026 20:53
@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Warning

Review limit reached

Only developers with an assigned seat can use this organization's usage-based review budget, and seats here are assigned manually. Ask an admin to assign a seat, or change the review continuation mode in Billing.

Next included review available in 21 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration
  • Configuration used: Path: .coderabbit.config.ts
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 6df42db1-f6c4-4670-8fbf-5f98417e73e1
📥 Commits

Reviewing files that changed from the base of the PR and between 6a3edae and 3f31329.

📒 Files selected for processing (6)
  • apps/server/src/orchestration-v2/ThreadPullRequestService.ts
  • apps/server/src/orchestration-v2/http.ts
  • apps/server/src/persistence/Sqlite.ts
  • apps/server/src/ws.ts
  • packages/shared/src/nodeSqliteClient.test.ts
  • packages/shared/src/nodeSqliteClient.ts

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Path: .coderabbit.config.ts
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 8e44f7f7-abb9-4825-b03c-847b0521778a
📥 Commits

Reviewing files that changed from the base of the PR and between e7aa57e and 6a3edae.

📒 Files selected for processing (3)
  • apps/server/src/orchestration-v2/http.ts
  • apps/server/src/persistence/Sqlite.ts
  • apps/server/src/ws.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 1 remain after this review.


📝 Walkthrough

Walkthrough

The SQLite client can route effects marked read-only to a reader worker for eligible file-backed databases. Server shell snapshot transactions now use this routing, and the tests cover worker startup, transaction behavior, and fallback to the writer.

Changes

SQLite read-only snapshots

Layer / File(s) Summary
Reader worker and read-only routing
packages/shared/src/nodeSqliteClient.ts
The client adds reader-worker configuration, a read-only effect wrapper, a worker-backed SQLite connection, and reader startup and shutdown handling.
Read routing and fallback validation
packages/shared/src/nodeSqliteClient.ts, packages/shared/src/nodeSqliteClient.test.ts
Marked read-only work uses the reader when available. Tests cover snapshot behavior, write rejection, nested transactions, startup failures, and fallback to the writer.
Server snapshot read routing
apps/server/src/persistence/Sqlite.ts, apps/server/src/orchestration-v2/*, apps/server/src/ws.ts
File-backed clients enable reader workers. Shell snapshot transactions use the read-only wrapper.

Priority: ⬆️ High

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Bug fix · Severity of issue fixed: High

Sequence Diagram(s)

sequenceDiagram
  participant getShellSnapshot
  participant NodeSqliteClient
  participant ReaderWorker
  getShellSnapshot->>NodeSqliteClient: mark snapshot transaction read-only
  NodeSqliteClient->>ReaderWorker: execute read request
  ReaderWorker-->>NodeSqliteClient: return query result
  NodeSqliteClient-->>getShellSnapshot: return snapshot data
Loading

Suggested reviewers: juliusmarminge

Merge Risk: ⚪ Minimal · up to 6a3ed

Eligible snapshots can use the reader worker, while unavailable-reader cases retain the existing writer path. No actionable merge-blocking regression is established.

Architecture Summary

Architecture risk: 🔵 Low · up to 6a3ed

The change affects 2 systems.

Changed systems: apps/server, packages/shared

Architecture concerns
No architecture-level concerns identified.

Review details

Systems and components

  • observed — apps/server (service) was modified; 4 changed files map to changed impact.
  • observed — packages/shared (library) was modified; 2 changed files map to changed impact.

Before / after behavior

  • observed — Modified behavior in apps/server/src/orchestration-v2/ThreadPullRequestService.ts: Adds the NodeSqliteClient import used to make snapshot reads read-only.
  • observed — Modified behavior in apps/server/src/orchestration-v2/ThreadPullRequestService.ts: Full active-thread snapshots now run through NodeSqliteClient.readOnly; the existing location: "active" and unsettledOnly selection is unchanged. Single-thread snapshots do not use this path.
  • observed — Modified behavior in packages/shared/src/nodeSqliteClient.test.ts: Adds worker-thread, deferred, fiber, logger, and test-clock imports used by the new reader-worker tests.
  • observed — Modified behavior in packages/shared/src/nodeSqliteClient.test.ts: Adds a mock that records workers and can simulate silent startup or a late ready message on termination, plus helpers for creating a WAL-backed reader client and counting entries.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly describes the main change: preventing large-database thread-list reads from freezing the server.
Description check ✅ Passed The description includes the required Problem, Change, Scope and approval, and Verification sections. It explains the issue, implementation, scope, test results, and known limitations.
Linked Issues check ✅ Passed [#14701] requires large thread-list reads to leave the event loop and writer available. The PR routes websocket, archived-thread, HTTP, and periodic discovery snapshots through read-only worker transa…
Out of Scope Changes check ✅ Passed The SQLite worker, caller markings, server configuration, and related tests support [#14701]. The supplied change summary identifies no unrelated changes.
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@saphid

saphid commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Review requested

Date (UTC) Reviewer Where
2026-10-03 Julius Discord DM

Logged so this PR shows when a maintainer was asked to review it.

@saphid
saphid force-pushed the perf/v2-sqlite-reader-worker branch from e7aa57e to 6a3edae Compare October 6, 2026 11:47
github-actions Bot and others added 2 commits October 7, 2026 03:33
…atabases

getShellSnapshot held the only SQLite connection on the main thread for
the whole read, blocking keepalives, requests and writes for seconds on
large databases. nodeSqliteClient gets a readerWorker option and a
readOnly marker: a transaction wrapped in readOnly runs on a second
connection in a worker thread. getShellSnapshot and the three routes that
wrap it in a transaction are marked.

The reader connection is opened read-only, because PRAGMA query_only can
be turned off by a statement on the same connection.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cannot start

The reader worker opened the database lazily, so a read-only open failure
broke every thread list read for good, and a database outside WAL mode
let a reader's shared lock stall the writer. The worker now opens the file
read-only at startup, checks for WAL, and reports ready before any read is
routed to it. If it cannot, does not report within readerStartupTimeout
(10 s), or a request is cancelled while it starts, reads use the main
connection as before; the reader turns off once, with one warning. A late
"ready" from a worker another read already retired is not handed out.

getShellSnapshot is no longer marked inside the store. Callers opt in
where a snapshot that keeps moving is fine: the three snapshot routes and
the once-a-minute pull request discovery pass, which guards its writes
with the snapshot's sequence. Restore-safety and project-deletion checks
read on the main connection, as on main.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@saphid
saphid force-pushed the perf/v2-sqlite-reader-worker branch from 6a3edae to 3f31329 Compare October 6, 2026 16:33
@saphid

saphid commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Review requested

Date (UTC) Reviewer Where
2026-10-06 Julius Discord DM

Logged so this PR shows when a maintainer was asked to review it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

macroscope-review Opt PRs made by unvouched contributors in for Macroscope review. Vouched contributors auto-reviews size:L 100-499 changed lines (additions + deletions). triage:keep-open Keeps this PR open despite not necessarily passing the contribution guide fully vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Server: the thread list read blocks the event loop and every other database call for seconds on large databases

2 participants