Skip to content

perf: storage phase 1 — batched index writes, log compaction, redb index, bounded client history, async cold reads - #591

Closed
Tryanks wants to merge 6 commits into
mainfrom
feat/storage-refactor
Closed

Tryanks wants to merge 6 commits into
mainfrom
feat/storage-refactor

Conversation

@Tryanks

@Tryanks Tryanks commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Phase 1 of the storage work in #535, as one PR. Closes #520, #521, #522, #523, #525; the first slice of #524 is included and the rest of #524 follows in this PR.

Each slice is one commit, in dependency order. The per-slice evidence is below; the measurements were taken on copies of the maintainer's data (2,950–2,960 thread logs, 9.44 GB) and the live data directory was only read.

1. Batch thread lifecycle commands; the store writer is the only index writer (#520, #521)

  • Protocol 9: ArchiveSessions, UnarchiveSessions, DeleteSessions { session_ids, remove_worktrees } replace the single-thread commands. Archive, unarchive, delete, delete project and the settle cascades persist once per command instead of once per thread.
  • External import prepares each thread off the host loop, installs its log through the store writer and indexes the run in one batch; the host no longer re-reads the index from disk afterwards (which dropped updates still queued in the writer). A project deleted mid-import no longer gets sessions written against it.
  • Not changed: deleting a thread still leaves its descendants indexed with a dangling parent (Deleting a thread leaves its descendant threads with a dangling parent #518).

2. Compact thread logs at turn end under a layout epoch (#522)

  • A log keeps only the last cumulative turn_changes_updated per turn and merges adjacent deltas of one item (first timestamp kept, at most 512 KiB of text per record). The rules live in crates/core/src/session/compaction.rs and follow the fold's own turn resolution; a removed record that raised the open turn's clock is kept unless the next record restores it. Compaction is idempotent.
  • Timeline derives PartialEq, private continuation state included; a rewrite is installed only if the compacted records as they will be read back fold equal to the originals, from a strict read (any blank, truncated, non-UTF-8 or unparseable line leaves the log untouched).
  • A rewritten log starts with {"layout_epoch":N}, replaced with the body by one fsync + rename. Protocol 10: a session cursor is (epoch, position) in subscriptions, pages, snapshots and live events, with SessionHistoryReset; a cursor from another layout gets a baseline the client installs in one step, keeping scroll position. A resident log keeps the layout it was loaded in, so connected clients see no reset at turn end.
  • Runs in the store writer's order after TurnCompleted and before a log leaves residency; a startup pass compacts cold headerless logs one at a time. The first rewrite keeps <id>.jsonl.orig (a hard link) until the next start.
  • Orchestrate reads use the resident fold or a writer barrier, so they no longer race the TurnCompleted append.
  • Evidence: on a copy of the data, a headless host's pass took 41 s: 2,104 rewritten (8.78 GB → 3.81 GB), 850 unchanged, 0 rejected; total 4.47 GB. Every rewritten log folds equal to its original (separate process and parser); a second pass changed nothing; on restart the originals were removed. Cost: a turn-end rewrite of a 125 MB log makes other sessions' appends wait 120–170 ms; peak RSS during the pass 2.26 GB (the 193 MB original). Real Claude Code turns rewrote the log after each turn while the subscribed client stayed on its epoch.
  • Left as is: the wire's window rules (superseded_turn_changes stubs, wire_window delta merge), since clients fold a turn-aligned suffix and positions must not move.

3. Keep the project and thread index in redb (#523)

  • tcode.redb replaces sessions.json; tables projects, sessions, meta, values are the rows' JSON. Only a host opens it. Each change writes the rows it touches in one transaction on the blocking pool; the writer batches back-to-back index writes.
  • redb 4.3's experimental-multiprocess in SingleWriter mode (maintainer's decision, Principle 8): another process can read while the host runs; a second host on the same data dir fails explicitly (headless exits 1; desktop shows a critical alert and exits 1); a relaunch waits up to 15 s.
  • One-time import: sessions.json is parsed with today's tolerance, written to tcode.redb.tmp in one transaction, read back and compared, then renamed to sessions.json.migrated. A row this build cannot decode is skipped and kept. The file is compacted at open once it has doubled.
  • Departure from the issue: redb refuses Durability::None in multi-process modes, so every commit is durable: one meta upsert went from 3.8 ms on the executor thread writing 2.19 MB with no fsync, to 7.0 ms on the blocking pool writing about 120 KB with two fsyncs. A data dir on a file system without byte-range locks fails to start with an explicit error.
  • Evidence: the maintainer's sessions.json (23 projects, 2,957 sessions) imports with every row equal field for field; SIGKILL during import ×100 and during commits ×60 all recover; second host, read-only open and relaunch shown with real processes. Desktop binary +808 KiB; Cargo.lock gains one edge and no crate.

4. Bound the history a client holds (#525)

  • The store keeps replicas for the selected thread and the 4 threads left most recently; a deleted thread's are dropped at once. Pages read far above the tail are dropped after 30 s at the tail or when the thread is left; the tail's own pages are kept, so nothing is refetched and the list does not move.
  • A count rather than the issue's 5-minute TTL: a non-selected thread receives no records, so keeping it only saves one baseline window; a TTL bounds age, not memory. 50 threads visited: 26.1 MB held before, 2.6 MB after.
  • Not changed: prepending a page still refolds the held window (38 ms with a 17 MB thread fully held); what is bounded is what accumulates afterwards.

5. Read a cold thread's log off the host mailbox (#524, first slice)

  • One hydration per session: the store writer opens the log in its queue order and records its length (no parse), so the read holds exactly the appends queued before it; the parse and fold run on the blocking pool. Records accepted meanwhile are held in order and take the positions after the snapshot's end; waiting windows go out first, ending where the snapshot ends, then the held records, so each subscriber sees every record once. A subscription's ack waits for its window (the mux drops a request-scoped event after the ack). The log is kept only if the session is still resident.
  • Measured: while a 125 MB log is opened, an unrelated session's round trips went from 340 ms median / 770 ms max to 0.03 ms / 0.2 ms; open latency itself is unchanged. A cold open now waits behind the writer queue.

Tests

Each slice's tests drive the host through the pipe or a real spawn_host and wait for the state they assert. Added coverage: batched commands persist and reopen; compaction rules and the gate (case tables), strict-read rejection, header round trip, .orig rule; cursors from another layout get a baseline, resident pinning across a rewrite, crash states through startup, stale cursor through the mux; legacy index imports, interrupted import states, undecodable rows kept, second host refused; thread release and page trimming; shared cold reads, the event that starts a read, leaving before the read completes, two mux clients. Removed: session_index_upserts_orders_and_removes_only_the_selected_session (store-level ordering, now the host's sort; coverage moved to the index and cascade tests) and reconnect_and_mismatched_tail_preserve_exactly_one_copy_of_each_record (half asserted the blank-and-refill the epoch change removes; exactly-once is covered by the reconnect test). Drivers fixed, assertions kept: tests that waited on status before asserting held history, called the synchronous snapshot in the selection turn, or used an ack as a fence for windows.

Checks run

On the final tree: cargo fmt --all --check, cargo clippy --workspace --all-targets --locked -- -D warnings, cargo nextest run --workspace --locked (908 passed, 6 skipped), cargo machete, Web (wasm32) and iOS simulator checks with -D warnings. Each slice also passed CI on all platforms as a stacked PR (#585–#590). Not run locally: Android, Windows, Linux. Visual checks on macOS: the epoch reset in both themes at wide and narrow widths; the startup-failure alert; the scroll-up and trim cycle.

…ore writer

Archive, unarchive and delete now take a list of threads and persist with
one index write per command; deleting a project and the settle cascades do
the same. External import prepares each thread off the host loop, installs
its log through the store writer and indexes the whole run in one batched
upsert, so nothing outside the writer rewrites sessions.json while the
host runs and the in-memory index is no longer replaced from disk.

PROTOCOL_VERSION 9: ArchiveSessions, UnarchiveSessions and DeleteSessions
replace the single-thread commands.

Closes #520
Closes #521
A thread's log is rewritten with only the last cumulative turn diff per
turn and with adjacent deltas of one item merged. The rules live in core
and follow the fold's own decisions; a rewrite is installed only when the
compacted records fold equal to the originals, private continuation state
included, and only from a strict read of the file.

Rewriting renumbers records, so a session cursor is now a layout epoch and
a position. The epoch is the log's first line and is replaced with the
body in one rename. A resident log keeps the layout it was loaded in, so
connected clients are not disturbed; a cursor from another layout gets a
baseline that replaces what the client holds in one step.

Compaction runs in the store writer's order after a turn completes and
before a log leaves residency, and a background pass at startup rewrites
existing cold logs. The first rewrite keeps the original as
<id>.jsonl.orig until the next start.

PROTOCOL_VERSION 10.

Closes #522
tcode.redb replaces sessions.json. Each index change writes the rows it
touches in one transaction on the blocking pool, and the store writer
commits back-to-back index writes together, instead of re-reading and
rewriting the whole file for every change.

The host opens the database as redb's single writer under
experimental-multiprocess: another process can read it while the host
runs, and a second host on the same data directory fails with an explicit
error instead of silently overwriting. A relaunch waits for the outgoing
instance to close it.

sessions.json is imported once into a temporary database, read back and
compared, then kept as sessions.json.migrated. A row this build cannot
decode is skipped and left in place rather than failing the whole index.

Closes #523
A client kept every record of every thread it had visited for its whole
lifetime. It now keeps replicas for the selected thread and the four
threads left most recently, and drops a deleted thread's at once. A
thread that is not selected receives no records, so holding it longer
only saves one baseline window.

Pages read far above the tail are dropped once the reader has stayed at
the tail for 30 seconds, or leaves the thread. The pages the tail itself
needed are kept, so nothing dropped is fetched again and the list does
not move.

Closes #525
archiving_the_viewed_child_returns_to_its_parent waited only for the
parent's status, which arrives on its own topic, so on a slow runner the
child was opened before any parent record was held and there was nothing
to keep. The driver now waits for the conversation to be on screen, as a
user would see it, and the test also requires the parent to resume from
its cursor.
Opening a cold thread, the first event of an uncached thread, and history
pages or full outputs of a thread nobody has open each parsed and folded
the whole log on the host mailbox, stalling every other session for as
long as that took.

A session now has at most one hydration. The store writer opens the log
in its queue order and records its length, so the read holds exactly the
appends queued before it; the parse and fold run on the blocking pool.
Records accepted meanwhile are held in order and take the positions after
the snapshot's end. When the log is resident the waiting subscription
windows go out first, each ending where the snapshot ends, then the held
records, so a subscriber sees every record once. A subscription's
acknowledgement waits for its window. The log is kept only if the
session is still resident.

Part of #524
@Tryanks

Tryanks commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Converted to draft: the maintainer has decided that event bodies go into the database too and that SQLite is to be re-evaluated against redb for the whole store, which supersedes the "no SQLite, bodies stay in files" decisions in #535/#523. Not to be merged in this form; the evaluation decides what of this is kept.

@Tryanks

Tryanks commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Superseded by #594, which moved metadata and event logs into a Turso database as the maintainer re-scoped the work. The five commits stay on feat/storage-refactor in case any part (turn-end compaction, client history bound) is wanted later.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

External import races the store writer on sessions.json

1 participant