Skip to content

feat: keep thread metadata and event logs in a Turso database - #594

Merged
Tryanks merged 4 commits into
mainfrom
feat/storage-turso
Oct 6, 2026
Merged

Tryanks merged 4 commits into
mainfrom
feat/storage-turso

Conversation

@Tryanks

@Tryanks Tryanks commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

sessions.json and every <id>.jsonl move into one Turso database, tcode.db, owned by SessionStore. Closes #520 and #523 (as re-scoped: event bodies go into the database too; the engine is Turso by the maintainer's decision). Supersedes draft #591.

Behaviour

  • Schema (PRAGMA user_version = 1): projects(id, body) and sessions(id, body) hold each row's serde JSON unchanged; events(session_id, position, line) holds each raw log line including its delimiter, so CRLF, blank lines, invalid UTF-8 and a missing final newline survive and read_event_log (export) is byte-identical to the source. read_events still parses tolerantly.
  • Writer. One writer connection; synchronous = FULL and, on Apple, fullfsync = ON, both read back at open. The store writer runs on its own thread and commits its queued database writes as one transaction (≤ 256 writes / ~4 MiB); settings, secrets, caches and Flush are ordered boundaries. A failed commit is reported once per write with its existing RuntimeError, never replayed; Flush now carries a Result, and the first failure makes every later flush fail, so a released log is kept in memory rather than evicted on a failed barrier. An external import, a Tcode-export import and a fork each commit events and metadata in one transaction.
  • Panics. Turso asserts invariants by panicking; every store operation runs under catch_unwind, a panic or an unconfirmed rollback puts the store into a shared failed state, further operations error, nothing is checkpointed, and the host shuts its providers down and reports the error.
  • Ownership. tcode.lock is taken with File::try_lock for the host's lifetime before any migration artifact is touched; Turso's own exclusive file lock is a second guard. A second host gets ResourceBusy naming the data dir: headless exits 1, the desktop shows a critical prompt and exits 1. With the macOS relaunch marker present the open retries for up to 15 s; the marker is only consumed after ownership is acquired. open_at stays directory-only, so subcommands never open the database.
  • Migration, on first start without tcode.db: build tcode.db.migrating (schema, user_version 0); import the index with today's tolerance (object or bare array, migrate_index once, duplicates fail the migration, a corrupt index is preserved as .corrupt-*); import every *.jsonl found in the directory (orphans kept) in ~8 MiB transactions and verify each thread segment by segment against the file; read the metadata back field for field; set user_version = 1 last; checkpoint (result checked, WAL must be 0 bytes), drop every handle, reopen and re-check counts, fsync the file; write the archive list, fsync, rename to tcode.db, fsync the directory; then move the sources into legacy/ (idempotent, never overwrites). Before the rename nothing is published and the next start discards the staging files; after it the next start only finishes the archival. A user_version of 0 or above 1, or an unopenable file, is refused with a message naming the file and the sqlite3 CLI; the file is never deleted. A fresh install creates its database through the same staged path.
  • Callers. Export, import, fork, delete, orchestrate reads and session search go through the store; search is keyed by a store-owned per-session mutation generation instead of the file's (len, mtime). Nothing opens sessions.json or <id>.jsonl any more except the migration. The startup whole-index rewrite is gone; reads are fallible internally (no protocol change).
  • Fixed on the way: deliver_child_callback read a child's transcript while its last appends were still queued; it now uses the cached fold or a store barrier.

Evidence

  • Real data, on an APFS clone of the maintainer's data dir (3,000 logs, 9.54 GB, 7.14 M lines; 24 projects, 3,000 sessions): migration 73 s, every thread byte-identical, tcode.db 10.6 GB, WAL 0 after shutdown, sqlite3 PRAGMA integrity_check = ok; restart in 1 s with nothing migrated; threads render in the desktop app; a second headless host and a second desktop instance are refused with exit 1.
  • SIGKILL through the real writer and barrier, 50 runs: 180,833 acknowledged records all present, gapless, no torn row; 48 runs crossed an auto-checkpoint; integrity check ok on a preserved db+WAL copy each time. SIGKILL during migration, 50 runs: either no tcode.db and sources intact, or a complete one with every source byte-identical in place or in legacy/.
  • Disk-full (64 MB image): the failing batch reported once per write, later flushes failed, the cached history was kept, the process recovered after space was freed and a restart read everything.
  • Forced panic inside a transaction: reported, store failed for every handle, providers shut down, no checkpoint.
  • Migration memory: one 200 MB transaction gave a 413 MB WAL and 549 MB RSS, hence the 8 MiB chunks (36 MB RSS).
  • Dependency: turso = "=0.8.1" without default features (no mimalloc global allocator, no FTS). Clean release build +126 s (+33 %), desktop binary +12.5 MB, 66 new packages in Cargo.lock. Windows targets need rc.exe at build time (present on the CI runner; a local cross-check from macOS needs a stub).

Tests

Added at the store API: a migration fixture with legacy bare events, CRLF, blank, unparseable and non-UTF-8 lines, an unterminated last line, an empty log, an orphan log and removed meta fields; legacy bare-array index; unparseable index; duplicate ids; interrupted migration (torn staging, complete staging before the rename); interrupted archival; fresh install; refused user_version 0/2 and garbage files; exact bytes through clone and replace; a real second process refused and a relaunch waiting; two simultaneous first launches. Runtime: a host on legacy files migrates once and serves the same timelines across a restart; failed writes are reported and keep the unsaved history (fails on the old evict-on-any-barrier code). Two ignored SIGKILL harnesses stay as evidence tools.

Changed: tests that read or wrote the files now use the store (fork compares read_event_log; the search test replaces content through the store; the icon test reopens the store per restart as a process would). Removed: the icon test's failure injection via a directory at sessions.json.tmp (no database equivalent; the disk-full run covers a failed commit); two store tests folded into the migration fixture.

Checks run

cargo fmt --all --check, cargo clippy --workspace --all-targets --locked -- -D warnings, cargo nextest run --workspace --locked (899 passed, 10 skipped), cargo machete, Web (wasm32) and iOS simulator checks; cargo xwin clippy for x86_64-pc-windows-msvc and cargo zigbuild for x86_64-unknown-linux-gnu on the services and runtime crates. Not run: Windows and Linux tests, Android. An intermittent nextest LEAK on unrelated tcode-services unit tests reproduces on main.

  • This change alters the wire protocol, so the next release needs a
    PROTOCOL_VERSION bump: a note was added under "Unreleased" above the
    constant in crates/protocol/src/lib.rs (the number itself changes only
    when the release is cut — CONTRIBUTING.md, principle 9).

Migration dialog (second commit)

The migration is a separate cancellable step (SessionStore::needs_migration / migrate(progress, cancel)) run under the ownership lock before the host starts; open() refuses an unmigrated directory. The desktop opens a window with a modal dialog first (gpui-base dialog and progress primitives as they are): title, one line of explanation, progress bar, phase and counts, and a Quit button; nothing else is reachable. Quit, ⌘Q or closing the window sets cancel, the migration stops at its next chunk or verification page, discards the staging files and the process exits 0; on success the kernel and shell start in the same process. A failure stays in the dialog and Quit exits 1. Headless prints progress at most once a second and treats SIGINT and SIGTERM as cancel (SIGTERM now gets the same graceful shutdown as SIGINT for the whole of serve).

No source is written, renamed or removed before the staged database is renamed to tcode.db; cancel is never checked after the last pre-publication point, so publish → archive always completes. One fix on the way: the build used to rename an unparseable sessions.json to .corrupt-*; it is now left alone and archived unchanged.

Looked at in the app on scratch data: light and dark, default and 360 pt width; Quit, ⌘Q and the close button mid-import and mid-verify left every source byte-identical and no tcode.db; the next start completed; a forced failure exited 1. Known: at 360 pt the zh-CN description wraps CJK punctuation onto its own line (gpui's line breaker); the failure detail is the English error text.

Tests: a cancelled migration (at four points) changes no source and the next one completes with the phases in order and done == total; headless SIGINT during migration exits 0 with the directory unchanged (real child process); existing migration tests now migrate before opening, as the app does. Workspace suite 901 passed.

sessions.json and every <id>.jsonl move into tcode.db, owned by
SessionStore. Metadata rows hold each project's and session's JSON as
before; event rows hold each raw log line, delimiter included, so export
is byte-identical. One writer connection commits the store writer's
queued writes as one transaction, with synchronous=FULL and, on Apple,
fullfsync, so an acknowledged write survives a crash; a failed commit is
reported for every write in it and a later flush cannot certify it. An
imported or forked thread commits its events and metadata together.

The first start builds tcode.db.migrating from the index and every log in
the data directory, verifies each thread byte for byte and the metadata
field for field, marks the database complete, checkpoints, syncs and
renames it, then moves the sources into legacy/. An interrupted
migration starts over; a complete database is never deleted or
overwritten, and a corrupt one is reported with the sqlite3 CLI as the
salvage path. A lock file owned for the host's lifetime keeps a second
host out of the data directory; a relaunch waits for the outgoing one.

Closes #520
Closes #523
…quit

The migration is its own cancellable step with progress, run under the
data directory's ownership lock before the host starts. The desktop opens
a modal dialog with a progress bar and a Quit button first and starts the
kernel and shell in the same process once it completes; Quit, ⌘Q or
closing the window stops the migration at its next chunk, discards the
staging files and exits. Headless prints progress and treats SIGINT and
SIGTERM as cancel. No source file is written, renamed or removed before
the staged database is published; an unparseable sessions.json is no
longer renamed during the build but archived unchanged afterwards.
@Tryanks
Tryanks enabled auto-merge October 6, 2026 06:47
The export test snapshotted every file in the data dir to prove an export
writes nothing; the running host now holds tcode.db and tcode.lock with
locks that Windows enforces, so the read failed there. The snapshot
skips those files and the test compares the event log and the index
through the store instead.
…tore files by name

The store's lock file cannot be read while it is held on Windows, so the
migration snapshot records store files by name instead of bytes. An
in-process restart right after a host stops can find the data directory
still locked: a spawn the stopped host started (a git status refresh)
keeps a copy of the lock descriptor until the child execs. The app
restarts as a new process, and the test now does the same.
@Tryanks

Tryanks commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Two more CI failures after the previous push, both settled at step 1 (driver, not code):

  • Windows: a_cancelled_migration_changes_no_source_and_the_next_one_completes read tcode.lock while the store held it; Windows's LockFileEx covers the whole file, so even the empty lock file cannot be read. The snapshot records the store's own files by name only; every assertion keeps its meaning (a leftover staging file still fails by name).
  • macOS runner: legacy_files_migrated_at_startup_serve_the_same_timelines restarted a host in-process right after shutdown and got ResourceBusy. The host had released the lock before stopped (verified with temporary logging); the holder was a child process the stopped host was mid-spawn on (a git status refresh): on macOS posix_spawn leaves the child a copy of the parent's descriptors until it execs, which under load took ~55 ms. Reproduced 53/100 under CPU load, 100/100 after the fix. The app only ever restarts as a new process, so the test now starts each run as its own process (the existing child-process pattern). No production change; the relaunch path already waits up to 15 s.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

External import races the store writer on sessions.json

1 participant