Repository navigation
[Bug]: Checkpoint refs are written non-atomically — an unclean restart leaves a 0-byte ref that breaks all git fetch/push in the repo #10905
Description
Activity
Triage
Confirmed as a real server/VCS bug on current
main(0.0.40). Not a duplicate. #9809 is adjacent, not a closer.What we verified
T3 does not write checkpoint ref files itself.
GitVcsDriver.checkpoints.captureCheckpointstages through an isolated temp index, thenwrite-tree/commit-treeinto the shared object store, then:git update-ref refs/t3/checkpoints/<base64-thread-id>/turn/<n> <oid>(
apps/server/src/vcs/GitVcsDriver.ts,apps/server/src/checkpointing/Utils.ts)git update-refis lockfile → write → rename. That is atomic for concurrent readers, not crash-durable. Git does not fsync loose refs (or loose objects) by default (core.fsynciscommitted,-loose-objecton most platforms). Git’s own notes oncore.fsync=referencedescribe this exact ext4 “empty files” failure: rename is journaled, data never hits disk, the dest is 0 bytes after remount. T3 never passes-c core.fsync=…. The paired 0-byte loose object is the same window oncommit-tree/write-tree.So the crash signature is right; “T3 writes the ref in place” is slightly off. The durable fix is still fsync-before-publish (Git
core.fsync=objects,referenceon capture, or the temp → fsync → rename pattern we already use inbootService.writeDurably). There is no startup scrub ofrefs/t3/**.A zero-length loose ref is read as
0000…0. Every local-ref walk then fails (fetch,push,gc,show-ref,rev-list,worktree add).ls-remoteis the right discriminator.git update-ref -dwill not remove a zero-SHA ref — the file has to be deleted.Why this looks like auth / network
GitCommandErrorstoresstderrLengthonly and substitutes a staticdetail(git fetch origin failed,Background Git fetch exited with a non-zero status.). That is #4380.executeGitalways calls innerexecutewithallowNonZeroExit: true; the span is onexecute. Traces can showexit: {_tag: "Success"}on polls that then fail. That matches the report.statusDetailsRemotelogs theGitCommandErrorand thenignoreCauses it, soVcsStatusBroadcaster.retainRemotePollercan treat the refresh as success. Backoff already exists (15s success / 30s→15min failure on the fetch cache; 30s→15min on the poller), but swallowing the error defeats the outer counter. ~13s is consistent with the 15sSTATUS_UPSTREAM_REFRESH_INTERVAL, not a missing backoff function.Related, not the same
Issue / PR Why it doesn’t close this #9809 (open; #9808 / #3646) Isolates objects and publishes before update-refin an uninterruptible Effect region. That is interrupt/tmp-pack cleanup, not OS-crash durability. Publish isrenamewith no fsync; stillupdate-refwithoutcore.fsync.#4380 Opaque git errors. Makes this much harder to diagnose; does not cause the 0-byte ref. #4750 Same class (crash mid-write bricks a subsystem) for connection-catalog.json.#961 Persisted-state corruption; list does not include refs/t3/**.#5489 Intermittent restore / index.lock, not crash → empty refs.No merged fix for this crash class.
Suggested fix (smallest first)
- Prevent: on checkpoint capture (and any T3
update-refunderrefs/t3*), pass-c core.fsync=objects,reference. - Recover: on project open / server start, drop zero-length and non-SHA files under
refs/t3/**. - Diagnose: Git command errors discard stderr, making failures opaque to callers #4380 — bounded stderr on
GitCommandError; don’t mark the git span Success when the wrapper fails. - Poller: treat a persistent
bad refas a local fault (scrub or stop fetching), don’t swallow it into the success interval.
The issue’s truncate-a-ref repro is enough to test (2) and (3) without a real crash.
Workaround (unchanged)
find .git/refs/t3 -type f -size 0 -delete find .git/objects -type f -size 0 -delete git fsck --connectivity-only
That turn’s checkpoint is gone; it was never a valid commit.
- Prevent: on checkpoint capture (and any T3
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 9, 2026 Thanks for taking the time to report this and provide the details. We revisited it during the orchestrator V2 cleanup.
The corrected diagnosis was missing crash durability, not direct in-place ref writes. Checkpoint capture now applies core.fsync=objects,reference and core.fsyncMethod=fsync to write-tree, commit-tree and update-ref, directly preventing the reported publication window.
I’m closing this based on the current source and the evidence in this thread.
Source verification only. Existing corrupt repositories are not automatically repaired; the ancillary diagnostics/recovery requests are not proven implemented.
If you still hit this on a current build, please reply with the app/server versions and the steps that reproduce it. We can reopen this if the original problem is still there.
Before submitting
Area
apps/server
Steps to reproduce
T3 writes checkpoint refs to
refs/t3/checkpoints/<base64-thread-id>/turn/<n>in the project repo. These appear to be written in place, non-atomically. If the host dies mid-write, the ref file survives the crash with its size committed but its contents never flushed — i.e. as 0 bytes. Git reads a zero-length loose ref as the all-zeros SHA, and from that moment every git command that enumerates local refs fails, which includes fetch, push, and gc.What happened here:
A thread was running and T3 wrote a checkpoint ref for the turn, plus a loose object.
~2 minutes later the host (a Hetzner box running the T3 server) restarted uncleanly.
On remount, both files were 0 bytes — the ext4 delayed-allocation crash signature. Both shared the pre-crash mtime and a size of 0:
Every subsequent worktree creation failed, and the background status poller failed on repeat every ~13s.
To reproduce deliberately, truncate any checkpoint ref in a project repo and then create a worktree:
Expected behavior
A crash mid-checkpoint should not be able to leave the repository in a state where all git remote operations fail. Checkpoint refs should be written atomically (temp file →
fsync→rename), which is what git itself does for exactly this reason. At minimum, a zero-length or non-SHA file underrefs/t3/**is never legitimate and could be dropped on startup.Actual behavior
All fetch/push/gc in the affected repo fail until the file is deleted by hand. Worktree creation surfaces:
This reads as a credentials or connectivity problem and sends you down the wrong path. It is neither —
git fetchdies while building its "have" list during negotiation, before opening a socket. The discriminator isgit ls-remote, which never walks local refs:git ls-remote origin HEAD(same URL, creds, network)git rev-list --all(no network at all)gh auth git-credential)HOME,PATH) vs. login shellgit verify-packDiagnosing this took a long detour through
git fsck, since #4380 means git's actual stderr (fatal: bad ref …) never reaches the user, the logs, or the trace spans. The spans are worse than silent here:GitVcsDriver.fetchRemoteForStatusrecordsexit: {"_tag": "Success"}inserver.trace.ndjsoneven on the polls that failed, so tracing actively points away from the fault.Related but distinct: #4750 is the same failure class (non-atomic write + crash → NUL/0-byte file bricks a subsystem) for
connection-catalog.json; this one is the git ref store. #961 tracks persisted-state corruption generally, but its list of persisted surfaces does not include the repo'srefs/t3/**. #5489 is a different checkpoint failure (index.lock contention, intermittent).Impact
Blocks work completely
Version or commit
0.0.40
Environment
Linux server (ext4, default
rw,relatime/data=ordered), git 2.53.0, desktop client on macOS connecting over Tailscale.Logs or stack traces
Background poller, repeating every ~13s:
What git actually says, once you run it by hand:
Workaround
Delete the empty ref and object, then re-fetch. Note
git update-ref -dwill not remove a zero-SHA ref — the file has to be removed directly:find .git/refs/t3 -type f -size 0 -delete find .git/objects -type f -size 0 -delete git fsck --connectivity-only # expect clean; dangling objects are fine git fetch --no-tags originAny checkpoint whose ref was zeroed is unrecoverable, but it was never a valid commit, so nothing real is lost — that turn just can't be rewound to.
Suggested fixes
fsync→rename). Same for any loose objects T3 writes directly. This is the actual fix.refs/t3/**on startup and drop zero-length / non-SHA ref files.retainRemotePollerafter repeated identical failures instead of retrying every ~13s forever.GitCommandError) would have turned this into a 30-second diagnosis.