Skip to content

[Bug]: Checkpoint refs are written non-atomically — an unclean restart leaves a 0-byte ref that breaks all git fetch/push in the repo #10905

Description

@phillipphoenix

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

T3 writes checkpoint refs to refs/t3/checkpoints/<base64-thread-id>/turn/<n> in the project repo. These appear to be written in place, non-atomically. If the host dies mid-write, the ref file survives the crash with its size committed but its contents never flushed — i.e. as 0 bytes. Git reads a zero-length loose ref as the all-zeros SHA, and from that moment every git command that enumerates local refs fails, which includes fetch, push, and gc.

What happened here:

  1. A thread was running and T3 wrote a checkpoint ref for the turn, plus a loose object.

  2. ~2 minutes later the host (a Hetzner box running the T3 server) restarted uncleanly.

  3. On remount, both files were 0 bytes — the ext4 delayed-allocation crash signature. Both shared the pre-crash mtime and a size of 0:

    0 bytes  .git/refs/t3/checkpoints/<base64-thread-id>/turn/1
    0 bytes  .git/objects/<xx>/<rest-of-sha>
    
  4. Every subsequent worktree creation failed, and the background status poller failed on repeat every ~13s.

To reproduce deliberately, truncate any checkpoint ref in a project repo and then create a worktree:

: > .git/refs/t3/checkpoints/<any-existing-thread>/turn/0
git fetch origin        # fatal, before any network access

Expected behavior

A crash mid-checkpoint should not be able to leave the repository in a state where all git remote operations fail. Checkpoint refs should be written atomically (temp file → fsync → rename), which is what git itself does for exactly this reason. At minimum, a zero-length or non-SHA file under refs/t3/** is never legitimate and could be dropped on startup.

Actual behavior

All fetch/push/gc in the affected repo fail until the file is deleted by hand. Worktree creation surfaces:

GitCommandError: Git command failed in GitVcsDriver.fetchRemote (<repo>): git fetch origin failed

This reads as a credentials or connectivity problem and sends you down the wrong path. It is neither — git fetch dies while building its "have" list during negotiation, before opening a socket. The discriminator is git ls-remote, which never walks local refs:

Check Result
git ls-remote origin HEAD (same URL, creds, network) succeeds
git rev-list --all (no network at all) fails, same fatal
DNS + TCP 443 to the forge OK
Credential helper (gh auth git-credential) OK
Server process env (HOME, PATH) vs. login shell identical
git verify-pack clean

Diagnosing this took a long detour through git fsck, since #4380 means git's actual stderr (fatal: bad ref …) never reaches the user, the logs, or the trace spans. The spans are worse than silent here: GitVcsDriver.fetchRemoteForStatus records exit: {"_tag": "Success"} in server.trace.ndjson even on the polls that failed, so tracing actively points away from the fault.

Related but distinct: #4750 is the same failure class (non-atomic write + crash → NUL/0-byte file bricks a subsystem) for connection-catalog.json; this one is the git ref store. #961 tracks persisted-state corruption generally, but its list of persisted surfaces does not include the repo's refs/t3/**. #5489 is a different checkpoint failure (index.lock contention, intermittent).

Impact

Blocks work completely

Version or commit

0.0.40

Environment

Linux server (ext4, default rw,relatime / data=ordered), git 2.53.0, desktop client on macOS connecting over Tailscale.

Logs or stack traces

Background poller, repeating every ~13s:

GitCommandError: Git command failed in GitVcsDriver.fetchRemoteForStatus (<repo>):
Background Git fetch exited with a non-zero status.
    at statusDetailsRemote (…/t3/dist/bin.mjs)
    at readRemoteStatus …
    at remoteStatus …
    at VcsStatusBroadcaster.refreshRemoteStatus …
    at VcsStatusBroadcaster.retainRemotePoller …

What git actually says, once you run it by hand:

$ git show-ref
fatal: git show-ref: bad ref refs/t3/checkpoints/<…>/turn/1 (0000000000000000000000000000000000000000)

$ git rev-list --all
fatal: bad object refs/t3/checkpoints/<…>/turn/1

$ git fsck --connectivity-only
error: refs/t3/checkpoints/<…>/turn/1: badRefContent:
error: refs/t3/checkpoints/<…>/turn/1: invalid sha1 pointer 0000000000000000000000000000000000000000
error: object file .git/objects/<xx>/<rest-of-sha> is empty

Workaround

Delete the empty ref and object, then re-fetch. Note git update-ref -d will not remove a zero-SHA ref — the file has to be removed directly:

find .git/refs/t3 -type f -size 0 -delete
find .git/objects -type f -size 0 -delete
git fsck --connectivity-only    # expect clean; dangling objects are fine
git fetch --no-tags origin

Any checkpoint whose ref was zeroed is unrecoverable, but it was never a valid commit, so nothing real is lost — that turn just can't be rewound to.

Suggested fixes

  1. Write checkpoint refs atomically (temp → fsync → rename). Same for any loose objects T3 writes directly. This is the actual fix.
  2. Validate refs/t3/** on startup and drop zero-length / non-SHA ref files.
  3. Back off retainRemotePoller after repeated identical failures instead of retrying every ~13s forever.
  4. Git command errors discard stderr, making failures opaque to callers #4380 (carry git's stderr on GitCommandError) would have turned this into a 30-second diagnosis.

Activity

  1. juliusmarminge commented on Sep 9, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed as a real server/VCS bug on current main (0.0.40). Not a duplicate. #9809 is adjacent, not a closer.

    What we verified

    T3 does not write checkpoint ref files itself. GitVcsDriver.checkpoints.captureCheckpoint stages through an isolated temp index, then write-tree / commit-tree into the shared object store, then:

    git update-ref refs/t3/checkpoints/<base64-thread-id>/turn/<n> <oid>
    

    (apps/server/src/vcs/GitVcsDriver.ts, apps/server/src/checkpointing/Utils.ts)

    git update-ref is lockfile → write → rename. That is atomic for concurrent readers, not crash-durable. Git does not fsync loose refs (or loose objects) by default (core.fsync is committed,-loose-object on most platforms). Git’s own notes on core.fsync=reference describe this exact ext4 “empty files” failure: rename is journaled, data never hits disk, the dest is 0 bytes after remount. T3 never passes -c core.fsync=…. The paired 0-byte loose object is the same window on commit-tree / write-tree.

    So the crash signature is right; “T3 writes the ref in place” is slightly off. The durable fix is still fsync-before-publish (Git core.fsync=objects,reference on capture, or the temp → fsync → rename pattern we already use in bootService.writeDurably). There is no startup scrub of refs/t3/**.

    A zero-length loose ref is read as 0000…0. Every local-ref walk then fails (fetch, push, gc, show-ref, rev-list, worktree add). ls-remote is the right discriminator. git update-ref -d will not remove a zero-SHA ref — the file has to be deleted.

    Why this looks like auth / network

    GitCommandError stores stderrLength only and substitutes a static detail (git fetch origin failed, Background Git fetch exited with a non-zero status.). That is #4380.

    executeGit always calls inner execute with allowNonZeroExit: true; the span is on execute. Traces can show exit: {_tag: "Success"} on polls that then fail. That matches the report.

    statusDetailsRemote logs the GitCommandError and then ignoreCauses it, so VcsStatusBroadcaster.retainRemotePoller can treat the refresh as success. Backoff already exists (15s success / 30s→15min failure on the fetch cache; 30s→15min on the poller), but swallowing the error defeats the outer counter. ~13s is consistent with the 15s STATUS_UPSTREAM_REFRESH_INTERVAL, not a missing backoff function.

    Related, not the same

    Issue / PR Why it doesn’t close this
    #9809 (open; #9808 / #3646) Isolates objects and publishes before update-ref in an uninterruptible Effect region. That is interrupt/tmp-pack cleanup, not OS-crash durability. Publish is rename with no fsync; still update-ref without core.fsync.
    #4380 Opaque git errors. Makes this much harder to diagnose; does not cause the 0-byte ref.
    #4750 Same class (crash mid-write bricks a subsystem) for connection-catalog.json.
    #961 Persisted-state corruption; list does not include refs/t3/**.
    #5489 Intermittent restore / index.lock, not crash → empty refs.

    No merged fix for this crash class.

    Suggested fix (smallest first)

    1. Prevent: on checkpoint capture (and any T3 update-ref under refs/t3*), pass -c core.fsync=objects,reference.
    2. Recover: on project open / server start, drop zero-length and non-SHA files under refs/t3/**.
    3. Diagnose: Git command errors discard stderr, making failures opaque to callers #4380 — bounded stderr on GitCommandError; don’t mark the git span Success when the wrapper fails.
    4. Poller: treat a persistent bad ref as a local fault (scrub or stop fetching), don’t swallow it into the success interval.

    The issue’s truncate-a-ref repro is enough to test (2) and (3) without a real crash.

    Workaround (unchanged)

    find .git/refs/t3 -type f -size 0 -delete
    find .git/objects -type f -size 0 -delete
    git fsck --connectivity-only

    That turn’s checkpoint is gone; it was never a valid commit.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Sep 9, 2026
  3. juliusmarminge commented on Oct 2, 2026

    @juliusmarminge
    Member

    Thanks for taking the time to report this and provide the details. We revisited it during the orchestrator V2 cleanup.

    The corrected diagnosis was missing crash durability, not direct in-place ref writes. Checkpoint capture now applies core.fsync=objects,reference and core.fsyncMethod=fsync to write-tree, commit-tree and update-ref, directly preventing the reported publication window.

    I’m closing this based on the current source and the evidence in this thread.

    Source reviewed.

    Source verification only. Existing corrupt repositories are not automatically repaired; the ancillary diagnostics/recovery requests are not proven implemented.

    If you still hit this on a current build, please reply with the app/server versions and the steps that reproduce it. We can reopen this if the original problem is still there.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions