Skip to content

connection-catalog.json is written non-atomically every ~3s, and a corrupt document bricks the app permanently #4750

Description

@Samy104

Summary

~/.t3/userdata/connection-catalog.json appears to be written in place, non-atomically, on a fixed ~3 second cadence. If the process dies mid-write (power loss, hard reset, OOM kill), NTFS commits the file's new size but not its data, leaving the file filled entirely with NUL bytes.

On the next launch, desktop.connectionCatalogStore.get throws a decode error and there is no fallback path — the app comes up with an empty project list and a connection that spins on "Connecting…" forever. It never self-heals, because the renderer only calls connectionCatalog.set after a successful get, so no subsequent write ever repairs the file.

The result is a completely unusable app from a single unlucky 3-second window, with no in-app indication of what's wrong.

Impact

After an unclean shutdown, the app launches into this state permanently:

  • Sidebar shows "No projects yet" despite projects existing in state.sqlite
  • The connection (in my case WSL/Ubuntu) is stuck on "Connecting…" indefinitely
  • Agents (Claude, Codex) can't be used, since no connection ever resolves
  • Nothing in the UI surfaces the underlying error — it just looks like data loss

Critically, no data is actually lost. state.sqlite is completely intact. Only a small piece of connection metadata is destroyed, but it takes the entire app down with it.

Reproduction

  1. Have T3 Code open with at least one project and one connection
  2. Kill power to the machine (or hard-kill the process) — the window is ~3s wide, so it lands fairly often
  3. On reboot, check ~/.t3/userdata/connection-catalog.json
  4. Relaunch T3 Code

Expected: app recovers, falls back to a default catalog, rediscovers local/WSL connections Actual: empty project list, connection hangs forever, no error shown

Evidence

The file is 100% NUL bytes

$ xxd ~/.t3/userdata/connection-catalog.json | head -3
00000000: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00000010: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00000020: 0000 0000 0000 0000 0000 0000 0000 0000  ................

size=1408 nul_bytes=1408

Every other JSON file under .t3/ parsed cleanly. This was the only corrupt file on disk.

Timeline from desktop.trace.ndjson

Time (UTC) | Event -- | -- 20:27:18 | last successful desktop.connectionCatalogStore.writeDocument 20:27:22 | file mtime — the write that got truncated by power loss 20:35:35 | first DesktopConnectionCatalogStoreDocumentDecodeError on relaunch every launch since | same error, 16 occurrences per session, zero recovery attempts

The error

DesktopConnectionCatalogStoreDocumentDecodeError: Failed to decode the desktop
connection catalog document at C:\Users\<user>\.t3\userdata\connection-catalog.json.
    at ...\resources\app.asar\apps\desktop\dist-electron\main.cjs:12811:227
    at desktop.connectionCatalogStore.get (...main.cjs:12...)

Caused by: SyntaxError: Unexpected token '\u0000'

It propagates all the way up the IPC chain — every one of these fails together:

desktop.connectionCatalogStore.get   Failure
desktop.ipc.connectionCatalog.get    Failure
desktop.ipc.method                   Failure
desktop.ipc.invoke                   Failure

The write cadence is a constant 3.01s

Gap histogram between consecutive writeDocument spans in a single session:

3.01s → 124    3.02s → 47    3.00s → 8    3.06s → 2    3.05s → 1

That's ~1,200 writes/hour of what is essentially static connection metadata. The spans carry no payload, so I can't confirm the content is unchanged between writes, but a perfectly constant 3.01s interval strongly suggests a polling timer that writes unconditionally rather than on change.

An atomic writer already exists elsewhere in the codebase

I found this sitting next to the corrupt file:

client-settings.json                                       803 bytes  (valid)
client-settings.json.59004.e6b320f650d14d91a69bd7885ebb0da5.tmp   663 bytes  (valid)

That <name>.<pid>.<hash>.tmp pattern is a write-to-temp-then-rename. The settings store was also interrupted by the same power loss — and its real file survived perfectly intact, leaving only an orphaned temp file behind. The connection catalog store, with no such sibling, was destroyed in place.

So the fix pattern is already present in the codebase; the catalog store just isn't using it.

Suggested fixes

In rough priority order:

  1. Write atomically. Reuse whatever the settings store does: write to connection-catalog.json.<pid>.<rand>.tmp, fsync, then rename() over the target. On POSIX and NTFS alike, rename is atomic — a torn write can then only ever leave an orphaned temp file, never a destroyed catalog.
  2. Never let a decode failure be fatal. Treat a malformed document exactly like a missing one: log a warning, move the bad file aside to connection-catalog.json.corrupt-<timestamp>, and fall back to the default catalog so local/WSL connections get rediscovered. A config file that can't be parsed should degrade, not brick.
  3. Only write when the content actually changed. Compare the serialized document to the last-written value and skip the write if identical. This collapses ~1,200 writes/hour to near zero, which shrinks the corruption window by orders of magnitude on its own — and is worth doing regardless of (1) for disk-wear and battery reasons.
  4. Surface the failure in the UI. Right now "No projects yet" + an infinite connection spinner is indistinguishable from real data loss. Even a toast saying "couldn't read connection settings" would have saved a lot of guesswork.

(1) and (2) are independent and both worth having: (1) prevents the corruption, (2) makes the app survive it if it happens anyway through some other route.

Workaround for anyone hitting this now

Quit T3 Code completely (including any lingering backend child process), then:

powershell
mv "$HOME\.t3\userdata\connection-catalog.json" "$HOME\.t3\userdata\connection-catalog.json.bak"

Relaunch. The catalog is regenerated, projects reappear from state.sqlite, and local/WSL connections are rediscovered. Nothing is lost — the corrupt file contained no recoverable bytes anyway. Any manually configured remote connections may need to be re-added.

Environment

  • OS: Windows 11, x64 (NTFS)
  • T3 Code: Alpha channel, t3@0.0.29
  • Backend: WSL backend enabled, distro Ubuntu (wslBackendEnabled: true)
  • Trigger: power surge / unclean shutdown

Possibly unrelated, but noticed while reading the logs

desktop.trace.ndjson records SSH remote pairing tokens in plaintext inside span events:

Connection string: http://127.0.0.1:3774
Token: <REDACTED>
Pairing URL: http://127.0.0.1:3774/pair#token=<REDACTED>

Since these trace files are exactly what users get asked to attach to bug reports, it may be worth redacting tokens at the log sink. Happy to file this separately if you'd prefer.

Activity

  1. TrueBurn commented on Aug 25, 2026

    @TrueBurn

    Hit this on 0.0.33, four weeks after you filed it. Windows 11, WSL2 Ubuntu, same three symptoms: no providers, no projects, no sessions.

    Some numbers that might help, plus a confirmation that the recovery in #5902 works.

    The file

    Mine was 860 bytes of NUL. Yours was 1408. After recovery the app wrote a fresh one at 860 bytes, so the size follows the catalog contents, and in both cases the size reached disk and the data didn't.

    The envelope:

    {"version":1,"encryptedCatalog":"<base64>"}

    So the outer document is what fails to decode, matching the description in #5902. Every other JSON file in .t3\userdata parsed fine.

    On the ~3s cadence

    I sampled the mtime with the app idle, nothing running:

    16:34:14  16:34:21  16:34:29  16:34:36  16:34:45
    16:34:53  16:35:01  16:35:10  16:35:19  16:35:26
    

    It rewrites every 6 to 9 seconds, the same 860 bytes, and it doesn't stop. No state change behind any of it.

    The part that surprised me: my corrupt file carried an mtime of 15:37, but I rebooted at 16:05. At that cadence the app wrote it another 240 times in between, and Windows flushed none of them, data or metadata. So the exposure isn't a few seconds around one write. I lost 28 minutes of writes to this path, which makes the odds of landing on it worse than the title suggests.

    Nothing else broke

    You're right that the data survives:

    • Windows state.sqlite, 59 MB: pragma quick_check returns ok
    • WSL state.sqlite, 344 MB: same

    For anyone stuck here before #5902 lands:

    1. Quit T3 Code
    2. mv ~/.t3/userdata/connection-catalog.json ~/.t3/userdata/connection-catalog.json.bak
    3. Start it again

    I had a valid catalog back within seconds, my projects and threads where I left them, and both environments reconnected on their own. From desktop.trace.ndjson:

    catalogStore.get Failure  last at 16:09:53
    catalogStore.get Success  last at 16:26:26
    

    I never re-added the WSL connection. desktop-settings.json still held wslBackendEnabled: true and wslDistro: "Ubuntu", so the app picked it up again.

    That's the same place #5902 gets to on its own by quarantining the document and treating the catalog as missing, so the Option.none path gives you everything back instead of a half-working app. The draft looks like it's sitting on the Effect.catchTag to Effect.catchTags convention check.

    The logs point the wrong way

    Worth a warning for whoever debugs this next. My child log was full of:

    Timed out after 60000ms waiting for desktop backend readiness at http://<wsl-ip>:3774/.well-known/t3/environment.
    

    WSL hadn't finished booting inside the 60 second probe, and I chased that first. Meanwhile the UI sat empty and both backends answered:

    GET http://127.0.0.1:3773/.well-known/t3/environment  -> 200
    GET http://<wsl-ip>:3774/.well-known/t3/environment   -> 200
    

    The catalog gate runs ahead of connection health, so a healthy backend still leaves you with an empty app.

    One more thing I noticed while testing: the probe doesn't retry after it gives up. An environment that comes up after those 60 seconds stays unavailable until you restart the app, valid catalog or not.

  2. Pjieter commented on Sep 14, 2026

    @Pjieter

    Same bug on 0.0.40 (desktop, Windows 11 x64 10.0.26200, WSL Debian backend enabled). Found with t3 triage.

    What happened
    The desktop app opens to "Still connecting / T3 Code could not confirm this workspace." with a Reload button that doesn't help. It has been broken since around Sep 5. Uninstalling and reinstalling didn't fix it, because the uninstaller leaves ~/.t3/userdata in place.

    Evidence

    • ~/.t3/userdata/connection-catalog.json: 860 bytes, all 860 are NUL, mtime 2026-09-05 14:32. That's the same size as the file in the earlier comment here.
    • desktop.trace.ndjson on every launch:
      desktop.connectionCatalogStore.get Failure
      DesktopConnectionCatalogStoreDocumentDecodeError: Failed to decode the desktop connection catalog document at ~/.t3/userdata/connection-catalog.json
      
    • The backends themselves are healthy. probeReadiness succeeds for both the local server (127.0.0.1:3773) and the WSL server (:3774/.well-known/t3/environment returns 200).
    • Every other JSON file in userdata is fine.

    How the corrupt file turns into this screen (traced in v0.0.40 source)

    1. DesktopConnectionCatalogStore.readDocument (apps/desktop/src/app/DesktopConnectionCatalogStore.ts:204-215) fails on the outer document decode. Only NotFound is handled.
    2. The IPC promise rejects, apps/web/src/connection/storage.ts:264-269 maps it to catalogError("load"), and loadUnlocked (:317) doesn't catch it. The web-side quarantine/reset path (:320-346) only covers bad strings that reach the renderer, so it never runs on desktop.
    3. EnvironmentRegistry.make (packages/client-runtime/src/connection/registry.ts:141) fails while the layer builds, and catalogValueAtom falls back to isReady: false for good.
    4. resolveFirstRunDecision stays "pending" while !catalogReady (apps/web/src/onboarding/firstRun.logic.ts:149-156). After the timeout, FirstRunGate shows the "Still connecting" recovery screen, and its Reload button just runs the same failing load again.

    So on desktop, the "Still connecting" screen is another symptom of this bug, on top of the empty sidebar and endless "Connecting" described above. Anyone searching that error message should end up here.

    Workaround
    Quit T3 Code, rename connection-catalog.json to .bak, and relaunch. Confirmed working: the app started normally again.

    Note that the writer on v0.0.40 already writes to a temp file and renames it, but it never fsyncs before the rename. On NTFS that still leaves a NUL-filled file after an unclean shutdown. #5902 (quarantine a corrupt document and treat it as missing) was closed without merging on Sep 7, and I found no fix on main.

    Posted via T3 Code triage by Claude Opus 5 (Claude Code).

  3. Pjieter commented on Sep 16, 2026

    @Pjieter

    Second time on this machine, now on 0.0.42. There's also one detail I missed before that I think explains why this bug keeps getting blamed on releases.

    Same setup as my 14/09 comment: Windows 11 x64 (10.0.26200), desktop app, WSL Debian backend. Found again with t3 triage.

    Both corruptions landed at a power cut

    Local time Event
    05/09 14:32:37 connection-catalog.json mtime, 860 bytes, all NUL
    05/09 14:33:00 Kernel-Power 41
    14/09 21:21:32 last of 1062 connectionCatalogStore.set spans that session
    14/09 21:21:36 connection-catalog.json mtime, 2168 bytes, all NUL
    14/09 21:21:45 Kernel-Power 41

    The 05/09 file is the one from my earlier comment. I renamed it aside, the app rebuilt a working catalog, and nine days later the next power cut destroyed the replacement. Two unclean shutdowns, two dead installs. Neither time did anything else in userdata fail to parse.

    My power situation is my own problem. But a 4-second write loop is what turns a 20-second outage into a permanently broken app, and that part isn't.

    The cadence on 0.0.40 measures 4s, unconditional

    1062 desktop.connectionCatalogStore.writeDocument spans in one 70-minute session. Gap histogram: 4s x1047, 3s x14. Nothing in the app state changed across those 70 minutes. About 900 writes an hour of a file whose contents never moved, which matches the 3s and 6-9s numbers others reported here.

    The damage stays invisible until the next restart

    This is the bit I hadn't worked out before. makeCatalogStore caches the decoded catalog in a Ref (apps/web/src/connection/storage.ts:322-359). After one successful load, loadUnlocked hands back the cached value and never reads disk again. So a running app carries on quite happily after its catalog file has been shredded, for as long as nobody restarts it.

    Mine ran for two days on a catalog that existed only in memory. What finally broke it was an update I triggered remotely from another machine:

    15:06:14Z  last desktop.updates.handleUpdateAvailable
    15:06:48Z  desktop.backendInstance.stop
    15:07:05Z  desktop.backendInstance.start
    15:07:09Z  desktop.ipc.connectionCatalog.get  fails with DocumentDecodeError
    

    connectionCatalog.get fired 5 times across 48 hours of trace, once per renderer start, and all 5 failed. I was convinced the 0.0.42 update had eaten my projects, so I rolled back to 0.0.40, which of course fixed nothing, because the bad file sits in userdata and no installer goes near it. I'd bet a fair share of the "the update deleted everything" reports are this: an old crash doing the damage, and whatever restarts the app getting the blame.

    Still unfixed in 0.0.42

    PR #5902 touched this exact file and closed unmerged on 07/09. In v0.0.42 (719a76ca1dbf), readDocument handles only NotFound (apps/desktop/src/app/DesktopConnectionCatalogStore.ts:186-215), so a corrupt document is still fatal.

    The recovery the web layer already has can't reach the desktop either. loadUnlocked does discard and rewrite a corrupt catalog (storage.ts:330-355), but that path needs backend.read to succeed and pass it a bad string. On desktop the decode happens in the main process, the IPC promise rejects, backend.read fails with catalogError("load"), and the recovery branch never runs. The desktop backend supplies no quarantine either (storage.ts:272-296), unlike the IndexedDB one at :304. Browser installs heal themselves from this. Desktop installs sit there bricked.

    Nothing was actually lost, either time

    state.sqlite still held all 5 projects, deleted_at null on every row, 33 threads, thread activity from the same day. A few hundred bytes of connection metadata took the whole app down with it.

    The workaround still works

    Quit the app fully, rename ~/.t3/userdata/connection-catalog.json aside, relaunch. The catalog gets rebuilt and the projects come back. Manually configured remote connections need re-adding.


    Filed from t3 triage on 0.0.42. Diagnosis by Claude Opus 5 (1M context) running in Claude Code.

  4. JDeffner commented on Sep 24, 2026

    @JDeffner

    Another occurrence on 0.0.42, Windows 10 Home x64, build 19045 (NTFS), investigated on 24 September 2026. This adds Windows 10 evidence and an isolated reproduction of the desktop-to-renderer failure path.

    After restarting the development desktop, T3 Code appeared to have no connected environment. The host server remained accessible from another client.

    Evidence from this machine

    • ~/.t3/userdata/connection-catalog.json: 1,552 bytes, all 1,552 NUL. A copy was preserved. The other seven top-level JSON files under 200 KB parsed successfully.
    • The same outer-document decode failure occurred on three desktop launches:
    2026-09-24T13:39:47.513Z  desktop.connectionCatalogStore.get  Failure
    2026-09-24T13:41:08.201Z  desktop.connectionCatalogStore.get  Failure
    2026-09-24T13:46:47.423Z  desktop.connectionCatalogStore.get  Failure
    
    DesktopConnectionCatalogStoreDocumentDecodeError:
    Failed to decode the desktop connection catalog document at
    ~/.t3/userdata/connection-catalog.json.
    Cause: SchemaError: Expected a valid JSON string
    
    • Each launch also recorded backend readiness. GET http://127.0.0.1:3773/.well-known/t3/environment returned HTTP 200, reporting version 0.0.42.
    • Read-only PRAGMA quick_check returned ok. The database still contained 38 nondeleted projects and 536 nondeleted threads, including the work in progress.
    • Windows recorded Kernel-Power event 41 at 15:12:42 CEST, then a user-initiated restart at 15:15:03. The catalog's recorded modification time is 15:12:10. Event 6008 reports an earlier unexpected shutdown at 14:54:09. An unclean shutdown is a plausible trigger; these timestamps do not prove the exact corruption mechanism or time.

    Reproduction without interrupting the live host

    Two isolated tests used the real desktop catalog store, a temporary T3 home, and the repository's test double for OS encryption:

    1. Write a synthetic 1,552-byte NUL catalog. Call get twice. Both calls fail with DesktopConnectionCatalogStoreDocumentDecodeError, leaving the fixture unchanged. Rename only that fixture to a quarantine filename. The next read returns Option.none().
    2. Connect that store to the real renderer catalog backend through a test IPC bridge. Reading fails with ConnectionTransientError, and no recovery write occurs. After quarantining the fixture, the renderer can load an empty catalog.

    Both tests passed, along with nine existing desktop catalog tests and 44 existing web storage/first-run tests. The desktop tests used an isolated configuration to resolve already-installed workspace dependencies. No source changes, live catalog reset, power-loss experiment, or full desktop UI recovery were performed.

    This confirms the distinction described above: the desktop read rejects before the renderer recovery branch can run. Registry construction then fails even though the local server is healthy.

    The installed app.asar already uses temporary-file write followed by rename. Its catalog write path has no explicit durable flush or backup. The current writer should therefore not be described as writing in place, and the missing flush remains a possible contributor rather than a proven cause here.

    Recovery should preserve the damaged file and let a healthy local environment start, with an explicit connection-settings error. Malformed documents need to remain distinct from permission failures and temporarily unavailable encryption keys so recovery does not overwrite recoverable credentials.

    Follow-up: proposed fix and verification

    PR #13468 contains the catalog recovery fix and its regression tests. Its description links back to this investigation. Subsequent checks verified recovery in an isolated built desktop app, preservation of the damaged bytes, native Windows encryption, and a full app restart. The PR records the test coverage and limits for anyone investigating an alternative fix. The original corruption trigger remains unproven; no physical power-loss experiment was performed.

  5. nasroykh commented on Oct 1, 2026

    @nasroykh

    Hit this on current main (148e6deea0, desktop dev build via start:desktop), so it is still unfixed there as of Oct 1.

    • Windows 11 Pro x64, build 26300, NTFS
    • Power cut on Sep 30 while the app was open. connection-catalog.json mtime is 23:23 that night.
    • The file is 904 bytes, all NUL. Every other JSON file in ~/.t3/userdata has no NUL bytes, and state.sqlite has a valid header.
    • On relaunch the backend reports ready on 127.0.0.1:3773, then the renderer's catalog read fails:
    Error occurred in handler for 'desktop:get-connection-catalog': {
      _tag: 'DesktopConnectionCatalogStoreDocumentDecodeError',
      catalogPath: 'C:\\Users\\<user>\\.t3\\userdata\\connection-catalog.json',
      ...
    }
    

    The local environment never connects, and projects and providers are missing from the UI.

    I wrote the quarantine-and-treat-as-missing change locally against main as a check. A regression test with a 904-byte zero-filled file fails without it and passes with it. That is the same approach as #13468, which also flushes the temp file before the rename, so I'd rather back that PR than open a duplicate.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions