Repository navigation
connection-catalog.json is written non-atomically every ~3s, and a corrupt document bricks the app permanently #4750
Description
Activity
Hit this on
0.0.33, four weeks after you filed it. Windows 11, WSL2 Ubuntu, same three symptoms: no providers, no projects, no sessions.Some numbers that might help, plus a confirmation that the recovery in #5902 works.
The file
Mine was 860 bytes of NUL. Yours was 1408. After recovery the app wrote a fresh one at 860 bytes, so the size follows the catalog contents, and in both cases the size reached disk and the data didn't.
The envelope:
{"version":1,"encryptedCatalog":"<base64>"}So the outer document is what fails to decode, matching the description in #5902. Every other JSON file in
.t3\userdataparsed fine.On the ~3s cadence
I sampled the mtime with the app idle, nothing running:
16:34:14 16:34:21 16:34:29 16:34:36 16:34:45 16:34:53 16:35:01 16:35:10 16:35:19 16:35:26It rewrites every 6 to 9 seconds, the same 860 bytes, and it doesn't stop. No state change behind any of it.
The part that surprised me: my corrupt file carried an mtime of 15:37, but I rebooted at 16:05. At that cadence the app wrote it another 240 times in between, and Windows flushed none of them, data or metadata. So the exposure isn't a few seconds around one write. I lost 28 minutes of writes to this path, which makes the odds of landing on it worse than the title suggests.
Nothing else broke
You're right that the data survives:
- Windows
state.sqlite, 59 MB:pragma quick_checkreturnsok - WSL
state.sqlite, 344 MB: same
For anyone stuck here before #5902 lands:
- Quit T3 Code
mv ~/.t3/userdata/connection-catalog.json ~/.t3/userdata/connection-catalog.json.bak- Start it again
I had a valid catalog back within seconds, my projects and threads where I left them, and both environments reconnected on their own. From
desktop.trace.ndjson:catalogStore.get Failure last at 16:09:53 catalogStore.get Success last at 16:26:26I never re-added the WSL connection.
desktop-settings.jsonstill heldwslBackendEnabled: trueandwslDistro: "Ubuntu", so the app picked it up again.That's the same place #5902 gets to on its own by quarantining the document and treating the catalog as missing, so the
Option.nonepath gives you everything back instead of a half-working app. The draft looks like it's sitting on theEffect.catchTagtoEffect.catchTagsconvention check.The logs point the wrong way
Worth a warning for whoever debugs this next. My child log was full of:
Timed out after 60000ms waiting for desktop backend readiness at http://<wsl-ip>:3774/.well-known/t3/environment.WSL hadn't finished booting inside the 60 second probe, and I chased that first. Meanwhile the UI sat empty and both backends answered:
GET http://127.0.0.1:3773/.well-known/t3/environment -> 200 GET http://<wsl-ip>:3774/.well-known/t3/environment -> 200The catalog gate runs ahead of connection health, so a healthy backend still leaves you with an empty app.
One more thing I noticed while testing: the probe doesn't retry after it gives up. An environment that comes up after those 60 seconds stays unavailable until you restart the app, valid catalog or not.
- Windows
Same bug on 0.0.40 (desktop, Windows 11 x64 10.0.26200, WSL Debian backend enabled). Found with
t3 triage.What happened
The desktop app opens to "Still connecting / T3 Code could not confirm this workspace." with a Reload button that doesn't help. It has been broken since around Sep 5. Uninstalling and reinstalling didn't fix it, because the uninstaller leaves~/.t3/userdatain place.Evidence
~/.t3/userdata/connection-catalog.json: 860 bytes, all 860 are NUL, mtime 2026-09-05 14:32. That's the same size as the file in the earlier comment here.desktop.trace.ndjsonon every launch:desktop.connectionCatalogStore.get Failure DesktopConnectionCatalogStoreDocumentDecodeError: Failed to decode the desktop connection catalog document at ~/.t3/userdata/connection-catalog.json- The backends themselves are healthy.
probeReadinesssucceeds for both the local server (127.0.0.1:3773) and the WSL server (:3774/.well-known/t3/environment returns 200). - Every other JSON file in userdata is fine.
How the corrupt file turns into this screen (traced in v0.0.40 source)
DesktopConnectionCatalogStore.readDocument(apps/desktop/src/app/DesktopConnectionCatalogStore.ts:204-215) fails on the outer document decode. Only NotFound is handled.- The IPC promise rejects,
apps/web/src/connection/storage.ts:264-269maps it tocatalogError("load"), andloadUnlocked(:317) doesn't catch it. The web-side quarantine/reset path (:320-346) only covers bad strings that reach the renderer, so it never runs on desktop. EnvironmentRegistry.make(packages/client-runtime/src/connection/registry.ts:141) fails while the layer builds, andcatalogValueAtomfalls back toisReady: falsefor good.resolveFirstRunDecisionstays"pending"while!catalogReady(apps/web/src/onboarding/firstRun.logic.ts:149-156). After the timeout,FirstRunGateshows the "Still connecting" recovery screen, and its Reload button just runs the same failing load again.
So on desktop, the "Still connecting" screen is another symptom of this bug, on top of the empty sidebar and endless "Connecting" described above. Anyone searching that error message should end up here.
Workaround
Quit T3 Code, renameconnection-catalog.jsonto.bak, and relaunch. Confirmed working: the app started normally again.Note that the writer on v0.0.40 already writes to a temp file and renames it, but it never fsyncs before the rename. On NTFS that still leaves a NUL-filled file after an unclean shutdown. #5902 (quarantine a corrupt document and treat it as missing) was closed without merging on Sep 7, and I found no fix on
main.Posted via T3 Code triage by Claude Opus 5 (Claude Code).
Second time on this machine, now on 0.0.42. There's also one detail I missed before that I think explains why this bug keeps getting blamed on releases.
Same setup as my 14/09 comment: Windows 11 x64 (10.0.26200), desktop app, WSL Debian backend. Found again with
t3 triage.Both corruptions landed at a power cut
Local time Event 05/09 14:32:37 connection-catalog.jsonmtime, 860 bytes, all NUL05/09 14:33:00 Kernel-Power 41 14/09 21:21:32 last of 1062 connectionCatalogStore.setspans that session14/09 21:21:36 connection-catalog.jsonmtime, 2168 bytes, all NUL14/09 21:21:45 Kernel-Power 41 The 05/09 file is the one from my earlier comment. I renamed it aside, the app rebuilt a working catalog, and nine days later the next power cut destroyed the replacement. Two unclean shutdowns, two dead installs. Neither time did anything else in
userdatafail to parse.My power situation is my own problem. But a 4-second write loop is what turns a 20-second outage into a permanently broken app, and that part isn't.
The cadence on 0.0.40 measures 4s, unconditional
1062
desktop.connectionCatalogStore.writeDocumentspans in one 70-minute session. Gap histogram: 4s x1047, 3s x14. Nothing in the app state changed across those 70 minutes. About 900 writes an hour of a file whose contents never moved, which matches the 3s and 6-9s numbers others reported here.The damage stays invisible until the next restart
This is the bit I hadn't worked out before.
makeCatalogStorecaches the decoded catalog in aRef(apps/web/src/connection/storage.ts:322-359). After one successful load,loadUnlockedhands back the cached value and never reads disk again. So a running app carries on quite happily after its catalog file has been shredded, for as long as nobody restarts it.Mine ran for two days on a catalog that existed only in memory. What finally broke it was an update I triggered remotely from another machine:
15:06:14Z last desktop.updates.handleUpdateAvailable 15:06:48Z desktop.backendInstance.stop 15:07:05Z desktop.backendInstance.start 15:07:09Z desktop.ipc.connectionCatalog.get fails with DocumentDecodeErrorconnectionCatalog.getfired 5 times across 48 hours of trace, once per renderer start, and all 5 failed. I was convinced the 0.0.42 update had eaten my projects, so I rolled back to 0.0.40, which of course fixed nothing, because the bad file sits inuserdataand no installer goes near it. I'd bet a fair share of the "the update deleted everything" reports are this: an old crash doing the damage, and whatever restarts the app getting the blame.Still unfixed in 0.0.42
PR #5902 touched this exact file and closed unmerged on 07/09. In v0.0.42 (
719a76ca1dbf),readDocumenthandles onlyNotFound(apps/desktop/src/app/DesktopConnectionCatalogStore.ts:186-215), so a corrupt document is still fatal.The recovery the web layer already has can't reach the desktop either.
loadUnlockeddoes discard and rewrite a corrupt catalog (storage.ts:330-355), but that path needsbackend.readto succeed and pass it a bad string. On desktop the decode happens in the main process, the IPC promise rejects,backend.readfails withcatalogError("load"), and the recovery branch never runs. The desktop backend supplies noquarantineeither (storage.ts:272-296), unlike the IndexedDB one at:304. Browser installs heal themselves from this. Desktop installs sit there bricked.Nothing was actually lost, either time
state.sqlitestill held all 5 projects,deleted_atnull on every row, 33 threads, thread activity from the same day. A few hundred bytes of connection metadata took the whole app down with it.The workaround still works
Quit the app fully, rename
~/.t3/userdata/connection-catalog.jsonaside, relaunch. The catalog gets rebuilt and the projects come back. Manually configured remote connections need re-adding.
Filed from
t3 triageon 0.0.42. Diagnosis by Claude Opus 5 (1M context) running in Claude Code.Another occurrence on 0.0.42, Windows 10 Home x64, build 19045 (NTFS), investigated on 24 September 2026. This adds Windows 10 evidence and an isolated reproduction of the desktop-to-renderer failure path.
After restarting the development desktop, T3 Code appeared to have no connected environment. The host server remained accessible from another client.
Evidence from this machine
~/.t3/userdata/connection-catalog.json: 1,552 bytes, all 1,552 NUL. A copy was preserved. The other seven top-level JSON files under 200 KB parsed successfully.- The same outer-document decode failure occurred on three desktop launches:
2026-09-24T13:39:47.513Z desktop.connectionCatalogStore.get Failure 2026-09-24T13:41:08.201Z desktop.connectionCatalogStore.get Failure 2026-09-24T13:46:47.423Z desktop.connectionCatalogStore.get Failure DesktopConnectionCatalogStoreDocumentDecodeError: Failed to decode the desktop connection catalog document at ~/.t3/userdata/connection-catalog.json. Cause: SchemaError: Expected a valid JSON string- Each launch also recorded backend readiness.
GET http://127.0.0.1:3773/.well-known/t3/environmentreturned HTTP 200, reporting version 0.0.42. - Read-only
PRAGMA quick_checkreturnedok. The database still contained 38 nondeleted projects and 536 nondeleted threads, including the work in progress. - Windows recorded Kernel-Power event 41 at 15:12:42 CEST, then a user-initiated restart at 15:15:03. The catalog's recorded modification time is 15:12:10. Event 6008 reports an earlier unexpected shutdown at 14:54:09. An unclean shutdown is a plausible trigger; these timestamps do not prove the exact corruption mechanism or time.
Reproduction without interrupting the live host
Two isolated tests used the real desktop catalog store, a temporary T3 home, and the repository's test double for OS encryption:
- Write a synthetic 1,552-byte NUL catalog. Call
gettwice. Both calls fail withDesktopConnectionCatalogStoreDocumentDecodeError, leaving the fixture unchanged. Rename only that fixture to a quarantine filename. The next read returnsOption.none(). - Connect that store to the real renderer catalog backend through a test IPC bridge. Reading fails with
ConnectionTransientError, and no recovery write occurs. After quarantining the fixture, the renderer can load an empty catalog.
Both tests passed, along with nine existing desktop catalog tests and 44 existing web storage/first-run tests. The desktop tests used an isolated configuration to resolve already-installed workspace dependencies. No source changes, live catalog reset, power-loss experiment, or full desktop UI recovery were performed.
This confirms the distinction described above: the desktop read rejects before the renderer recovery branch can run. Registry construction then fails even though the local server is healthy.
The installed
app.asaralready uses temporary-file write followed by rename. Its catalog write path has no explicit durable flush or backup. The current writer should therefore not be described as writing in place, and the missing flush remains a possible contributor rather than a proven cause here.Recovery should preserve the damaged file and let a healthy local environment start, with an explicit connection-settings error. Malformed documents need to remain distinct from permission failures and temporarily unavailable encryption keys so recovery does not overwrite recoverable credentials.
Follow-up: proposed fix and verification
PR #13468 contains the catalog recovery fix and its regression tests. Its description links back to this investigation. Subsequent checks verified recovery in an isolated built desktop app, preservation of the damaged bytes, native Windows encryption, and a full app restart. The PR records the test coverage and limits for anyone investigating an alternative fix. The original corruption trigger remains unproven; no physical power-loss experiment was performed.
Hit this on current
main(148e6deea0, desktop dev build viastart:desktop), so it is still unfixed there as of Oct 1.- Windows 11 Pro x64, build 26300, NTFS
- Power cut on Sep 30 while the app was open.
connection-catalog.jsonmtime is 23:23 that night. - The file is 904 bytes, all NUL. Every other JSON file in
~/.t3/userdatahas no NUL bytes, andstate.sqlitehas a valid header. - On relaunch the backend reports ready on
127.0.0.1:3773, then the renderer's catalog read fails:
Error occurred in handler for 'desktop:get-connection-catalog': { _tag: 'DesktopConnectionCatalogStoreDocumentDecodeError', catalogPath: 'C:\\Users\\<user>\\.t3\\userdata\\connection-catalog.json', ... }The local environment never connects, and projects and providers are missing from the UI.
I wrote the quarantine-and-treat-as-missing change locally against
mainas a check. A regression test with a 904-byte zero-filled file fails without it and passes with it. That is the same approach as #13468, which also flushes the temp file before the rename, so I'd rather back that PR than open a duplicate.
Summary
~/.t3/userdata/connection-catalog.jsonappears to be written in place, non-atomically, on a fixed ~3 second cadence. If the process dies mid-write (power loss, hard reset, OOM kill), NTFS commits the file's new size but not its data, leaving the file filled entirely with NUL bytes.On the next launch,
desktop.connectionCatalogStore.getthrows a decode error and there is no fallback path — the app comes up with an empty project list and a connection that spins on "Connecting…" forever. It never self-heals, because the renderer only callsconnectionCatalog.setafter a successfulget, so no subsequent write ever repairs the file.The result is a completely unusable app from a single unlucky 3-second window, with no in-app indication of what's wrong.
Impact
After an unclean shutdown, the app launches into this state permanently:
state.sqliteCritically, no data is actually lost.
state.sqliteis completely intact. Only a small piece of connection metadata is destroyed, but it takes the entire app down with it.Reproduction
~/.t3/userdata/connection-catalog.jsonExpected: app recovers, falls back to a default catalog, rediscovers local/WSL connections Actual: empty project list, connection hangs forever, no error shown
Evidence
The file is 100% NUL bytes
Every other JSON file under
.t3/parsed cleanly. This was the only corrupt file on disk.Timeline from
desktop.trace.ndjsonThe error
It propagates all the way up the IPC chain — every one of these fails together:
The write cadence is a constant 3.01s
Gap histogram between consecutive
writeDocumentspans in a single session:That's ~1,200 writes/hour of what is essentially static connection metadata. The spans carry no payload, so I can't confirm the content is unchanged between writes, but a perfectly constant 3.01s interval strongly suggests a polling timer that writes unconditionally rather than on change.
An atomic writer already exists elsewhere in the codebase
I found this sitting next to the corrupt file:
That
<name>.<pid>.<hash>.tmppattern is a write-to-temp-then-rename. The settings store was also interrupted by the same power loss — and its real file survived perfectly intact, leaving only an orphaned temp file behind. The connection catalog store, with no such sibling, was destroyed in place.So the fix pattern is already present in the codebase; the catalog store just isn't using it.
Suggested fixes
In rough priority order:
connection-catalog.json.<pid>.<rand>.tmp,fsync, thenrename()over the target. On POSIX and NTFS alike, rename is atomic — a torn write can then only ever leave an orphaned temp file, never a destroyed catalog.connection-catalog.json.corrupt-<timestamp>, and fall back to the default catalog so local/WSL connections get rediscovered. A config file that can't be parsed should degrade, not brick.(1) and (2) are independent and both worth having: (1) prevents the corruption, (2) makes the app survive it if it happens anyway through some other route.
Workaround for anyone hitting this now
Quit T3 Code completely (including any lingering backend child process), then:
Relaunch. The catalog is regenerated, projects reappear from
state.sqlite, and local/WSL connections are rediscovered. Nothing is lost — the corrupt file contained no recoverable bytes anyway. Any manually configured remote connections may need to be re-added.Environment
t3@0.0.29wslBackendEnabled: true)Possibly unrelated, but noticed while reading the logs
desktop.trace.ndjsonrecords SSH remote pairing tokens in plaintext inside span events:Since these trace files are exactly what users get asked to attach to bug reports, it may be worth redacting tokens at the log sink. Happy to file this separately if you'd prefer.