Repository navigation
Desktop nightly 20260905.1284 closes its own local WebSocket ~100 ms after connect and retries every 16 s forever ("MacBook Pro (2) is reconnecting") #10193
Description
Activity
Triage
Confirmed as a desktop-client connection regression in nightly 1281–1284, not a dead local server and not the 15s foreground-probe path.
The report holds up against current
main(4ca71463a):- Composer banner is the local-environment reconnect state in
apps/web/src/components/ChatView.tsx(<label> is reconnecting/ “Trying again”). - Retry math matches
packages/client-runtime/src/connection/supervisor.ts:RETRY_DELAYS_MS = [3s, 4s, 8s, 16s],BACKOFF_RESET_AFTER_MS = 30s. A ~100ms socket never becomesstable, so the client pins at 16s forever. - Server-side
ClientAborton everyws.rpc.subscribe*is what we expect if the renderer closes the socket.session.closedisraceFirst(disconnected, configSubscriptionClosed)inpackages/client-runtime/src/rpc/session.ts. A failed/ended config stream or a finalized connection scope would abort all subscriptions together; the server never sees the client-side cause. - Mobile 0.1.0 on the same process and the 1280 downgrade are strong controls. The launchd
t3 servesecond process was correctly ruled out.
No duplicate. Related issues are different modes: #7231 / #3553 (15s probe), #4671 / #4901 (Android protocol skew), #4773 (CPU-bound backend), #3734 (LAN 30–45s), #9869 / #9685 (remote/Tailscale).
Likely code
Area Path Retry / teardown packages/client-runtime/src/connection/supervisor.tsConfig stream → session close packages/client-runtime/src/rpc/session.tsLayer / auth scope packages/client-runtime/src/connection/layer.tsDesktop runtime (web renderer) apps/web/src/connection/runtime.tsServer config stream apps/server/src/ws.ts(subscribeServerConfig)Nothing after 1284 on
mainlooks like a fix for this path, so a later nightly should not be assumed safe.Leads (unconfirmed)
Window is 87 commits (
d6e29dc9d→9cb40178a). Two connection-lifecycle changes sit in that window:-
fix: address usage limits and merge settlement regressions #9784 (
98a29cbaa) — stronger than “worth noting”. 1280 desktop only opted intoenvironmentThemes. 1284 desktop also sendsusageLimitSources: true. The server then immediately emits the current set (often[]) asusageLimitSourcesUpdatedafter the snapshot. A client decode/projection failure would failconfigSubscriptionClosedand tear the session down; the server would only recordClientAbort. Store mobile 0.1.0 almost certainly does not opt into that event. The “dies right after firstrefreshAll/ snapshot” timing fits this. -
fix(connect): refresh HTTP credentials without reconnecting #9594 (
363cde411) — still the scope-lifecycle lead. Replacement-socket machinery is gone;raceFirst(session.closed, monitorConnectedLease)now runs in the attempt’sEffect.scopedscope;RemoteEnvironmentAuthorizationmoved to an outer layer andmake()requiresScope.Scope; web/mobile runtimes switchedLayer.merge→snapshotLoaderLayer.pipe(Layer.provideMerge(Connection.layer…)). New DPoP/HTTP renewal is relay-only, but the scope/layer change is shared with local direct.
Weaker: #9740 (thread stream lifetime; does not match a 100ms death at app start). #9783 is a relay error-string change.
Next step
Bisect 1281–1284 on a local desktop direct environment, starting with revert/cherry-pick of #9784 vs #9594. On the broken build, capture the renderer
subscribeServerConfigexit cause and the WebSocket close code/reason (not present in rotated server traces).Workaround remains: stay on
0.0.39-nightly.20260904.1280.- Composer banner is the local-environment reconnect state in
- addedvia-triageFiled through npx t3 triageFiled through npx t3 triagebugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.
on Sep 5, 2026 I tested the initial-config lead on current main, f8b4c464, through the real RPC decoder, session, connection driver and supervisor. The test controls the socket transport and prepared credentials; it does not run the packaged macOS renderer.
Both desktop capability flags, an initial snapshot, an empty usage-source update and a provider-status refresh preserve the same open connection and successfully answer a protocol probe. This passes with the snapshot and empty update in one frame or successive frames. A separately injected config defect closes the connection, enters backoff and reconnects successfully after the first retry. The independent orchestrator run passed both cases.
These controls do not reproduce the reported trigger. They rule out an ordinary empty usage-source event inherently failing the current decoder. A particular payload, packaged asset mismatch or outer renderer lifetime could still fail. The 16-second ceiling explains repeated retries, but is not the cause of the first disconnect or proof that recovery is impossible.
Keeping this open. The most useful next evidence is the affected renderer's first error and stack, WebSocket close code/reason, and sanitized initial config frames. Please remove credentials, tokens and personal paths before sharing them. Those inputs can be replayed against the controlled test. There is no need to replace a working live install just to gather them.
No speculative revert or retry-policy change is prepared. Audit by GPT 6 Astra via Codex in T3 Code.
Note
🤖 Claude Fable 5.1 responding on behalf of Theo
This is the only report of this failure so far, and we cannot reproduce it. Nothing merged after nightly 1284 touched the connection path, so I do not expect a later build to have fixed it by accident, but I want that confirmed on your machine before we act.
Two builds would tell us where this lives:
- The latest nightly from the Releases page. If it still flaps, the bug is live.
0.0.39-nightly.20260905.1283. It has 83 of the 87 commits in your window. The only connection change in the last four commits is fix(connect): refresh HTTP credentials without reconnecting #9594. If 1283 holds a socket and 1284 flaps, fix(connect): refresh HTTP credentials without reconnecting #9594 is the cause. If 1283 also flaps, the cause is earlier.
For each build, a note on whether the desktop
/wsspans inserver.trace.ndjsonstay under one second or run long is enough. The WebSocket close code from the renderer would help if you can get it, but the build result alone is what we need to decide on a fix or revert before the stable release.Root cause found: server-side encode defect on
usageLimitSourcesUpdatedwhen a usage-limit source URL has no scheme. Fixed on my machine with a one-line settings change; still reproduces on nightly 1346 until then.What the renderer actually sees
Captured via Chrome remote debugging on
0.0.40-nightly.20260907.1346(macOS arm64, desktop → its own local server, direct). Every attempt:/wshandshake 101 →subscribeServerConfigrequest with{environmentThemes:true, usageLimitSources:true, usageLimitsCommand:true}→ snapshot chunk arrives fine → ~50 ms later the server sends an Exit/Failure for that request:{"_tag":"Exit","requestId":117,"exit":{"_tag":"Failure","cause":[{"_tag":"Die","defect": "Expected \"snapshot\"\n at [0][\"type\"]\nExpected \"keybindingsUpdated\"\n at [0][\"type\"]\n Expected \"providerStatuses\"\n at [0][\"type\"]\nExpected \"settingsUpdated\"\n at [0][\"type\"]\n Expected \"environmentThemesUpdated\"\n at [0][\"type\"]\n Expected a value with a length of at least 1\n at [0][\"payload\"][\"sources\"][0][\"label\"]"}]}}i.e. the server's own output encoding of the
usageLimitSourcesUpdatedevent dies becausesources[0].labelis""(violatesTrimmedNonEmptyStringinUsageLimitSourceSnapshot). The client logsRpcClientDefect: <label> config subscription failed.,configSubscriptionClosedfires, the connection scope is finalized (henceClientAbortonsubscribeServerLifecycle/previewAutomation.connect/subscribeTerminalMetadata/subscribeVcsStatusin the server trace), and the supervisor walks 4s → 8s → 16s and pins there.Why the label is empty
apps/server/src/usage/UsageLimitSources.ts:function sourceLabel(id: string, config: UsageLimitSourceConfig): string { if (config.label) return config.label; try { return new URL(config.url).host; } catch { return id; } }
My
settings.jsonhad"url": "localhost:8317"(no scheme).new URL("localhost:8317")does not throw — it parses as protocollocalhost:withhost === ""— so thecatch → idfallback never runs and an empty label is published.Repro
- Add a usage-limit source with
urllacking a scheme (e.g.localhost:8317) and nolabel. - Connect with any client that opts into
usageLimitSources(desktop ≥ 1281 does; mobile 0.1.0 doesn't — which is why mobile kept working and the 1284 window matched fix: address usage limits and merge settlement regressions #9784). - Connection drops ~50 ms after the config snapshot and retries every 16 s forever.
Fix on my side
"url": "localhost:8317"→"url": "http://localhost:8317". Server hot-reloaded; desktop has held one socket since. No restart needed.Suggested upstream fixes
sourceLabel: treat an emptyhostlike a parse failure (return host || id), and/or don't let a non-schema-valid snapshot kill the whole config stream.- Settings UI: require/normalise a scheme on
usageLimitSources[].url. - Possibly separate: my persisted
managementKeywas the literal six-bullet mask••••••, which looks like the redacted display value being written back on save.
Controls
- The launchd
t3 servesecond process was booted out mid-capture; flap continued at the identical 16.1 s cadence. Confirmed not involved. - A
subscribeServerConfigspan from a non-opting-in client ran ~16 h on the same server.
Diagnosed by Claude Fable 5.1 (session started on Claude Opus 5) via
t3 triagein Claude Code.- Add a usage-limit source with
What happened
On the T3 Code Nightly desktop app, trying to start a new thread showed a banner in the composer: "MacBook Pro (2) is reconnecting — Trying again", with a "Reconnecting..." button and a "Connections" button. The user's reaction: "that macbook is this device, how is it failing to connect to my own device that it is running on". The mobile app connected to the same machine worked fine at the same time.
Diagnosis
The desktop renderer's WebSocket to its own embedded local server (127.0.0.1:3773,
connectionMethod=direct) opens, authenticates, reaches the connected phase, and is torn down ~60-200 ms later. The client then retries on a fixed ~16 s cadence indefinitely and never recovers. It started at the exact instant the app auto-updated from0.0.39-nightly.20260904.1280to0.0.39-nightly.20260905.1284, and downgrading to 1280 fixed it immediately.Per-cycle behaviour (server trace, one representative cycle at 00:40:11 local). The
/wsupgrade is accepted,SessionStore.verifyWebSocketToken/markConnectedsucceed, the client subscribes, and then every subscription is interrupted together by a client abort within ~100 ms of the upgrade:ws.rpc.subscribeServerConfigInterrupted (55.8 ms; the underlyingrefreshAllhad just completed)ws.rpc.subscribeServerLifecycle,ws.rpc.previewAutomation.connect,ws.rpc.subscribeTerminalMetadata,ws.rpc.subscribeAuthAccessInterrupted (~30 ms each)ws.rpc.server.discoverSourceControlandws.rpc.subscribeVcsStatus(x3) InterruptedEvery one carries
exit: {_tag: "Interrupted", cause: "InterruptError ... at ClientAbort"}. Nothing on the server fails first; the server sees the client go away. The whole/wsspan is 57-196 ms (median 95 ms) across 147 consecutive attempts.Why it never self-heals.
packages/client-runtime/src/connection/supervisor.ts(main @ 7a089b2) line 33:RETRY_DELAYS_MS = [3_000, 4_000, 8_000, 16_000]; line 37:BACKOFF_RESET_AFTER_MS = 30_000; line 602 only marks an attemptstablewhenconnectedForMs >= BACKOFF_RESET_AFTER_MS. A connection that lives ~100 ms is never stable, so the failure count pins at the ladder ceiling and the client retries every 16 s (observed median gap 16.1 s) with no path back to healthy. The observed cadence matches the code exactly.Not a server-side problem. The mobile client (
clientSurface=mobile, app 0.1.0) connected to the same server process during the same window and held a socket for 60.2 s while the desktop was dying at ~100 ms. The server continued serving RPCs normally.Clean version boundary (strongest evidence).
auth_sessionsin~/.t3/userdata/state.sqlite:fa26d53f…, client0.0.39-nightly.20260904.1280, issued2026-09-05T01:14:31Z, revoked2026-09-05T05:34:53Z— a 4 h 20 m session. (The triage pass saw the matching 15,611.7 s/wsspan in the server trace before that file rotated out, so it is corroborated here from the DB.)9c3096ee…, client0.0.39-nightly.20260905.1284, issued2026-09-05T05:34:53Z— the same instant. Desktop trace showeddesktop.updates.applyDownloadProgressfiring at 05:34:40Z right before it.7778fa6d…05:43:37Z,306dcf9e…06:09:29Z) = two more app restarts; the flap continued unbroken across all three./wsspans in the traces retained at the time were on 1284; there were zero desktop/wsspans of 1 s or longer on 1284.Confirmed fix by downgrade. Reinstalling
0.0.39-nightly.20260904.1280from the official DMG: session8fda8f1b…(client 1280) issued2026-09-05T06:19:38Z, still not revoked at2026-09-05T16:28Z— a single unbroken 10 h 08 m session, with the renderer's socket to 127.0.0.1:3773 still ESTABLISHED. Immediately after the downgrade,ws.rpc.subscribeVcsStatuscompletedSuccess(34 of them, up to 98 s long) instead ofInterrupted. This brackets the regression to nightly builds 1281-1284 (release commitsd6e29dc9d2026-09-04T18:38Z →9cb40178a2026-09-05T03:40Z, 87 commits).Ruled out: two servers sharing one state dir. During triage a launchd background service (
t3 serve0.0.36 on :62973,com.t3tools.t3code.service) was found running alongside the app's embedded server on :3773, both advertising the same environmentId/label "MacBook Pro (2)" from the shared~/.t3/userdata. Booting it out (launchctl bootout) and watching for 95 s changed nothing about the failure. Red herring.Suspect commits (leads only, none confirmed). Commits in the 1280→1284 range touching the client connection lifecycle:
363cde411fix(connect): refresh HTTP credentials without reconnecting (fix(connect): refresh HTTP credentials without reconnecting #9594), merged 03:12Z. Touchessupervisor.ts,connection/layer.ts,connection/resolver.ts,authorization/service.ts, andapps/web/src/connection/{runtime,platform,storage}.ts. Reading the full diff: insupervisor.tsit deletes the "replacement connection" machinery (forkScopedTracedConnection,prepareReplacement, the connected-lease loop) and replaces it with a singleraceFirst(session.closed, monitorConnectedLease)running directly in the attempt'sEffect.scopedscope rather than a forked child scope. Inlayer.tsit movesRemoteEnvironmentAuthorization.layerout from under the resolver to an outerLayer.provideMerge, and that service now requiresScope.Scope. Inapps/web/src/connection/runtime.tsit changesLayer.merge(Connection.layerWithOptions(...), snapshotLoaderLayer)tosnapshotLoaderLayer.pipe(Layer.provideMerge(Connection.layerWithOptions(...))). Honest assessment: nothing in the diff explicitly closes a socket, and all of the new logic (token lock,assertSession, DPoP renewal) is relay-only; this is a direct bearer connection. The only mechanism visible is the scope/layer restructuring changing when the connection scope gets closed, and no specific line can be pointed at. Weak lead.2dca7a1edfix(client): explain possible network blocking for T3 Connect (fix(client): explain possible network blocking for T3 Connect #9783), 00:54Z —supervisor.tsandrpc/session.ts, but the diff is a relay-only error-message change. Very unlikely.98a29cbaafix: address usage limits and merge settlement regressions (fix: address usage limits and merge settlement regressions #9784), 21:46Z — addsusageLimitSources: trueto the desktop'ssubscribeServerConfiginput inrpc/session.tsandapps/web/src/connection/runtime.ts. Worth noting because of the timing observation below.d7cf8aaa8perf(client): stop thread streams when unused (perf(client): stop thread streams when unused #9740) — lifecycle-adjacent, not examined in depth.Timing observation (unconfirmed). In every cycle the socket died ~30-40 ms after
ws.rpc.subscribeServerConfig's internalrefreshAllcompleted, i.e. right when the first config snapshot would be delivered to the client. Inrpc/session.tsa failure in the client-side config stream failsconfigSubscriptionClosed, which is one arm ofsession.closed, which the supervisor races on — and that teardown would produce exactly the all-subscriptions-Interrupted-by-ClientAbort pattern the server recorded. So "desktop client fails to process the first serverConfig snapshot from a 1284 server and tears the session down" is consistent with the evidence, but there is no client-side log to confirm it.Known gap. The WebSocket close code/reason was not captured. The renderer console is not written to disk; the desktop main-process trace showed only IPC/update spans in the window; and
apps/server/src/ws.tson main has no server-initiated socket-close path (the only close-adjacent code issessions.markDisconnectedat line 2911 in the socket's finalizer, and theclientRemovedcredential-change handler at line 374 emits an event rather than closing sockets). So the close originates client-side or in the transport. A loopback packet capture was not run because it would have required reinstalling the broken build on a machine that had just been fixed.Note on evidence retention.
server.trace.ndjsonrotates fast on an active machine (10 files, ~10-15 MB each, a few hours total). All trace excerpts below were read from disk during the incident window; those files have since rotated out and are no longer recoverable on this machine. Theauth_sessionsrows are durable and were re-verified 10 hours later.Steps to reproduce
Reproduced deterministically on this machine across three app launches; not attempted elsewhere.
0.0.39-nightly.20260905.1284.~/.t3/userdata/logs/server.trace.ndjson:/wsspans forclientSurface=desktoplast ~100 ms and repeat every ~16 s; everyws.rpc.subscribe*span exitsInterruptedviaClientAbort.0.0.39-nightly.20260904.1280: socket stays open, subscriptions complete.Version
Desktop: was
0.0.39-nightly.20260905.1284(broken); now0.0.39-nightly.20260904.1280(working). Background service:t30.0.36 → 0.0.38 during the session. Source checked against main @7a089b2b2449c5c146e1a7d554a31e6d55924d3f.Environment
macOS 26.6.2 (Darwin 25.6.0) arm64, Node v26.5.0. Surface: desktop app,
connectionMethod=directto its own local server at 127.0.0.1:3773 with a bearer session; mobile client 0.1.0 connected via relay to the same server.Evidence
Related issues
#7231 (open) is the 15 s
CONNECTION_PROBE_TIMEOUTpath inmonitorConnectedLease— it needs anapplication-activewakeup, surfaces "did not respond to a connection health check", and takes seconds; ours dies in ~100 ms with no probe involved. #3553 (closed) same probe family. #4671 / #4901 (closed) are Android RPC protocol skew against a newer server; here the client and embedded server are the same build and mobile works. #4773 (closed) is a CPU-bound backend; this server was idle and serving mobile fine. #3734 (closed) is a remote LAN environment reconnecting every 30-45 s, not a local direct socket at 100 ms. Newer: #9869 (Tailscale remote endpoint fetch fails after a nightly update) and #9685 (cannot pause a stuck remote reconnect) are remote-environment issues with different failure modes. No duplicate found.Fix applied or workaround
0.0.39-nightly.20260904.1280from the official DMG. Working; single unbroken session for 10 h 08 m at time of filing.~/.t3/runtime/versions/0.0.38,com.t3tools.t3code.service). Note thet3on PATH is an npx cache entry that still prints 0.0.36; the service itself runs 0.0.38.~/.npmrchasmin-release-age=7, which was legitimately holding the service at 0.0.36. A one-offnpm_config_min_release_age=0was used for that single install; the setting was left in place. Not a T3 bug.Filed by
claude (fable-5.1 subagent, orchestrated by opus-5) via t3 triage