Skip to content

[Bug]: After a preview automation timeout, the session fails over to a client on another machine that doesn't have the thread open #13540

Description

@QarthO

Filed by Claude (Opus 5.5, running inside T3 Code) on behalf of @QarthO, at his request. The investigation and reproduction were done by the agent on his machine; he confirmed what he saw on screen at each step.

Before submitting

Area

apps/server (preview automation broker), with visible effects in apps/web

Summary

When a preview automation request times out, the broker disconnects the host and the agent's session fails over to another connected client. That other client can be a desktop app on a different computer that doesn't have the thread open. It then loads the preview there, so local-only URLs (localhost, 127.0.0.1, *.orb.local, etc.) fail with "This site can't be reached". Because tab state is shared, the failure also shows up in the preview the user is actually watching. From the user's side, the preview looks like it breaks every time the agent touches it.

Steps to reproduce

  1. Run T3 Code on machine A (the server runs here too).
  2. Open the desktop app on machine B and connect it to machine A's server. Leave it on the new-thread screen; don't open the thread.
  3. On machine A, open a thread and a preview tab on a URL that only resolves on machine A (for example http://localhost:3000, or an OrbStack https://myapp.orb.local).
  4. Cause a preview automation timeout on machine A. The easiest way is [Bug]: preview_resize always times out (and evicts the automation host) when the main window is zoomed, because the guest webview reports innerWidth × hostZoom #12319: zoom the UI in once (View > Zoom In), then have the agent call preview_resize with a freeform size. It times out after 15 s.
  5. Have the agent call preview_evaluate or preview_navigate on the same tab.

Expected behavior

After the timeout, the agent keeps controlling the preview on machine A (the one showing the thread), or gets a clear error. It shouldn't silently move to a client on another machine that doesn't have the thread open.

Actual behavior

  • preview_evaluate returns chrome-error://chromewebdata/, and its results don't match what machine A shows. The page it reports has a different creation time from the one on screen.
  • preview_navigate reports success, but the page is an error page. On machine A the preview flashes the real page for a split second, then shows "This site can't be reached … ERR_NAME_NOT_RESOLVED" (for 127.0.0.1 URLs it's the same, pointing at machine B's own localhost).
  • Clicking Reload on machine A always works immediately, but the agent's next preview action breaks it again.
  • Public URLs (e.g. https://example.com) work, which makes it look like a DNS or network problem on machine A. It isn't: curl and other browsers on machine A reach the site fine.
  • It persists until machine B's app is closed. After closing it, the same tab immediately responded with machine A's page, and the same preview_resize (with the zoom reset) succeeded straight away.

Evidence it's machine B handling the commands: machine A's desktop.trace.ndjson has no PreviewManager.automationEvaluate / navigate spans for any of the agent's calls during the failure. As soon as machine B's app was closed, they appear again. The server trace shows the handover at the timeout:

21:59:12  preview_resize (freeform 992x800)      -> times out (machine A's UI zoomed)
21:59:27  PreviewAutomationBroker.closeConnection / disconnect
21:59:28  PreviewAutomationBroker.acquireConnection / connect   (machine A reconnects)
          ... every later preview call goes to machine B until its app is closed

Why it happens (from reading apps/server/src/mcp/PreviewAutomationBroker.ts)

  1. On timeout, awaitResponse disconnects the host (as described in [Bug]: Optional 500 ms preview metadata timeout disconnects the automation host #12273 / [Bug]: preview_resize always times out (and evicts the automation host) when the main window is zoomed, because the guest webview reports innerWidth × hostZoom #12319).
  2. The next request looks for a host in the environment. Hosts that own the target tab are sorted first, but a host without the thread or tab is still eligible. If machine A hasn't reconnected yet (it came back about a second later here), machine B is the only candidate and gets the assignment.
  3. The assignment is then pinned: "a live assignment … is not silently moved to a newer client". So even after machine A reconnects, is focused and owns the visible tab, the session stays on machine B. (fix(preview): use the visible browser for new agent sessions #13064's "prefer the client showing the tab" only applies when there's no live assignment, and at the moment of failover machine A wasn't connected to be preferred.)
  4. Machine B's preview host loads the tab URL on its own machine and network. The resulting load failure is written to the shared tab state, which is why machine A's panel shows the error too.

Possible solutions

  • Only fail over to hosts that have the thread's preview open. Treat ownsTargetTab as a requirement rather than a sort key. If no host qualifies, return a clear "no host has this preview open" error instead of picking any client.
  • Let a reconnecting host reclaim its session (the follow-up fix(preview): use the visible browser for new agent sessions #13064 left out). If the host that just lost its assignment reconnects within a short grace period, or if a focused host owns the visible target tab while the pinned host doesn't, move the assignment back instead of keeping it pinned to the other client.
  • Make the handover visible. Include which client or device handled a command in tool results and errors (today it's only failed on client preview-<id>). Also consider keeping one client's load failure out of the shared tab state that other clients render.

Fixing #12319 and not evicting hosts on timeouts (#12273) would remove the usual trigger, but any other timeout can still cause this handover.

Impact

Major degradation or frequent failure. The agent can't use the preview at all, and it looks like the agent is breaking the user's preview, which is confusing to debug.

Version or commit

T3 Code (Nightly) 0.0.43-nightly.20260924.2223 (Electron 44.4.2, Chrome 152) on machine A. Machine B was the Windows desktop app.

Environment

Machine A: macOS, Apple Silicon, running the server and the desktop app; the dev site is served by OrbStack at an *.orb.local HTTPS address. Machine B: Windows desktop app connected to machine A's server over the network, idle on the new-thread screen. Agent provider: Claude Code.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions