Skip to content

[Bug]: Preview automation host disappears mid-session ("No preview automation host is available for <op> in environment …") and never reconnects #12146

Description

@GoblinRules

Summary

During a long agent session the preview automation host silently disappears. Every preview tool call (preview_open, preview_status, preview_evaluate, preview_navigate) then fails with:

No preview automation host is available for <op> in environment 3327ef5d-c1f1-4d71-8797-edb75589bfe1.

The T3 Code window is still open, the collaborative browser tab is still visible to the user, and the user is actively looking at it, but the agent can no longer drive it. The host never reconnects for the rest of the session. The only recovery is presumably restarting the app (not tried, session in progress).

Environment

  • T3 Code (Nightly) 0.0.43-nightly.20260916.1811, Windows 10 Pro 10.0.19045
  • Agent: Claude Code via the claudeAgent provider (Fable 5.1)
  • Preview tabs used in the session: ~11 (tab_1 … tab_b), created with preview_open (reuseExistingTab: false several times)

Timeline (single session, ~12 h)

  1. Preview automation worked for hours on http://localhost:5173 (a dev server) and a Portainer instance: navigate, click, type, evaluate, snapshot.
  2. preview_snapshot started failing intermittently on tall pages with PreviewAutomationExecutionError and once PreviewAutomationTimeoutError, while preview_evaluate on the same tab kept working. Closing/reopening tabs did not change that.
  3. Some preview_evaluate calls failed with Preview automation evaluate failed on client preview-f3e97ff86f65d5f30253a4d560504301. — after that the affected tab id was dead, but a new tab (reuseExistingTab:false) worked.
  4. Later, on a new tab pointed at https://api.slack.com/apps (signed-in Slack app settings), a sequence of preview_click + preview_evaluate worked, then one preview_evaluate returned No preview automation host is available for evaluate in environment ….
  5. From that point on, every preview tool (preview_open, preview_status with and without tabId) fails with the same message. device_list reports "Agent device access is turned off for this environment" (expected, it was never enabled).
  6. The user confirms the T3 window and the Slack tab are still open and usable by hand.

Expected

Either the automation host reconnects (or a new host is attached to the environment) and tool calls resume, or the tool error tells the agent/user what to do (e.g. "reopen the preview panel", "restart the app") so the session can recover without guessing.

Actual

Permanent No preview automation host is available … for the rest of the session; no hint about recovery; the user sees a working browser tab while the agent cannot reach it.

Possibly related

Happy to run a debug build or collect logs from ~/.t3/userdata/logs if you point me at which file the preview host writes to (the provider events.*.log only contains agent events).

Activity

  1. juliusmarminge commented on Sep 16, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed as a real host-lifetime bug, complementary to #11167. The visible preview tab is not the automation host. After the desktop host stream drops, the broker has nobody to route to, and the client does not open a new previewAutomation.connect for the rest of the session.

    What the report matches

    No preview automation host is available for <op> in environment … is PreviewAutomationNoAvailableHostError from PreviewAutomationBroker.invoke when clients has no live host for that environment (or a pinned assignment is dead and there is no failover target). That is environment-wide: preview_status and preview_open fail the same way because they are also invoke.

    The reported preview-f3e97ff… id is a desktop host (createPreviewAutomationClientId → preview-<hex>). PreviewAutomationHosts only mounts on Electron (isElectron && previewBridge?.automation). The collaborative browser tab is a separate webview. The window and Slack tab staying usable by hand is consistent with the RPC host being gone.

    device_list being off is unrelated.

    Why this can become permanent (especially on this nightly)

    #11381 (merged 2026-09-15, in 0.0.43-nightly.20260916.1811) implemented the #11167 suggestion: an unanswered broker deadline (default 15s) now disconnects that host, shuts down its request queue, and clears assignments. Hosts that reply with PreviewAutomationTimeoutError stay registered.

    The #11381 regression only proves failover when a second host is already connected. A typical desktop session has one host. Evicting it leaves clients empty.

    The desktop subscription does not heal that:

    • previewAutomation.connect is an acquire/release stream. Queue shutdown ends it (tests see an interrupt).
    • subscribe() in packages/client-runtime/src/rpc/client.ts resubscribes on environment session change (and optional retryExpectedFailureAfter / resubscribe, which automation does not set). It does not reopen a stream that completed while the websocket session stayed up.
    • PreviewAutomationHosts comment about “subscription runtime own reconnects” means WS session reconnect, not “broker evicted us, connect again.”
    • focusHost cannot recreate a host (does not create host state from focus updates without a live stream).

    So: timeout or any other queue/stream drop → sole host gone → every later preview_* is NoAvailableHostError until the renderer remounts the host or the environment session reconnects (in practice, restart the app).

    That also explains the earlier, milder failures in this session:

    #11295 is not this failure (that one bricks provider history with huge inline PNGs). Large/tall snapshots are still a plausible trigger for an unanswered broker deadline.

    Not duplicates

    Issue / PR Why not this
    #11167 / #11381 Opposite hole: zombie host stayed selected. #11381 evicts; this report is eviction without re-register.
    #3713 / #8486 Tab poisoned; preview_open still works.
    #6355 / #7200 Thread-switch CDP detach; broker still has a host.
    #6317 macOS last-window close destroys the renderer. Reporter is Windows, window still open.
    #11295 / #11624 Session-history image size, not host routing.
    #8981 (open) Desktop CDP/budget hardening, not connect-stream re-register.
    #7870 Tools vanish after server restart (MCP credentials).

    Suggested fix

    1. Desktop host (primary): if previewAutomation.connect ends while PreviewAutomationHosts is still mounted, open a new connect stream (same pattern as session resubscribe / retryExpectedFailureAfter). Do not wait for a websocket reconnect or app restart.
    2. Error text: PreviewAutomationNoAvailableHostError should say the collaborative tab is not the host, and to recover: keep the desktop window open and reload/restart T3 Code so the automation host can register again.
    3. Broker: eviction/failover must work when the evicted host was the only one. Either wait for that client to reconnect before failing forever, or treat “no hosts” as a distinct recoverable condition. Do not keep a dead assignment.

    Optional: preview_status should say registered vs responsive and which clientId would take the next command (#11167 leftover).

    Workaround

    Restart T3 Code (or otherwise remount the desktop renderer / reconnect the environment websocket). Opening another preview tab will not help once preview_status itself returns no host. includeImage: false on snapshots avoids the #11295 class; it does not restore a dropped host.

    Logs (if you still have the session)

    There is no preview-host file under ~/.t3/userdata/logs. Provider events.*.log will not show this. The useful artifact is:

    ~/.t3/userdata/logs/server.trace.ndjson

    Filter for PreviewAutomationBroker.invoke, PreviewAutomationBroker.disconnect, PreviewAutomationBroker.connect, and EnvironmentRpc.subscribe with rpc.method previewAutomation.connect. A disconnect / queue shutdown followed by no new connect while the window stayed up is the smoking gun.

    No code change from this triage.

  2. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 16, 2026
  3. yesiamdaniel commented on Sep 17, 2026

    @yesiamdaniel

    +1 running into this a lot right now. Making it hard for agents to browser use reliably...

  4. GoblinRules commented on Sep 17, 2026

    @GoblinRules
    Author

    Reproduced again after a full restart of T3 Code (same nightly). The preview worked for ~12 tool calls (open, navigate, evaluate, click), then one preview_evaluate failed with Preview automation evaluate failed on client preview-…, and a few calls later every operation returned No preview automation host is available for <op> in environment … again. preview_open with reuseExistingTab=false does not recover it. The drop happened right after a cross-origin navigation (api.slack.com → app.slack.com redirect), same as the first occurrence.

  5. m13v commented on Sep 17, 2026

    @m13v

    the 'attach a fresh host to the environment' half of the ask is where it bites. a new host resumes tool calls but comes up blind to the ~11 tabs the dead one owned, so preview_open works again while your signed-in slack tab and the localhost:5173 session are orphaned in the old process. you'd trade a clear 'no host available' for calls that succeed against a host that can't see the tab you're actually looking at. the reconnect worth building rehydrates the existing tab handles, not just the environment binding.

  6. MOEG-5 commented on Sep 17, 2026

    @MOEG-5

    Additional confirmed occurrence on Linux x86_64, T3 Code Desktop 0.0.42 AppImage, commit 719a76c, Codex provider, on 17 September 2026.

    I still had the server trace and checked the lifecycle sequence requested above. The source was server.trace.ndjson plus its rotated siblings, not provider logs. Times below are UTC.

    Time Evidence
    19:19:22.841 Broker invoke starts; eventually fails with PreviewAutomationTimeoutError for evaluate after 15,002.7 ms.
    19:19:37.841 PreviewAutomationBroker.disconnect and closeConnection start in the same trace as that timed-out invoke.
    19:19:37.842 ws.rpc.previewAutomation.connect stream ends with Interrupted; its cleanup also calls disconnect.
    19:19:40.449 onward Repeated PreviewAutomationNoAvailableHostError for type/evaluate/status/open.
    Until 19:29:18.510 No new PreviewAutomationBroker.connect or acquireConnection in the inspected trace interval. Seven successful focusHost calls are recorded during the gap. The application window/tabs remained visible to the user.
    19:29:18.510 New short connect and acquireConnection spans, correlating with the user restarting T3. The gap was about 9 minutes 41 seconds.
    19:30:08.713 A status deadline produces another disconnect. A concurrent status request receives PreviewAutomationClientDisconnectedError. The RPC stream ends with Interrupted at 19:30:08.725.
    19:30:17.015 preview_open fails with PreviewAutomationNoAvailableHostError.
    19:31:16.401 Another new connect/acquireConnection pair, correlating with a second user restart.

    The attached sanitized lifecycle excerpt preserves names, start/end times, durations, result/error summaries, RPC method attributes, and consistently pseudonymized trace/span/parent IDs. It contains no URLs, page contents, account details, local paths, original client/environment IDs, or stack traces.

    Trace interpretation: this server build names the RPC stream span ws.rpc.previewAutomation.connect, with rpc.method=previewAutomation.connect. No completed EnvironmentRpc.subscribe span was present in the retained server traces. The no-reconnect finding uses the absence of the short broker connect/acquireConnection spans plus continuing NoAvailableHost failures; it does not rely on absence of a completed long-lived subscription span. Successful focusHost calls do not prove registration: the installed implementation can return success without recreating a missing host. The duplicate disconnect calls at each outage are the timeout path plus stream cleanup, not four separate outages.

    Read-only inspection of the installed bundle also confirms that an unanswered broker deadline calls disconnect and shuts down that connection's queue, consistent with the triage diagnosis. The initial reason evaluate/status failed to answer is still unknown; this report does not claim a confirmed tab, website, or concurrency trigger.

    A further code-path concern: src/mcp/toolkits/preview/tools.ts performs an automatic broker status call after evaluate/click/type, with timeoutMs: 500, to obtain page URL metadata for the tool icon. It catches the error, but that status deadline would already have disconnected the host. This is verified in the installed code, not the demonstrated trigger here: the observed unanswered deadlines were 15 seconds.

    No reproduction was attempted on the completed external forms. No browser or application state was changed for this trace inspection.

    sanitized-host-lifecycle.txt

  7. Gal-WPsite commented on Sep 17, 2026

    @Gal-WPsite

    Posting on behalf of Gal Hadad (@Gal-WPsite), at his request; I am Claude (Fable 5.1) running inside T3 Code, and the traces below come from his Mac.

    Another confirmed occurrence, with a deterministic trigger this time: macOS, T3 Code (Alpha) 0.0.42, local environment, 2026-09-17 23:50 Asia/Jerusalem.

    Lifecycle from server.trace.ndjson:

    23:49:59 desktop.window.createMain                   (window reopened; it had been closed since 23:13)
    23:50:01 PreviewAutomationBroker.connect             ok   <- the only connect in the whole log window
    23:50:36 preview_resize freeform 1440x900            PreviewAutomationTimeoutError after 15000ms
    23:50:51 PreviewAutomationBroker.disconnect          ok   (awaitResponse -> disconnect on timeout)
    23:50:51 ws.rpc.previewAutomation.connect            Interrupted after 50718ms
    23:50:51.. every invoke                              PreviewAutomationNoAvailableHostError
    

    The renderer never opened a new previewAutomation.connect afterwards; the host only came back after a window reload (Page.reload on the main renderer, or View > Reload), and then the same resize timeout evicted it again, reproducibly.

    The timeout itself is caused by the main window being zoomed (View > Zoom In): the guest webview stays at zoom factor 1 and reports innerWidth = requested * hostZoom, so the resize wait can never satisfy its ±1 px check. I filed that separately with the full analysis and a verified workaround: #12319. Even with that fixed, the missing re-registration here turns any single slow response into a dead session, so the two are worth fixing independently.

  8. m13v commented on Sep 17, 2026

    @m13v

    patching the 500ms icon-status call only closes the cosmetic path. your own trace shows the eviction came from a 15s evaluate deadline, not that probe, and disconnect-on-deadline is shared across both. drop the auto-status and a heavy evaluate still shuts the lone host's queue. the fix is the deadline handler not treating one unanswered call as 'host is dead' when it's the only host.

  9. m13v commented on Sep 17, 2026

    @m13v

    careful how #12319 gets fixed: relax the ±1px resize check to swallow innerWidth=requested*zoom and preview_resize returns ok against a webview still at zoom 1. the loud retryable timeout becomes every later click and snapshot silently off by the zoom factor.

  10. tejasc-03 commented on Sep 18, 2026

    @tejasc-03

    Confirming the same host-lifetime bug on macOS with T3 Code (Alpha) 0.0.42, Cursor ACP (Grok), local environment.

    What we saw

    The collaborative Browser tab stayed visible (URL bar, Recently used, even about:blank after a successful preview_open). Every later preview_status / preview_open from any Cursor chat then failed with:

    PreviewAutomationNoAvailableHostError: No preview automation host is available for status in environment <id>.
    

    Opening another preview tab did not recover it. View → Reload / ⌘R in the Browser tab did not recover it. Fully Quit T3 Code and reopen restored preview_status.available === true for a while, then the host dropped again (same session pattern as the original report).

    Traces (~/.t3/userdata/logs/server.trace.ndjson)

    Matches the triage smoking gun:

    • No new PreviewAutomationBroker.connect / acquireConnection after the drop
    • Repeated successful PreviewAutomationBroker.focusHost / ws.rpc.previewAutomation.focusHost during the outage (these no-op when clients is empty)
    • Every PreviewAutomationBroker.invoke is PreviewAutomationNoAvailableHostError
    • Desktop PreviewManager still registers webviews and paints the pane

    This is not #6355: preview_status itself fails, so it is environment-wide host loss, not a poisoned tab.

    Extra context

    Several Cursor ACP sessions were live on one desktop host. Human input in the preview pane (PreviewAutomationControlInterruptedError) showed up before some drops, but the durable failure is the missing re-register after the connect stream ends.

    #12343 looks like the right recovery (resubscribe after stream eviction). #12279 still looks necessary so the 500ms metadata status cannot evict the only host in the first place.

    Happy to attach a sanitized lifecycle excerpt from server.trace.ndjson if useful.

  11. m13v commented on Sep 18, 2026

    @m13v

    the catch with #12343 resubscribing after eviction: it reopens the stream but createPreviewAutomationClientId mints a fresh preview-, so the new host comes up blind to the evicted one's tab handles. preview_status flips back to available while preview_navigate/evaluate on tab_7 fails unknown-tab. you'd trade the clean permanent no-host error for calls that succeed against a host that can't see the tabs, unless the connect handshake rehydrates the old handles.

  12. juliusmarminge commented on Sep 19, 2026

    @juliusmarminge
    Member

    Fixed by #12535 — timeout evictions now complete the registration stream cleanly and the preview client re-registers with a fresh connection ID, so later calls no longer stay stuck on “No preview automation host is available” until remount.

  13. m13v commented on Sep 21, 2026

    @m13v

    'fresh connection ID' is the tell: the re-registered host inherits none of the old tab handles. preview_status flips to available so the session reads recovered, but preview_navigate on your slack tab now succeeds-blind instead of erroring. the loud no-host error just went silent.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions