Skip to content

[Bug]: SSH reconnect kills a busy remote server, cancelling all running agents (2s reuse probe) #16477

Description

@bloom-street

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/desktop (SSH environment launch script), with apps/server startup recovery

Steps to reproduce

  1. On laptop A, run T3 Code desktop with an SSH environment pointing at machine B. The launch script starts a managed server on B (~/.t3/ssh-launch/<key>/, managed).
  2. Start several long agent runs on B, with background shells and subagents, and leave B under heavy CPU load (load average about 2-3x core count from test suites).
  3. Close the lid on laptop A, on battery. macOS DarkWakes it about every 15 minutes, and the desktop app reconnects the SSH environment on each wake.
  4. Each reconnect reruns REMOTE_LAUNCH_SCRIPT on B.

Expected behavior

A live, healthy server that is only slow to answer is reused. Running agents and their background work survive a client reconnect.

Actual behavior

On each reconnect the reuse probe (REMOTE_REUSE_READY_TIMEOUT_MS = 2e3, SSH_READY_PROBE_TIMEOUT_MS = 1e3) does not get an answer in time on the loaded host. The script hits elif ! wait_ready ...; then kill "$REMOTE_PID" and starts a new server. The new server's V2 orchestration recovery then stops every provider session (stoppedSessions: 8) and emits run.background-work-cancelled for every running background shell and subagent.

In one night this happened 29 times, about every 15 minutes from 00:34 to 08:29. Each server restart was about 10-15 s after a laptop DarkWake, for example:

Laptop DarkWake Server log Relay client stopped then Listening on
00:33:44 00:33:57 → 00:34:06
00:46:23 00:46:33 → 00:46:40
01:02:08 01:02:20 → 01:02:24
01:21:42 01:21:59 → 01:22:03
01:37:27 01:37:37 → 01:37:42

No agent work survived the night. The server never crashed; every shutdown was a clean SIGTERM.

Related cases in the same script:

  • RUNNER_CHANGED=1, when a client on a different nightly reconnects after an auto-update, also kills a busy managed server and cancels all of its work without warning.
  • If server-runtime.json points at another live server (for example the desktop app's own server on B), the next reconnect kills the managed server even though it is healthy and running agents.

Impact

Major degradation or frequent failure

Version or commit

0.0.46-nightly.20261004.2657 (overnight), then 0.0.46-nightly.20261005.2702. macOS on both ends.

Suggested fixes

  • Do not kill a managed server whose pid is alive just because a 2 s readiness probe failed. Retry with a longer budget (for example 30-60 s) or check liveness via the pid and server-runtime file before deciding.
  • On RUNNER_CHANGED with active runs, defer the restart or ask the user, rather than SIGTERMing immediately.
  • Consider making a client reconnect never restart the server at all. Restarts only from explicit user action or update.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions