Skip to content

SSH remote: client kills its own managed server on every reconnect (default-runtime adoption check) #15986

Description

@understory-finch

Summary

When the desktop app connects to a remote host over SSH, every reconnect after the first SIGTERMs the healthy server that the client itself launched, then starts a new one. Every provider session (Claude, etc.) on the remote dies on each reconnect.

On my Mac host this caused 22 server restarts in ~18h on v0.0.45. Each one followed a tunnel reconnect (laptop sleep, Wi-Fi blip, app restart).

Root cause

This is in REMOTE_LAUNCH_SCRIPT in packages/ssh/src/tunnel.ts (lines from v0.0.45; the logic is unchanged on main @ cf3e714):

  1. The client starts the managed server with --base-dir "$HOME/.t3" (L705) and writes managed to $MANAGED_FILE.
  2. Because of that base dir, the managed server writes its own pid and port to ~/.t3/userdata/server-runtime.json. That is the same file the script treats as the "default runtime" (DEFAULT_RUNTIME_FILE, L550).
  3. On the next launch, resolve_default_runtime_port finds that file. The pid is alive and the origin is localhost, so it passes.
  4. The port answers wait_ready, and REMOTE_MANAGED = managed (L646). The script assumes an external service has appeared and kills PID_TO_STOP, then marks the state external. But that pid is the client's own managed server, the same pid as in $PID_FILE.
  5. The external branch re-probes the port. It just killed that server, so the probe fails, the state is cleared, and a fresh server is spawned (L695-709).

tunnel.test.ts (around L300-312) asserts the kill-then-adopt ordering. It does not cover the case where the runtime file belongs to the managed server itself.

Evidence (remote host)

  • ~/.t3/ssh-launch/<hash>/pid == server-runtime.json.pid, and managed = managed.
  • server.log shows a new Listening on http://127.0.0.1:3773 22 times with no error before any of them.
  • server.trace.ndjson: ProviderService finalizer (runStopAll → stopSessions) at 10:16:58, last span 10:17:00. ssh-launch files rewritten and new server up at 10:17:07. This is a clean SIGTERM followed by an immediate relaunch.

Suggested fix

Only adopt or kill when the default runtime is not the managed server. For example, skip the kill branch when DEFAULT_RUNTIME_PID = REMOTE_PID and treat it as reuse. Alternatively, have the server record whether it is service-managed in server-runtime.json and check that.

Workaround

Run t3 service install on the remote host and remove ~/.t3/ssh-launch/<hash>/{pid,managed}. The client then adopts the service as external and never kills it.

Environment: desktop client → remote macOS (Darwin 24.5, arm64) over Tailscale SSH, t3 0.0.45 release archive.

Activity

  1. caitlon commented on Oct 6, 2026

    @caitlon

    Same here on v0.0.45, macOS host reached from a MacBook over SSH: 25 server restarts between 2026-10-05 16:08 and 2026-10-06 08:30 UTC. Each one starts within 10 s of an incoming SSH login from the laptop, mostly sleep/wake overnight. On the host, ~/.t3/ssh-launch/<key>/managed reads managed and pid matches the pid in ~/.t3/userdata/server-runtime.json, so the default-runtime check adopts and kills the client's own server, exactly as described above.

    Two consequences not mentioned yet:

    • Background subagents die silently. Every restart stops all Claude sessions (Session stopped, graceful), and the next turn resumes with "Background agent … didn't finish before the previous session ended". In one evening that killed five running subagents (a builder, a code reviewer, and the same review agent twice), each mid-task, with nothing in the UI at the time.
    • Other clients lose the server too. A phone connected to the same host is cut off on every laptop reconnect. Once, at 04:12 UTC, the old server was killed and no new one started until 04:23, so the phone had no server for about ten minutes.

    The phone itself never caused a restart: none of its 37 turns changed the model or effort, and switching devices mid-session without a laptop reconnect left the session alive. So the trigger is purely the laptop's SSH launch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions