Skip to content

[Bug]: Service update leaves T3 Connect relay targeting stale origin port #7458

Description

@Dawsson

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Run T3 Code as the installed user systemd service on a remote Linux VM and expose it through T3 Connect.
  2. Upgrade the desktop app and allow the managed VM runtime/service to update. In this case, the macOS desktop app was 0.0.33, the shell CLI remained 0.0.32, and the managed runtime had installed 0.0.33.
  3. Attempt to reconnect to the remote environment from the desktop app.
  4. Inspect t3 service status, listeners, and boot-service.log.

The service reported needs an update or repair. The active T3 server listened on 127.0.0.1:3773, while the existing managed Cloudflare tunnel continued forwarding to an old ephemeral origin port (127.0.0.1:35765).

Expected behavior

A managed runtime/service update should atomically restart or reconcile the background service and relay connector. The tunnel origin should always point to the currently active T3 server port, and saved clients should reconnect without manual operator intervention.

Actual behavior

The desktop app could not connect and repeatedly displayed:

Failed to connect. Reconnecting...
Reason: Relay environment endpoint is unavailable: endpoint_request_failed

The relay itself was provisioned and t3 connect status showed exposure enabled, but Cloudflared repeatedly failed to reach its stale origin. The server logs also showed invalid saved-session/bootstrap credentials during reconnect attempts.

Impact

Blocks work completely

Version or commit

Desktop app 0.0.33; managed VM runtime/service 0.0.33; shell CLI 0.0.32 before repair

Environment

macOS desktop client connecting to a Linux Coder VM over T3 Connect; Node v24.19.0; user-level systemd service

Logs or stack traces

$ t3 service status
T3 Code service
  Status: needs an update or repair
  Next: Run `npx t3@latest service update`.

$ ss -lntp
LISTEN 127.0.0.1:3773  # active T3 server

cloudflared: Unable to reach the origin service: dial tcp 127.0.0.1:35765: connect: connection refused

Workaround

Run npx t3@latest service update with access to the user systemd bus, then explicitly restart t3code.service. After restart, T3 0.0.33 listened on port 3773, Cloudflared was relaunched against the live origin, and an outside-in relay health request returned HTTP 200.

One additional wrinkle: from a non-interactive Coder SSH shell, the first service repair could not access the user systemd bus (Failed to connect to bus: No medium found). Setting XDG_RUNTIME_DIR=/run/user/$(id -u) and DBUS_SESSION_BUS_ADDRESS=unix:path=$XDG_RUNTIME_DIR/bus allowed the service command to inspect the installed unit. This may be worth handling or surfacing explicitly.

Activity

  1. Dawsson commented on Aug 19, 2026

    @Dawsson
    Author

    I apologize for the lack of human input on this, I have no idea where to even start with this issue

  2. Dawsson commented on Aug 19, 2026

    @Dawsson
    Author

    This reproduced again roughly 80 minutes after the first repair, and it also affected a separate desktop-managed macOS environment at the same time.

    New evidence:

    • The Linux/Coder T3 service remained active and healthy on 127.0.0.1:3773.
    • Its managed Cloudflared process remained active, but the relay configuration had changed to another dead ephemeral origin (127.0.0.1:37337).
    • Restarting only t3code.service caused CloudManagedEndpointRuntime to create a new tunnel/connector and restored the public relay health endpoint to HTTP 200.
    • Separately, the macOS 0.0.33 desktop backend was healthy on local port 3773, but its T3 Connect endpoint returned the same endpoint_request_failed error after the desktop app and connector had been running continuously for about two days.
    • Restarting T3 Code desktop triggered CloudManagedEndpointRuntime.reconcileConfig, replaced the managed connector, and restored that relay health endpoint to HTTP 200 as well.

    This recurrence suggests the issue is not limited to a partially upgraded systemd service. It appears that relay configuration can drift to a dead ephemeral origin while the actual backend and connector processes remain healthy. No environment IDs, relay hostnames, tunnel IDs, or credentials are included here.

  3. vidhunv1 commented on Aug 24, 2026

    @vidhunv1

    I'm having this same issue, it happens very frequently and have to restart the service to recover. There is some config difference between the server port and the remotely managed tunnel. For me, it was 32913.

  4. tayyabfareed commented on Aug 27, 2026

    @tayyabfareed

    Reproduced on a desktop-managed macOS environment, including after updating and restarting into T3 Code Nightly 0.0.36-nightly.20260827.1205.

    This case differs slightly from the stale local-origin-port case in the original report: the active managed tunnel is healthy and demonstrably reaches the current backend on port 3773, but relay status/connect still returns endpoint_request_failed without a request arriving at the tunnel.

    Reproduction / recovery attempts

    1. Confirmed only one embedded T3 server and one managed cloudflared process were running; removed a duplicate backend/tunnel created during troubleshooting.
    2. Updated from Nightly 0.0.34 through 0.0.36-nightly.20260827.1205.
    3. Performed a complete unlink rather than only disabling the tunnel:
      • disabled Publish agent activity;
      • disabled T3 Connect;
      • verified npx t3 connect status --json returned linked: false, cloudUserId: null, relayUrl: null, and publishing off.
    4. Re-enabled T3 Connect, producing a fresh managed tunnel, then restored activity publishing.
    5. Fully quit/restarted T3 Code and retried with a freshly reloaded relay-environment descriptor.

    The failure persists:

    Failed to connect. Reconnecting...
    Reason: Relay environment endpoint is unavailable: endpoint_request_failed
    

    Diagnostics

    • T3 backend is listening on the expected port 3773.
    • Managed cloudflared 2026.5.2 has four active HA connections.
    • cloudflared_tunnel_request_errors is 0.
    • A direct HTTPS request to the managed hostname reaches the current backend and returns HTTP 200.
    • A direct unauthenticated POST to /api/t3-connect/health reaches the backend and returns the expected HTTP 400, proving the public hostname routes to the live origin.
    • After the fresh link/restart, relay status/connect failures do not increment cloudflared_tunnel_total_requests and do not appear in the server trace. This suggests failure/stale state before the relay request reaches the managed endpoint, rather than an origin-port mismatch.
    • The same account's mobile client initially showed invalid_dpop; the complete unlink/relink should have replaced that stale credential, but desktop relay connectivity remains blocked by endpoint_request_failed.

    Latest trace ID from Nightly 1205:

    32aecf918ca826b6378bf8dd4036fae9
    

    Earlier trace IDs from the same investigation:

    a83519abcd17650f591f6da5a1b86b5c
    d5b14aefa9a3f3d6bd09603782221724
    

    No environment IDs, managed hostnames, tunnel IDs, account IDs, or credentials are included here. This was collected through npx t3 triage --agent codex plus direct local checks.

  5. sheepbun commented on Sep 6, 2026

    @sheepbun

    Reproduced on a plain single-machine Linux setup — no VM/Coder hop and no desktop app involved, so the drift isn't limited to the multi-host update scenarios above.

    Environment

    • T3 Code 0.0.38 (stable), user-level systemd service (t3code.service, Restart=always)
    • Linux x64 (Arch, kernel 7.2.3-arch1-2), Node v26.8.1
    • Launch method: plain npx t3 in an interactive terminal (which also manages the background service)

    What happened

    The background service crash-looped for several restart cycles after a fresh npx t3 install (unrelated root cause — the node-pty install-script issue tracked in #7475), then finally came up and stayed running. That surviving process landed on a different local port than the last time the environment had linked with the relay (the CLI's own port picker falls back to a random available port whenever the default 3773 is occupied, which it was during the crash-loop churn).

    From then on:

    • The T3 server was demonstrably alive and listening on its new port.
    • boot-service.log showed the startup path logging "T3 Connect desired link reconciled on startup" with no error — i.e. the reconcile call it makes on every boot did run and did not fail.
    • Despite that, cloudflared kept logging Unable to reach the origin service ... dial tcp 127.0.0.1:<old-port>: connect: connection refused against the previous session's port, indefinitely (observed for 20+ minutes with zero recovery, until we intervened).
    • A remote client (a second T3 Code instance connecting through T3 Connect) saw exactly Relay could not reach the environment endpoint (endpoint_request_failed) the whole time.

    Fix that worked: systemctl --user restart t3code.service. The restart spawned a brand-new relay-client/tunnel connector (new tunnel ID observed in the logs) which immediately registered against the current port with no further warnings. Same workaround already documented above for the VM/desktop cases.

    Possible mechanism (from reading apps/server/src/cloud/ManagedEndpointRuntime.ts and apps/server/src/cloud/http.ts on this version): the per-boot reconcile (reconcileDesiredCloudLink) does re-provision the relay side with the current origin, but CloudManagedEndpointRuntime.reconcileConfig's dedup key (runtimeConfigKey) only hashes providerKind + connectorToken + tunnelId + tunnelName — it never includes the origin host/port. So when the same tunnel identity is reused across a reconcile, the already-running local cloudflared tunnel run process is left untouched even if the origin it should be forwarding to has changed, and there's nothing else here that forces it to reconnect or verifies it actually converged on the new target. That would explain why a full service restart (which tears down and respawns the connector unconditionally) reliably fixes it, while the routine startup-time reconcile does not.

    Diagnosed via t3 triage (Claude Sonnet 5).

  6. KrtinShet commented on Oct 1, 2026

    @KrtinShet

    Reproduced on macOS (launchd service), 0.0.44, with a deterministic trigger

    Environment: T3 Code 0.0.44 (CLI, service runtime and desktop app), macOS arm64 (Darwin 27.0.0), Node v26.8.2, cloudflared 2026.5.2 (managed install). Background service via launchd (t3 serve on 127.0.0.1:3773), desktop app with local environment disabled, linked to T3 Connect with a managed tunnel.

    Repro:

    1. Have the background service running and linked to T3 Connect (tunnel origin 127.0.0.1:3773, remote clients work).
    2. In a terminal, run bare t3 as the same user with the default T3CODE_HOME. It starts a second web-mode server on an ephemeral port against the same state dir.
    3. Wait ~10 s, then Ctrl-C it.
    4. Remote clients now show T3 Connect · Reconnecting: Relay could not reach the environment endpoint (endpoint_request_failed). indefinitely.

    Diagnosis: the ad-hoc server shares the service's secrets and environment link, so its startup path (registerManagedCloudTunnelRecovery(localOrigin), apps/server/src/server.ts:868) asks the relay to reconcile the tunnel ingress to its own ephemeral port. When it exits, the ingress stays on that port. The service only registers its origin at startup, so it never corrects it, and its cloudflared child keeps forwarding to the dead port.

    Evidence (UTC, 2026-10-01):

    • 09:58:14 service: T3 Connect managed tunnel recovery registered (port 3773)
    • 09:58:14–09:58:48 trace spans with server.mode=web, server.port=51865; environment.cloud.registerManagedCloudTunnelRecovery succeeds at 09:58:24
    • 09:58:48 releaseManagedTunnelOnShutdown fails: upstream_unavailable, trace ID 20b251b9f693c559ae637a4922753a65
    • 09:58:54–10:03:35 service's cloudflared: Unable to reach the origin service … dial tcp 127.0.0.1:51865: connect: connection refused … originService=http://127.0.0.1:51865
    • 10:03:08–10:03:37 second bare t3 on port 52657; registerManagedCloudTunnelRecovery fails with relay returned HTTP 504 (Ray ID a43aa3258f08fd81-SIN), yet from 10:03:39 the ingress is originService=http://127.0.0.1:52657, so the origin change was applied despite the 504
    • lsof: only t3 serve listens, on 127.0.0.1:3773; t3 connect status shows exposure enabled and link provisioned

    Workaround: t3 service restart re-registers 3773. Verified here: after the restart at 10:16:16 the service logged T3 Connect managed tunnel recovery registered at 10:17:24 and cloudflared reported no further origin errors.

    Suggestions: an ad-hoc t3 should not re-register the managed tunnel origin when a service owns the T3 home (overlaps #6097), and/or the service should re-reconcile its origin when its connector reports repeated connection refused to a port that is not its own.

    Produced via t3 triage by Claude Opus 5.5 (claude-opus-5-5) in Claude Code.

  7. DigitalWestern commented on Oct 11, 2026

    @DigitalWestern

    Reproduced on Linux (Fedora 44, x64) with the systemd service on 0.0.46-nightly.20261008.2849. In my case the stale port has a concrete trigger: a second T3 server launched over SSH by the macOS desktop app.

    What happens

    1. The t3code.service server listens on 127.0.0.1:3773 and owns the T3 Connect managed tunnel.
    2. The macOS desktop app connects to this host over SSH and starts its own server via ~/.t3/ssh-launch/<id>/run-t3.sh. It shares the same T3 home, so it also shares the same T3 Connect link and tunnel.
    3. 3773 is busy, so findAvailablePort (apps/server/src/cli/config.ts) gives the SSH-launched server 3774. On startup it calls registerManagedTunnelRecovery with origin 127.0.0.1:3774, and the relay switches the tunnel's ingress to 3774.
    4. The SSH-launched server exits when the Mac disconnects. Nothing points the relay back at 3773. The service server only re-registers its origin at its own startup.
    5. From then on, cloudflared (still running under the 3773 server) logs dial tcp 127.0.0.1:3774: connect: connection refused for every request, and mobile shows Relay could not reach the environment endpoint (endpoint_request_failed).

    Evidence

    • The service log shows the 3774 ingress errors on 18 separate days between Sep 6 and Oct 11, matching the times the SSH-launched server was started.
    • The SSH-launched server's server.log shows Listening on http://127.0.0.1:3774 followed by T3 Connect managed tunnel recovery registered (last at 23:17, it stopped at 23:24). The phone failed from then on.
    • Requests reached 3773 successfully in the window where the service server had registered last. So the relay ingress flips to whichever server registered most recently.
    • systemctl --user restart t3code.service fixes it right away: the public /api/t3-connect/health reaches the backend again (HTTP 400 for an unauthenticated probe) and the 3774 errors stop.

    Repro

    1. Run T3 as a service on Linux host A (port 3773) with T3 Connect enabled.
    2. From the macOS desktop app, add host A as an SSH environment, so it launches a second server on 3774.
    3. Close the desktop app or let the Mac sleep, so the SSH server exits.
    4. Open host A from the mobile app through T3 Connect: endpoint_request_failed.

    Possible directions: the SSH-launched server shouldn't run the managed tunnel or register an origin when another server on the host already owns the link. Or the tunnel owner could re-assert its origin when cloudflared reports origin connection failures.

    Mobile trace IDs: f68867e720322ef82c798d215ec5567e, e8e8d71d4868e441ae24c8e1f0937fde

    Diagnosed and written by Claude Opus 5.5 (Claude Code) via t3 triage.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions