Repository navigation
[Bug]: Service update leaves T3 Connect relay targeting stale origin port #7458
Description
Activity
I apologize for the lack of human input on this, I have no idea where to even start with this issue
Reacted by Sebastian ZelonkaThis reproduced again roughly 80 minutes after the first repair, and it also affected a separate desktop-managed macOS environment at the same time.
New evidence:
- The Linux/Coder T3 service remained active and healthy on
127.0.0.1:3773. - Its managed Cloudflared process remained active, but the relay configuration had changed to another dead ephemeral origin (
127.0.0.1:37337). - Restarting only
t3code.servicecausedCloudManagedEndpointRuntimeto create a new tunnel/connector and restored the public relay health endpoint to HTTP 200. - Separately, the macOS 0.0.33 desktop backend was healthy on local port 3773, but its T3 Connect endpoint returned the same
endpoint_request_failederror after the desktop app and connector had been running continuously for about two days. - Restarting T3 Code desktop triggered
CloudManagedEndpointRuntime.reconcileConfig, replaced the managed connector, and restored that relay health endpoint to HTTP 200 as well.
This recurrence suggests the issue is not limited to a partially upgraded systemd service. It appears that relay configuration can drift to a dead ephemeral origin while the actual backend and connector processes remain healthy. No environment IDs, relay hostnames, tunnel IDs, or credentials are included here.
- The Linux/Coder T3 service remained active and healthy on
I'm having this same issue, it happens very frequently and have to restart the service to recover. There is some config difference between the server port and the remotely managed tunnel. For me, it was 32913.
Reproduced on a desktop-managed macOS environment, including after updating and restarting into T3 Code Nightly
0.0.36-nightly.20260827.1205.This case differs slightly from the stale local-origin-port case in the original report: the active managed tunnel is healthy and demonstrably reaches the current backend on port 3773, but relay status/connect still returns
endpoint_request_failedwithout a request arriving at the tunnel.Reproduction / recovery attempts
- Confirmed only one embedded T3 server and one managed
cloudflaredprocess were running; removed a duplicate backend/tunnel created during troubleshooting. - Updated from Nightly
0.0.34through0.0.36-nightly.20260827.1205. - Performed a complete unlink rather than only disabling the tunnel:
- disabled Publish agent activity;
- disabled T3 Connect;
- verified
npx t3 connect status --jsonreturnedlinked: false,cloudUserId: null,relayUrl: null, and publishing off.
- Re-enabled T3 Connect, producing a fresh managed tunnel, then restored activity publishing.
- Fully quit/restarted T3 Code and retried with a freshly reloaded relay-environment descriptor.
The failure persists:
Failed to connect. Reconnecting... Reason: Relay environment endpoint is unavailable: endpoint_request_failedDiagnostics
- T3 backend is listening on the expected port
3773. - Managed
cloudflared2026.5.2has four active HA connections. cloudflared_tunnel_request_errorsis0.- A direct HTTPS request to the managed hostname reaches the current backend and returns HTTP 200.
- A direct unauthenticated POST to
/api/t3-connect/healthreaches the backend and returns the expected HTTP 400, proving the public hostname routes to the live origin. - After the fresh link/restart, relay status/connect failures do not increment
cloudflared_tunnel_total_requestsand do not appear in the server trace. This suggests failure/stale state before the relay request reaches the managed endpoint, rather than an origin-port mismatch. - The same account's mobile client initially showed
invalid_dpop; the complete unlink/relink should have replaced that stale credential, but desktop relay connectivity remains blocked byendpoint_request_failed.
Latest trace ID from Nightly 1205:
32aecf918ca826b6378bf8dd4036fae9Earlier trace IDs from the same investigation:
a83519abcd17650f591f6da5a1b86b5c d5b14aefa9a3f3d6bd09603782221724No environment IDs, managed hostnames, tunnel IDs, account IDs, or credentials are included here. This was collected through
npx t3 triage --agent codexplus direct local checks.- Confirmed only one embedded T3 server and one managed
Reproduced on a plain single-machine Linux setup — no VM/Coder hop and no desktop app involved, so the drift isn't limited to the multi-host update scenarios above.
Environment
- T3 Code
0.0.38(stable), user-level systemd service (t3code.service,Restart=always) - Linux x64 (Arch, kernel
7.2.3-arch1-2), Nodev26.8.1 - Launch method: plain
npx t3in an interactive terminal (which also manages the background service)
What happened
The background service crash-looped for several restart cycles after a fresh
npx t3install (unrelated root cause — the node-pty install-script issue tracked in #7475), then finally came up and stayed running. That surviving process landed on a different local port than the last time the environment had linked with the relay (the CLI's own port picker falls back to a random available port whenever the default3773is occupied, which it was during the crash-loop churn).From then on:
- The T3 server was demonstrably alive and listening on its new port.
boot-service.logshowed the startup path logging "T3 Connect desired link reconciled on startup" with no error — i.e. the reconcile call it makes on every boot did run and did not fail.- Despite that,
cloudflaredkept loggingUnable to reach the origin service ... dial tcp 127.0.0.1:<old-port>: connect: connection refusedagainst the previous session's port, indefinitely (observed for 20+ minutes with zero recovery, until we intervened). - A remote client (a second T3 Code instance connecting through T3 Connect) saw exactly
Relay could not reach the environment endpoint (endpoint_request_failed)the whole time.
Fix that worked:
systemctl --user restart t3code.service. The restart spawned a brand-new relay-client/tunnel connector (new tunnel ID observed in the logs) which immediately registered against the current port with no further warnings. Same workaround already documented above for the VM/desktop cases.Possible mechanism (from reading
apps/server/src/cloud/ManagedEndpointRuntime.tsandapps/server/src/cloud/http.tson this version): the per-boot reconcile (reconcileDesiredCloudLink) does re-provision the relay side with the current origin, butCloudManagedEndpointRuntime.reconcileConfig's dedup key (runtimeConfigKey) only hashesproviderKind+connectorToken+tunnelId+tunnelName— it never includes the origin host/port. So when the same tunnel identity is reused across a reconcile, the already-running localcloudflared tunnel runprocess is left untouched even if the origin it should be forwarding to has changed, and there's nothing else here that forces it to reconnect or verifies it actually converged on the new target. That would explain why a full service restart (which tears down and respawns the connector unconditionally) reliably fixes it, while the routine startup-time reconcile does not.Diagnosed via
t3 triage(Claude Sonnet 5).- T3 Code
Reproduced on macOS (launchd service), 0.0.44, with a deterministic trigger
Environment: T3 Code 0.0.44 (CLI, service runtime and desktop app), macOS arm64 (Darwin 27.0.0), Node v26.8.2, cloudflared 2026.5.2 (managed install). Background service via launchd (
t3 serveon127.0.0.1:3773), desktop app with local environment disabled, linked to T3 Connect with a managed tunnel.Repro:
- Have the background service running and linked to T3 Connect (tunnel origin
127.0.0.1:3773, remote clients work). - In a terminal, run bare
t3as the same user with the defaultT3CODE_HOME. It starts a second web-mode server on an ephemeral port against the same state dir. - Wait ~10 s, then Ctrl-C it.
- Remote clients now show
T3 Connect · Reconnecting: Relay could not reach the environment endpoint (endpoint_request_failed).indefinitely.
Diagnosis: the ad-hoc server shares the service's secrets and environment link, so its startup path (
registerManagedCloudTunnelRecovery(localOrigin),apps/server/src/server.ts:868) asks the relay to reconcile the tunnel ingress to its own ephemeral port. When it exits, the ingress stays on that port. The service only registers its origin at startup, so it never corrects it, and its cloudflared child keeps forwarding to the dead port.Evidence (UTC, 2026-10-01):
- 09:58:14 service:
T3 Connect managed tunnel recovery registered(port 3773) - 09:58:14–09:58:48 trace spans with
server.mode=web, server.port=51865;environment.cloud.registerManagedCloudTunnelRecoverysucceeds at 09:58:24 - 09:58:48
releaseManagedTunnelOnShutdownfails:upstream_unavailable, trace ID20b251b9f693c559ae637a4922753a65 - 09:58:54–10:03:35 service's cloudflared:
Unable to reach the origin service … dial tcp 127.0.0.1:51865: connect: connection refused … originService=http://127.0.0.1:51865 - 10:03:08–10:03:37 second bare
t3on port 52657;registerManagedCloudTunnelRecoveryfails withrelay returned HTTP 504(Ray IDa43aa3258f08fd81-SIN), yet from 10:03:39 the ingress isoriginService=http://127.0.0.1:52657, so the origin change was applied despite the 504 lsof: onlyt3 servelistens, on127.0.0.1:3773;t3 connect statusshows exposure enabled and link provisioned
Workaround:
t3 service restartre-registers 3773. Verified here: after the restart at 10:16:16 the service loggedT3 Connect managed tunnel recovery registeredat 10:17:24 and cloudflared reported no further origin errors.Suggestions: an ad-hoc
t3should not re-register the managed tunnel origin when a service owns the T3 home (overlaps #6097), and/or the service should re-reconcile its origin when its connector reports repeatedconnection refusedto a port that is not its own.Produced via
t3 triageby Claude Opus 5.5 (claude-opus-5-5) in Claude Code.- Have the background service running and linked to T3 Connect (tunnel origin
Reproduced on Linux (Fedora 44, x64) with the systemd service on
0.0.46-nightly.20261008.2849. In my case the stale port has a concrete trigger: a second T3 server launched over SSH by the macOS desktop app.What happens
- The
t3code.serviceserver listens on127.0.0.1:3773and owns the T3 Connect managed tunnel. - The macOS desktop app connects to this host over SSH and starts its own server via
~/.t3/ssh-launch/<id>/run-t3.sh. It shares the same T3 home, so it also shares the same T3 Connect link and tunnel. - 3773 is busy, so
findAvailablePort(apps/server/src/cli/config.ts) gives the SSH-launched server 3774. On startup it callsregisterManagedTunnelRecoverywith origin127.0.0.1:3774, and the relay switches the tunnel's ingress to 3774. - The SSH-launched server exits when the Mac disconnects. Nothing points the relay back at 3773. The service server only re-registers its origin at its own startup.
- From then on, cloudflared (still running under the 3773 server) logs
dial tcp 127.0.0.1:3774: connect: connection refusedfor every request, and mobile showsRelay could not reach the environment endpoint (endpoint_request_failed).
Evidence
- The service log shows the 3774 ingress errors on 18 separate days between Sep 6 and Oct 11, matching the times the SSH-launched server was started.
- The SSH-launched server's
server.logshowsListening on http://127.0.0.1:3774followed byT3 Connect managed tunnel recovery registered(last at 23:17, it stopped at 23:24). The phone failed from then on. - Requests reached 3773 successfully in the window where the service server had registered last. So the relay ingress flips to whichever server registered most recently.
systemctl --user restart t3code.servicefixes it right away: the public/api/t3-connect/healthreaches the backend again (HTTP 400 for an unauthenticated probe) and the 3774 errors stop.
Repro
- Run T3 as a service on Linux host A (port 3773) with T3 Connect enabled.
- From the macOS desktop app, add host A as an SSH environment, so it launches a second server on 3774.
- Close the desktop app or let the Mac sleep, so the SSH server exits.
- Open host A from the mobile app through T3 Connect:
endpoint_request_failed.
Possible directions: the SSH-launched server shouldn't run the managed tunnel or register an origin when another server on the host already owns the link. Or the tunnel owner could re-assert its origin when cloudflared reports origin connection failures.
Mobile trace IDs:
f68867e720322ef82c798d215ec5567e,e8e8d71d4868e441ae24c8e1f0937fdeDiagnosed and written by Claude Opus 5.5 (Claude Code) via
t3 triage.- The
Before submitting
Area
apps/server
Steps to reproduce
t3 service status, listeners, andboot-service.log.The service reported
needs an update or repair. The active T3 server listened on127.0.0.1:3773, while the existing managed Cloudflare tunnel continued forwarding to an old ephemeral origin port (127.0.0.1:35765).Expected behavior
A managed runtime/service update should atomically restart or reconcile the background service and relay connector. The tunnel origin should always point to the currently active T3 server port, and saved clients should reconnect without manual operator intervention.
Actual behavior
The desktop app could not connect and repeatedly displayed:
The relay itself was provisioned and
t3 connect statusshowed exposure enabled, but Cloudflared repeatedly failed to reach its stale origin. The server logs also showed invalid saved-session/bootstrap credentials during reconnect attempts.Impact
Blocks work completely
Version or commit
Desktop app 0.0.33; managed VM runtime/service 0.0.33; shell CLI 0.0.32 before repair
Environment
macOS desktop client connecting to a Linux Coder VM over T3 Connect; Node v24.19.0; user-level systemd service
Logs or stack traces
Workaround
Run
npx t3@latest service updatewith access to the user systemd bus, then explicitly restartt3code.service. After restart, T3 0.0.33 listened on port 3773, Cloudflared was relaunched against the live origin, and an outside-in relay health request returned HTTP 200.One additional wrinkle: from a non-interactive Coder SSH shell, the first service repair could not access the user systemd bus (
Failed to connect to bus: No medium found). SettingXDG_RUNTIME_DIR=/run/user/$(id -u)andDBUS_SESSION_BUS_ADDRESS=unix:path=$XDG_RUNTIME_DIR/busallowed the service command to inspect the installed unit. This may be worth handling or surfacing explicitly.