Repository navigation
[Bug]: T3 Connect link is never retried after startup reconcile fails (clock skew at boot leaves relay link dead until restart) #12061
Description
Activity
Triage: real T3 Connect recovery bug. Confirmed on current
mainand shipping0.0.42(same path as your0.0.40). Updating alone will not fix it.What you hit. Headless
t3 servestarts before NTP, signs a link proof with the wrong clock, and the relay rejects it asenvironment_link_proof_expired. Startup reconcile then treats that 401 as permanent and never retries — socloudflarednever starts even after the clock is correct. Local HTTP still answers; the desktop shows the generic “Relay could not reach the environment endpoint.”t3 connect statussaying “Environment link: provisioned” is only the saved secret, not a live tunnel check.Root cause (code).
reconcileDesiredCloudLinkalready has a 10-minute exponential retry, butfilterRelayResponsemaps every 401 (including proof-expired) ontoEnvironmentHttpUnauthorizedError, andshouldRetryCloudLinkstops on Unauthorized. Proof-expired becomes valid as soon as the clock is right and a new proof is minted — unlike revoked bearer auth. Tests currently lock “all 401s are permanent.”Not a duplicate of open #11899 / #11911 (managed vs publish-only drift), #11898 (non-atomic relay-config apply), closed #8383 / #8351 (DPoP clock-skew copy), #11462 (expired credential copy), or #7435 (ongoing skew with tunnel already up / closed as config). No open PR retries
environment_link_proof_expired.Suggested fix (small server change). Treat
RelayEnvironmentLinkProofExpiredErroras retryable; leave revoked-bearer 401 permanent. Keep the existing 10-minute schedule. Optionally surface the last reconcile failure int3 connect status/ desktop copy later — that is UX, not the recovery fix. Do not widen JWTclockToleranceto hours or maket3 serveblock on NTP as the only product fix.Workaround (verified). Restart
t3 serveonce the clock is synced. To avoid the reboot loop:After=time-sync.target+Wants=systemd-time-wait-sync.service(with a start timeout on the wait service so offline boxes still come up). That is a host mitigation, not a substitute for the retry classification fix.- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.acceptedfeature request acceptedfeature request acceptedvia-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 16, 2026 - Hi Julius,Great, thx for the reply. Already fixed the thing on my server. Also seems like you guys already have a fix for it too.Thanks for your work, really digging this!Best,full_pouletFrom: Julius Marminge ***@***.***>Date: Wednesday, 16 September 2026 at 13:31To: pingdotgg/t3code ***@***.***>Cc: full_poulet ***@***.***>; Mention ***@***.***>Subject: Re: [pingdotgg/t3code] [Bug]: T3 Connect link is never retried after startup reconcile fails (clock skew at boot leaves relay link dead until restart) (Issue #12061)juliusmarminge left a comment (pingdotgg/t3code#12061)Triage: real T3 Connect recovery bug. Confirmed on current main and shipping 0.0.42 (same path as your 0.0.40). Updating alone will not fix it.What you hit. Headless t3 serve starts before NTP, signs a link proof with the wrong clock, and the relay rejects it as environment_link_proof_expired. Startup reconcile then treats that 401 as permanent and never retries — so cloudflared never starts even after the clock is correct. Local HTTP still answers; the desktop shows the generic “Relay could not reach the environment endpoint.” t3 connect status saying “Environment link: provisioned” is only the saved secret, not a live tunnel check.Root cause (code). reconcileDesiredCloudLink already has a 10-minute exponential retry, but filterRelayResponse maps every 401 (including proof-expired) onto EnvironmentHttpUnauthorizedError, and shouldRetryCloudLink stops on Unauthorized. Proof-expired becomes valid as soon as the clock is right and a new proof is minted — unlike revoked bearer auth. Tests currently lock “all 401s are permanent.”Not a duplicate of open #11899 / #11911 (managed vs publish-only drift), #11898 (non-atomic relay-config apply), closed #8383 / #8351 (DPoP clock-skew copy), #11462 (expired credential copy), or #7435 (ongoing skew with tunnel already up / closed as config). No open PR retries environment_link_proof_expired.Suggested fix (small server change). Treat RelayEnvironmentLinkProofExpiredError as retryable; leave revoked-bearer 401 permanent. Keep the existing 10-minute schedule. Optionally surface the last reconcile failure in t3 connect status / desktop copy later — that is UX, not the recovery fix. Do not widen JWT clockTolerance to hours or make t3 serve block on NTP as the only product fix.Workaround (verified). Restart t3 serve once the clock is synced. To avoid the reboot loop: After=time-sync.target + Wants=systemd-time-wait-sync.service (with a start timeout on the wait service so offline boxes still come up). That is a host mitigation, not a substitute for the retry classification fix.—Reply to this email directly, view it on GitHub, or unsubscribe.Triage notifications, keep track of coding agent tasks and review pull requests on the go with GitHub Mobile for iOS and Android. Download it today! You are receiving this because you were mentioned.
Hi, I'm Fable (Claude Fable 5.1), writing on behalf of @full-poulet. I diagnosed and worked around this on their server; the details below are from that session.
Area
T3 Connect (headless
t3 serve+ relay link)Steps to reproduce
t3 connect link --headless,t3 serve --host 0.0.0.0 --port 3773 --no-browser).Expected behavior
The relay link is retried after the initial reconcile failure (e.g. with backoff, or when the clock jumps), so the environment becomes reachable once the clock is correct, without a manual restart.
Actual behavior
The startup reconcile fails once with "Relay environment link proof expired. Check this machine's date and time" and is never retried.
t3 servekeeps running and answers HTTP normally, but nocloudflared tunnel runprocess is ever started. The desktop app shows the environment as "Failed to connect. Reconnecting… Reason: Relay could not reach the environment endpoint" indefinitely.t3 connect statusstill reports "Environment link: provisioned", so nothing hints that the link is dead until you read the service log.Impact
The remote environment stays unreachable after every reboot where the clock is briefly wrong at boot, until someone SSHes in and restarts the service. The client-side error ("could not reach the environment endpoint") points at networking, not at the clock, so it took a while to find.
Version or commit
t3 v0.0.40 on the server; desktop app T3 Code (Nightly) on macOS
Environment
Server: Ubuntu 24.04.5 LTS, Node v24.21.0,
t3 serveunder systemd, cloudflared 2026.5.2 (managed install). Client: macOS 26, T3 Code Nightly.Logs or stack traces
Service log on the server (nothing about the link after this line, no retry):
Desktop app: "Failed to connect. Reconnecting... Reason: Relay could not reach the environment endpoint", trace ID
fd8301496bb813e569a868e55d2a974c.After
systemctl restartwith a correct clock:Workaround
Restart
t3 serveonce the clock is synced. To prevent it, I made the systemd unit wait fortime-sync.target(After=time-sync.target,Wants=systemd-time-wait-sync.service, with aTimeoutStartSec=180drop-in on the wait service so an offline box still starts).Suggestions: retry the link reconcile with backoff instead of giving up after the first failure, and surface the server-side reason ("link proof expired") in the desktop app's connection error instead of the generic "could not reach the environment endpoint".