Skip to content

[Bug]: T3 Connect stays down after a network change: cloudflared keeps running with zero connections and is never restarted #16258

Description

@leeonfield

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. On macOS, run a full-tunnel VPN client whose DNS answers with session-scoped virtual IPs. Here it is a corporate ZTNA/VPN client: while it is on, region1.v2.argotunnel.com resolves to virtual addresses from the client's own range that only route through its utun interface, instead of Cloudflare's published 198.41.192.x addresses.
  2. Run the T3 Code background service with T3 Connect enabled and confirm a phone can open the environment. cloudflared registers 4 QUIC connections to the virtual addresses it resolved at startup.
  3. Put the Mac to sleep, move it to another network, and wake it. Here the VPN's network extension restarted on wake and rebuilt its address map: region1/region2.v2.argotunnel.com now resolve to different virtual addresses, and the old ones no longer route anywhere.
  4. Try to open the environment from the phone.

I don't have a VPN-free repro yet. The core condition is that the addresses cloudflared resolved at startup stop working while DNS starts returning new ones.

Expected behavior

When the managed relay client is still running but has had no registered tunnel connection for a sustained period, T3 should replace it. A fresh cloudflared re-resolves the edge and reconnects. Doing that by hand restored the tunnel within 3 seconds (see Workaround).

Proposed semantics for maintainers to decide before any PR:

  • Signal: use cloudflared's own readiness endpoint. Spawn it with --metrics 127.0.0.1:0, read the bound address from its Starting metrics server on <addr>/metrics line, and poll /ready. It returns 200 when readyConnections > 0 and 503 otherwise. If that line never appears, do nothing, which keeps today's behavior.
  • Policy: poll every 30 s and count only explicit 503 responses; request errors don't count. After 6 consecutive 503s (about 3 minutes of awake time, since timers pause during sleep), kill the child and let the existing superviseConnector restart it. Double the threshold after each restart that doesn't reconnect, up to about 30 minutes. Reset it once /ready returns 200.
  • Open question: the exit path also queues a managed-tunnel recovery request. A watchdog restart fixes a transport problem, not a credential one, so it may be better to skip that request.
  • Out of scope: reporting running before registration ([Bug]: T3 Connect treats a spawned but unreachable tunnel as 'running' #7447). Also the stale QUIC MTU case (T3 Connect: new connections fail after the host joins a VPN because the managed cloudflared keeps a stale QUIC MTU #15897), where /ready still reports 4 connections.

Actual behavior

Addresses are replaced with documentation ranges: 192.0.2.x stands for the virtual addresses cloudflared resolved at startup, and 203.0.113.x for the ones DNS returned after the network change. 198.41.192.7 is a real, published Cloudflare edge address.

The tunnel stayed down for 2 h 47 min, until I restarted cloudflared by hand:

  • cloudflared kept dialing only the 20 edge addresses it had resolved at startup, 36 hours earlier (192.0.2.120–192.0.2.139). The log has 461 failed to dial to edge with quic: timeout: no recent network activity errors and no successful registration.
  • curl http://127.0.0.1:20241/ready returned {"status":503,"readyConnections":0,...}.
  • The environment's relay hostname returned HTTP 530, and the phone could not connect.
  • T3 did nothing. The process never exited and no registration was rejected, so neither recovery path ran. t3 connect status reports saved setup, not live state, so nothing on the host showed the outage either.

The network was not blocking Cloudflare Tunnel. I checked from the same Mac at the same time, using a raw QUIC version-negotiation probe and an openssl s_client handshake:

Destination UDP 7844 (QUIC) TCP 7844 (TLS)
Current DNS answer for region1/region2 (e.g. 203.0.113.143) reply in ~10 ms Cloudflare presents its Origin certificate
Published edge IP 198.41.192.7 reply in ~8 ms Cloudflare presents its Origin certificate
Address cloudflared kept dialing (192.0.2.121) no reply connection reset

Why it never recovers:

  1. cloudflared resolves edge addresses only once, in NewSupervisor (supervisor.go). edgediscovery.Edge never refreshes them (edgediscovery.go), so it only rotates through the startup set. With normal DNS that set is Cloudflare's stable published IPs, so this rarely shows. With virtual-IP DNS the set can go stale while the process runs.
  2. T3 restarts the connector only in two cases: when the process exits (superviseConnector), or after 4 rejected registrations (observeConnectorOutput). Transport errors are only logged, so a live connector with zero connections is kept forever.

Any DNS that hands out session-scoped virtual IPs, such as other ZTNA clients or fake-IP modes in proxy tools, should hit the same problem. I have only reproduced it with one corporate VPN client.

Related: #7447 (a spawned but unreachable tunnel is reported as running) and #15897 (stale QUIC MTU after joining a VPN). All three treat a live cloudflared process as a working tunnel. This report covers a tunnel that worked and then went stale.

Impact

Blocks work completely

Version or commit

T3 Code desktop 0.0.45 and t3 service 0.0.45. Code links point to main @ f5eb250.

Environment

macOS 27.0 (Apple Silicon); cloudflared 2026.6.0 from Homebrew (picked up from PATH); corporate full-tunnel ZTNA/VPN client;

Logs or stack traces

# 2026-10-04 06:32Z, home network: initial registration (lines trimmed, IDs redacted)
INF Registered tunnel connection connIndex=0 ip=192.0.2.127 protocol=quic

# 2026-10-05 19:00Z, after sleep + network change: every connection drops
ERR failed to accept incoming stream requests error="failed to accept QUIC stream: timeout: no recent network activity" connIndex=0 ip=192.0.2.127

# 19:00Z-21:47Z: only the startup addresses are retried (461 times)
ERR Failed to dial a quic connection error="failed to dial to edge with quic: timeout: no recent network activity" connIndex=0 ip=192.0.2.121

# Meanwhile DNS returns different addresses
$ dscacheutil -q host -a name region1.v2.argotunnel.com
203.0.113.141 ... 203.0.113.150

# 21:47Z, after `kill <cloudflared pid>`
WARN Relay client exited; restarting
INF Registered tunnel connection connIndex=0 ip=203.0.113.143 protocol=quic
INF Registered tunnel connection connIndex=1 ip=203.0.113.204 protocol=quic
INF T3 Connect managed tunnel recovered

Screenshots, recordings, or supporting files

No response

Workaround

Kill only the managed cloudflared process: kill <pid>, where <pid> comes from pgrep -P <t3 serve pid> cloudflared. T3 restarts it right away, the new process re-resolves the edge, and all 4 connections registered within 3 seconds. t3 service restart also works, but it restarts the whole server.


I'm happy to send a focused PR once the intended behavior is agreed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions