Skip to content

T3 Connect: new connections fail after the host joins a VPN because the managed cloudflared keeps a stale QUIC MTU #15897

Description

@bramley-jetcharge

What happened

The host is a MacBook running the T3 Code desktop app with T3 Connect on, and the client is the Android app. The phone can't open a new connection to the host while the host is on a full-tunnel work VPN. The errors shown include "Your cloud sign-in changed", "Your DNS or firewall may be blocking T3 Connect" and "Could not connect relay environment". A phone session opened before the VPN came on keeps working. Turning the VPN off lets the phone connect right away. The VPN is needed on the host for work, so turning it off isn't a real option.

Diagnosis

The cause is QUIC path MTU in the managed cloudflared. It is not DNS, a firewall, or sign-in.

  1. T3 started cloudflared tunnel run while the Mac was on its home network (MTU 1500). Each of the 4 QUIC connections probed up to quic_client_mtu = 1344 (max_udp_payload = 1360).
  2. The VPN (UniFi Identity Enterprise VPN) then became the default route. Its utun interface has MTU 1280. QUIC doesn't lower its MTU when that happens, so full-size packets were silently dropped. 3 of the 4 connections showed smoothed_rtt 0 and hundreds of lost_packets{reason="time_threshold"}, but /ready still reported readyConnections: 4 and status 200.
  3. Small exchanges still fit, so the tunnel looked healthy. Every relay POST /api/t3-connect/mint-credential reached the host and returned 200, and existing sessions stayed up. Anything bigger stalled until it timed out. Through the public tunnel hostname, GET /api/auth/session (230 B) answered in about 0.1s, while GET /favicon.ico and POST /api/auth/websocket-ticket hung for over 10s. The same requests answer in under 1 ms on 127.0.0.1:3773.
  4. While the phone was failing, the host saw a mint-credential every ~25s, all 200. Only some were followed by POST /oauth/token, and the connection never got past setup. That's why the client reports a connect/network error even though the host's traces are all Success.
  5. Killing cloudflared (T3 respawned it, ManagedEndpointRuntime.ts) fixed it straight away. The new connections probed over the VPN and settled on quic_client_mtu = 1252. Every request above then completed in about 0.1s, including the 82 KB favicon, and the phone connected with the VPN still on.

Nothing in apps/server/src/cloud/ (checked at main @ a1d9d72) reacts to the host's default route or interface changing. cloudflared is always started with tunnel --no-autoupdate --loglevel info --output default run and the default QUIC transport, and it keeps the MTU it picked on the old network for its whole life. Related symptom: #7447 (a spawned but unhealthy tunnel is reported as running), except here registration succeeded and the connections degraded afterwards.

Possible directions, for maintainers to judge: restart the managed connector when the default route or interface changes; detect connections that are registered but stalled (high loss, zero RTT) and recycle them; or let users force --protocol http2 (TCP adapts through MSS).

Steps to reproduce

  1. On a network with MTU 1500, start the T3 Code desktop app with T3 Connect on. Confirm the Android app can connect.
  2. Connect the host to a full-tunnel VPN whose interface MTU is 1280 (WireGuard-style). Confirm curl -s http://127.0.0.1:20241/metrics | grep quic_client_mtu still shows 1344.
  3. Force-quit the Android app and reopen it. Opening the environment fails with the errors above.
  4. Through the tunnel hostname, curl https://<tunnel-host>/api/auth/session succeeds quickly, while curl https://<tunnel-host>/favicon.ico hangs.
  5. pkill -f "cloudflared tunnel", wait about 30s for T3 to respawn it, and retry. quic_client_mtu is now 1252 and the phone connects.

Version

Desktop 0.0.45 (macOS app "T3 Code (Alpha)"), Android app 1.3.0

Environment

macOS 15.7.3 (Apple silicon), managed cloudflared 2026.9.3, UniFi Identity Enterprise VPN (full tunnel, utun MTU 1280)

Evidence

# cloudflared metrics while broken (VPN on, connector started before the VPN)
quic_client_mtu{conn_index="0..3"} 1344
quic_client_max_udp_payload{conn_index="0..3"} 1360
quic_client_smoothed_rtt{conn_index="0"} 14   # conn 1..3: 0
quic_client_lost_packets{conn_index="0",reason="time_threshold"} 109
quic_client_lost_packets{conn_index="1",reason="time_threshold"} 125
quic_client_lost_packets{conn_index="2",reason="time_threshold"} 68
quic_client_lost_packets{conn_index="3",reason="time_threshold"} 231
/ready -> {"status":200,"readyConnections":4}

# through the public tunnel hostname, while broken
GET  /api/auth/session            200 0.11s 230B
GET  /favicon.ico                 000 10.0s (timeout)   # 82338B, 0.001s on 127.0.0.1
POST /api/auth/websocket-ticket   000 10.0s (timeout)   # 401 in 0.0008s on 127.0.0.1

# host server.trace.ndjson while the phone was failing (all spans Success)
15:11:04 POST /api/t3-connect/mint-credential 200
15:11:17 POST /api/t3-connect/mint-credential 200
15:11:29 POST /api/t3-connect/mint-credential 200
15:11:42 POST /api/t3-connect/mint-credential 200
15:12:09 POST /api/t3-connect/mint-credential 200
...

# after killing cloudflared and T3 respawning it (VPN still on)
quic_client_mtu{conn_index="0..3"} 1252
GET  /api/auth/session            200 0.08s
GET  /favicon.ico                 200 0.11s 82338B
GET  /                            200 0.09s 19493B
POST /api/auth/websocket-ticket   401 0.08s (x5)

Related issues

#5744 (native clients time out "despite healthy relay tunnel" after a host restart). It may be the same failure if that host's connector started on a path with a different MTU; that report never pinned down the cause. #14264 shows the same "DNS or firewall" message, but there the cause is on the phone (iOS Connectivity Assist). Here the cause is on the host. #7447 is about a tunnel that never registers; here it registered and degraded later.

Fix applied or workaround

pkill -f "cloudflared tunnel" after connecting the VPN. T3 respawns the connector, which re-probes MTU on the new path. The fix only lasts until the network path changes again.

Filed by

Claude Code (claude-opus-5-5), following .github/triage/PLAYBOOK.md

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions