Skip to content

Desktop server is unresponsive for minutes after launch: two idle fd read streams hold half the libuv pool #16932

Description

@nemesiscodex

What happened

After updating the desktop app (Nightly) and restarting, the app showed " is reconnecting", my theme didn't load, and a "Work provider status: Claude provider status has not been checked in this session yet" toast appeared. This lasted 1.5–5 minutes after every launch, across five launches, then recovered on its own. The server process was alive and listening the whole time.

Diagnosis

The desktop-spawned server permanently loses two of its four libuv worker threads, so two slow operations are enough to stall every fs and DNS call in the process.

  • The server reads two channels inherited from the desktop app with NodeFS.createReadStream("", { fd }):
    • apps/server/src/resourceTelemetry/DesktopTelemetryReceiver.ts:469
    • apps/server/src/preview/DesktopBrowserChannel.ts:103
  • On a pipe or socket fd, fs.ReadStream issues fs.read on the thread pool, and that worker blocks in read() for as long as the channel is idle.
  • sample of the server during stalls, in three separate launches, showed exactly two libuv-worker threads in read for 100% of the samples. In two of those, the other two workers were in uv_getaddrinfo for the whole sample. The main thread was idle in kevent and the process was at about 0.3% CPU.
  • While stalled, /, /health and /.well-known/t3/environment hang until the client gives up. ServerSecretStore.get took up to 63 s and openStaticFile was interrupted after 3.5 s.
  • The delays are queueing, not slow work. Of 979 spans over 3 s, 560 ended within 0.35 s of a 5-second multiple. runGitCommand spans of 5, 20 and 25 s were for commands such as git config --get, which take 0.05 s outside the server (60 concurrent: 0.16 s).
  • DesktopBrowserChannel.ts was added on 2026-10-06 in feat(preview): run the browser on the environment server #15328. That would take the desktop server from one permanently blocked worker to two. I updated on 2026-10-07; I don't know which build I ran before, so the regression link is inferred.

Confirmation: after launchctl setenv UV_THREADPOOL_SIZE 16 and relaunching, the server showed 16 libuv-worker threads and /health answered in about 3 ms at 51 s after launch, when earlier launches were still hung.

Not identified: which hostnames the two blocked getaddrinfo calls were resolving. Lookups of github.com and api.github.com from a fresh Node process, and from the app's own binary, take 6–250 ms.

Ruled out along the way: worktree count (deleted all 40, no change), defaultAutoPull (set to false, no change), and database size (statev2.sqlite is 6.7 GB, but a short fs_usage capture showed only ordinary SQLite reads).

Steps to reproduce

I don't have a minimal repro for the app itself. What I observed:

  1. Run desktop Nightly 0.0.46-nightly.20261007.2774 on macOS with many projects and threads (42 projects here), so startup runs many git, PR-sync and settlement jobs.
  2. Quit and reopen the app.
  3. For the next 1.5–5 minutes, curl -m5 http://127.0.0.1:3773/health times out while the server is idle. sample <server pid> shows two workers in read.

The mechanism reproduces standalone on Node v24.21.0: a child process that opens fs.createReadStream("", { fd }) on four inherited pipes that nobody writes to leaves a later fs.readFile pending past 3 s. With zero or two such streams it takes 3–4 ms.

Version

Desktop Nightly 0.0.46-nightly.20261007.2774 (611132c). Source lines above are from main at f570bd2 (nightly 2787), which still has both readers.

Environment

macOS (Darwin 25.6.0), arm64. t3 triage ran under Node v26.8.2. Tailscale MagicDNS and Bitdefender are installed.

Evidence

sample <server pid> (stalled, 3 launches):
  libuv-worker  read                                   100% of samples
  libuv-worker  read                                   100% of samples
  libuv-worker  uv_getaddrinfo -> mdns_addrinfo -> kevent
  libuv-worker  uv_getaddrinfo -> mdns_addrinfo -> kevent   (2 of 3 launches)
  main thread   uv__io_poll -> kevent (idle)

http.server GET /health                       10.0s Interrupted
http.server GET /.well-known/t3/environment   10.0s Interrupted (repeated)
ServerSecretStore.get                         62.8s Success
spans >3s ending within 0.35s of a 5s multiple: 560 of 979

After UV_THREADPOOL_SIZE=16: 16 libuv-worker threads, /health 200 in 3 ms at t+51s

Related issues

#15971: same pool exhaustion with a different trigger (reads blocked on a TCC-protected file). Its triage comment already notes the 4-thread pool is a hard limit. This report is about two workers being held by design on every desktop launch.

Fix applied or workaround

launchctl setenv UV_THREADPOOL_SIZE 16, then relaunch the app. It lasts until logout. A likely fix is reading those fds with net.Socket instead of fs.createReadStream, so they don't use the pool.

Filed by

claude (opus-5-5) via Claude Code, running npx t3 triage. An earlier part of the investigation ran on sonnet-5-5.

Activity

  1. juliusmarminge commented on Oct 7, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Thanks for the thorough diagnosis. Both idle fd readers you identified already have open fixes that adopt the pipe as a read-only net.Socket instead of holding a threadpool worker:

    Those PRs were opened for the exit-hang symptom of the same mechanism. With both merged, the desktop server should keep its full libuv pool. I've linked this issue from both, and it'll close when #16626 lands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions