Repository navigation
[Bug]: ⚠️ Claude provider: every turn gets stuck showing "Working..." forever after the CLI already completed — stop button doesn't recover it #4452
Description
Activity
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.needs-triageIssue needs maintainer review and initial categorization.Issue needs maintainer review and initial categorization.
on Jul 24, 2026 +1
Hitting this without the network angle: no proxy, direct Anthropic API (
HTTP(S)_PROXY,ANTHROPIC_BASE_URL,CLAUDE_CODE_USE_BEDROCK/_VERTEXall unset), Arch Linux, nightly0.0.29-nightly.20260727.915,claudeAgent+claude-opus-5, CLI 2.1.220.Not 100% of turns here.
queryCurrentContextUsagespans from one session: 409 ms, 8.3 s, 262 s. Unbounded latency rather than a deterministic hang.The 262 s one never resolved on its own — 11:35:40 + 262.207 s = 11:40:02.2, matching
provider.stopAllat09:40:02.258Z. Shutdown killed the CLI, the promise rejected into the existingcatch, which is why the span exitsSuccess.The part that matters more: it freezes every thread on the server, not just the stuck one. Two threads created during the window had
turn-start-requestedaccepted and never started — noprovider/<threadId>.log, no CLI spawned. One was codex, in a different project:1754 750668bd turn-start-requested {"instanceId":"claudeAgent_vanilla","model":"claude-opus-5"} 1758 8d9f306d turn-start-requested {"instanceId":"codex","model":"gpt-5.6-sol"}ProviderCommandReactorprocesses everything on one sequential fiber: singlemakeDrainableWorker(ProviderCommandReactor.ts:1085), andDrainableWorker.ts:50isTxQueue.take+Effect.foreverwith no concurrency. Bothturn-start-requestedandturn-interrupt-requestedqueue behind the hung call — which is also why Stop does nothing, it never gets dispatched.On the fix:
Effect.timeoutalone won't do it,Effect.promiseis uninterruptible. NeedsEffect.disconnect:.pipe(Effect.timeout("10 seconds"), Effect.disconnect, Effect.orElseSucceed(() => undefined))
Separately worth containing: any unbounded await in any adapter reproduces the same server-wide freeze.
Same failure mode on Linux with no proxy and no Bedrock/Vertex, and I have a
dataset that pins the delay to an exact value rather than "hangs forever".Environment
- Linux x64, T3 Code nightly channel, builds spanning
0.0.32→0.0.34
(AppImage; some of them rebuilt locally from the nightly source tree with
unrelated screenshot/image patches only). - Provider:
claudeAgent,claudeCLI2.1.231→2.1.233. - Direct Anthropic API.
HTTP_PROXY,HTTPS_PROXY,ANTHROPIC_BASE_URL,
CLAUDE_CODE_USE_BEDROCK,CLAUDE_CODE_USE_VERTEXall unset, no
apiKeyHelper, no custom base URL.
The delay is a hard 601s, not unbounded
Parsed every
claude/result/success→turn.completedpair across all
per-thread provider logs (~/.t3/userdata/logs/provider/*.log*): 55 turns
total, 20 of them blocked.The distribution is strictly bimodal: either 0-3s, or 601s. Nothing in
between.day turns median max blocked >60s Aug 10 13 0.0s 601.6s 3 Aug 13 2 601.5s 601.5s 2 Aug 14 5 601.0s 602.3s 3 Aug 15 22 0.8s 602.3s 8 Aug 16 12 2.7s 608.8s 4 Blocked values, verbatim: 601.0, 601.1, 601.2, 601.2, 601.3, 601.3, 601.4,
601.4, 601.4, 601.5, 601.5, 601.5, 601.6, 601.6, 602.3, 602.3, 608.8, 194.3.That clustering suggests a 600s timeout somewhere on the control path rather
than a promise that never settles. In this environment the turn does eventually
complete, but the composer is locked and the thread reads "Working" for the full
ten minutes.Confirmation that
getContextUsageis what gates itCleanest single case, one provider log:
15:33:16.818 claude/result/success <- CLI finished, answer already streamed 15:43:25.653 thread.token-usage.updated <- +608.8s 15:43:25.653 turn.completed <- same millisecondAcross all 20 blocked turns, the first event to arrive after the stall is:
14x thread.token-usage.updated 4x claude/system/background_tasks_changed 4x claude/system/init 3x claude/assistant 3x claude/system/commands_changedthread.token-usage.updatedlanding in the same millisecond as
turn.completedin the majority of cases is consistent with the reported root
cause:completeTurn()awaits the context-usage control request before emitting
turn.completed.It stalls unrelated threads too (confirms the earlier comment)
One window with several concurrent threads (one
codex, fourclaudeAgent),
largest gap between consecutiveorchestration_eventsper stream over ~4h:thread provider events largest gap mean gap A codex 1786 (long turn) 8.4s B claudeAgent 408 609s 6.5s C claudeAgent 87 602s 13.5s D claudeAgent 657 1342s 5.5s E claudeAgent 5 1033s 206.6s So the stall is not confined to the thread whose turn is blocked.
Possibly relevant load on the control channel
The same CLI control channel also carries
claude/system/commands_changed
payloads of ~150 KB (~390 entries) in this setup, plus
background_tasks_changed. Two of the five distinct unblock events are exactly
those, so a busy control channel may be what pushes the context-usage request
into its timeout.Happy to run any read-only query or share more aggregate log parsing if useful.
- Linux x64, T3 Code nightly channel, builds spanning
Thanks for the detailed report and the follow-up timing data. We believe this is fixed by PR #5891 and PR #7719.
The first PR adds a one-second timeout around the
getContextUsagerequest, so a stalled request can no longer block turn completion. The second PR repairs orphaned running sessions during startup, including the stale state that remained after a restart.I'm closing this as fixed as part of an automated pass on all open issues. If this still happens in a build that includes both PRs, please reply with the T3 Code version and the timing between
claude/result/successandturn.completed, and we can reopen it.
Before submitting
Area
apps/desktop
Steps to reproduce
Reproduction
state.sqlite:projection_thread_sessions.statusstaysrunningwith a staleactive_turn_ideven though the provider log showsclaude/result/success/terminal_reason:completedfor that turn.Expected behavior
After the Claude provider CLI reports a completed turn (
claude/result/success,terminal_reason: completed),the thread session should transition back to
readyand accept new messages/turns.Actual behavior:
The CLI subprocess completes and streams the correct assistant response, but T3 never emits a
turn.completeddomain event for it. The thread session stays
status: runningwith a staleactive_turn_idforever. The UIshows "Working for Xs" indefinitely, the composer stays locked, and clicking Stop only records an interrupt
request — it does not recover the session. This happens on ~100% of turns in my environment (reproduces even on
a one-word "hi" with no tools involved). Confirmed via the local
state.sqlite(projection_thread_sessionsnever gets a
status:"ready"follow-up event) and per-thread provider logs(
~/.t3/userdata/logs/provider/<threadId>.log), which show the CLI reporting success while T3's server-sideevent log shows nothing after
thread.activity-appended.Likely cause (from apps/server/dist/bin.mjs in app.asar): ClaudeAdapter's
completeTurn()awaitscontext.query.getContextUsage()(from@anthropic-ai/claude-agent-sdk) before emittingturn.completed.That SDK call has no timeout on its underlying control-request Promise. In my environment (corporate network,
all traffic through an internal HTTPS proxy, Claude via AWS Bedrock) this call appears to hang, which blocks
turn.completedfrom ever being emitted — even though the CLI already finished and streamed its answer.Actual behavior
Environment
claudeAgentadapter), modelseu.anthropic.claude-sonnet-5/eu.anthropic.claude-opus-4-8via AWS BedrockclaudeCLI 2.1.218, invoked by T3 with--setting-sources=user,project,localSummary
Every single turn, even a trivial one-word prompt like "hi", gets stuck showing "Working for Xm Ys" indefinitely in the UI. Clicking Stop does not recover it — it only records a
thread.turn-interrupt-requesteddomain event but the session never actually returns toready. This happens on every new thread/session, 100% reproducible, not an occasional race.Root cause (confirmed from logs + DB + app.asar source)
claudeCLI subprocess spawned by T3 does complete successfully. The per-thread provider log (~/.t3/userdata/logs/provider/<threadId>.log) always shows a clean:turn.completeddomain event. Querying the local event store confirms it:projection_thread_sessionsfor the thread staysstatus='running'withactive_turn_idpointing at the turn that the CLI already finished, forever (until the idle reaper eventually kills it ~30 min later, or the user restarts the app — which does not fix it either, since it's persisted instate.sqlite).thread.turn-interrupt-requested(confirmed viaorchestration.command.thread.turn.interruptspans inserver.trace.ndjson) — it never flips the session back toreadyon a turn the provider has already completed and exited.Where it appears to break (apps/server/dist/bin.mjs, extracted from app.asar)
In the Claude adapter's
completeTurn()(called fromhandleResultMessagewhen aresultSDK message withsubtype:"success"arrives):context.query.getContextUsage()comes from@anthropic-ai/claude-agent-sdk'sQuery.getContextUsage():which goes through the SDK's internal
request()control-protocol call:This Promise has no timeout. If the CLI subprocess never sends back a
control_responsefor theget_context_usagecontrol request (e.g. because it's blocked on an internal network call under a corporate proxy, or any other stall), theawaitinqueryCurrentContextUsagehangs forever. SincecompleteTurn()awaits this before emittingturn.completed, the whole turn-completion pipeline silently deadlocks even though the CLI already streamed the finalresultmessage and the user-visible answer.This would explain why:
assistantmessage handling, beforecompleteTurnruns).ready, so the UI shows "Working..." forever.await.getContextUsage) that runs on every single turn completion, not with something related to turn content.I don't have 100% certainty this exact control request is what's hanging (I was not able to attach a debugger to the running
claudeCLI process to confirm), but the DB/log evidence rules out everything beforecompleteTurn(), and the code path is the only awaited, unbounded, non-timed-out promise between "CLI reports success" and "turn.completed is emitted."Suggested fix
context.query.getContextUsage()inqueryCurrentContextUsage(apps/server ClaudeAdapter), falling back to the last known usage snapshot on timeout — the existingtry/catchonly handles rejections, not hangs.awaiton an SDK control-request Promise insidecompleteTurn()should be wrapped in a race against a timeout, since a stuck one currently blocks the entire terminal lifecycle event for that turn.Impact
Blocks work completely
Version or commit
T3 Code (Nightly) 0.0.29-nightly.20260724.892
Environment
No response
Logs or stack traces
Screenshots, recordings, or supporting files
No response
Workaround
Manually patching the local DB (
~/.t3/userdata/state.sqlite, tableprojection_thread_sessions) — settingstatus='ready',active_turn_id=NULLfor the affected thread — unsticks that specific thread, but the same thing recurs on the very next new session, since the root cause is a per-turn stall, not corrupted state.