Repository navigation
[Bug]: A resumed Claude subagent's child thread stays empty when a later run resumes it #16261
Description
Activity
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.needs-triageIssue needs maintainer review and initial categorization.Issue needs maintainer review and initial categorization.
on Oct 5, 2026 Note
Grok responding on behalf of Julius.
Triage
Thanks @Artem-B for the detailed repro and the run-by-run sequence. It made this much easier to trace.
What I found
On current main (
c138163c8c), the child thread's items are dropped before they're stored, so a reload can't bring them back.- Child-thread items are written with
runId: null(makeSubagentConversationArtifacts, and Claude tool calls under a subagent inensureToolCallStarted). routeProviderEventkeeps aturn_item.updatedonly when the run owns thatrunIdor the item's thread is already inownedThreadIds.- A run gains the child thread in only two places:
app_thread.created(when it launches the subagent), orrelatedThreadIdsat run start, wherecanRouteRelatedSubagentis true only forcompletedsubagents.
That lines up with your sequence. Run 11 launches the subagent and owns the child thread. Run 12 starts while it's still
running, so it doesn't take the thread. The subagent completes during run 12, and the SendMessage reopen reattributessubagent.runIdto run 12 (updateClaudeSubagentNode). The parent card keeps updating becausesubagent.updatedis accepted viaownsRun, but nothing addschildThreadIdto run 12'sownedThreadIds. Run 13 then starts while the subagent isrunningagain, so it doesn't take the thread either. The resume prompt, the Bash calls, andDONE_Care all dropped.The existing
claude_background_subagent_lifecycletest doesn't cover this overlap, since its resume runs only after the subagent completed and the wake run settled. This also looks distinct from #13668, which handled a resume run starting after the subagent was already completed.Likely fix area
One option is enrolling
childThreadIdintoownedThreadIdswhenrouteProviderEventaccepts asubagent.updatedviaownsRun(event.subagent.runId). Two things to watch with that approach:- In
updateClaudeSubagentNode, the child rootnode.updatedand the SendMessage prompt are emitted beforesubagent.updated, so enrolling only on that event would still drop the prompt unless the reattributedsubagent.updatedis routed first. - If the launch fiber is still alive for other background work, it still owns the child thread. It would need to release the thread when
runIdmoves, or the same events could be stored twice (the reasoncanRouteRelatedSubagentexists).
Your two control cases (the same run launching and resuming the subagent, and a later run resuming only after completion while the parent is idle) are worth keeping as regression checks.
A maintainer will decide on the fix direction.
- Child-thread items are written with
- addedvia-triageFiled through npx t3 triageFiled through npx t3 triageand removedneeds-triageIssue needs maintainer review and initial categorization.Issue needs maintainer review and initial categorization.
on Oct 5, 2026 Another way to hit this, with event-log evidence from a real session: stopping a Claude subagent and then resuming it with SendMessage. Besides the empty child thread, this also leaves the child thread's footer reading "Cancelled" while the parent's Agents list shows it Working.
Sequence (from
orchestration_events)Run What happened 17 starts The subagent is completed, so run 17 takes over its child thread.17 SendMessage resumes it. The child root turn goes to runningand the prompt reaches the child thread.18 starts The subagent is still running, so run 18 doesn't take over the child thread.(17's fiber) The parent stops the subagent. The parent's subagent node and the child root turn both go to cancelled, and run 17's fiber ends.18 SendMessage resumes it. updateClaudeSubagentNodemoves the subagent to run 18 and emits the parent's subagent node, the child rootnode.updated, the resume prompt andsubagent.updated. Only the parent-side events are stored: sequence 89294 (parent noderunning) is followed directly by 89295 (subagent.updated), with nothing for the child thread in between. On the earlier resume in run 17, the child root update and prompt were stored (89173/89174).Visible result
- Parent Agents list: Working, with fresh progress. It reads the parent-side subagent record.
- Child thread: the composer bar reads "Cancelled · Runs on its own".
deriveProviderSubagentStatusreads the child's runless root turn, which never got therunningupdate. The thread stops at the previous prompt. This happens on web and mobile.
Two points for the fix
- Stopped subagents fail even without the run overlap.
canRouteRelatedSubagentonly passescompleted, and its comment says "An interrupted, failed or cancelled one is never resumed". Claude's SendMessage does resume them. So stop, then a new run, then SendMessage, loses the child thread even when the new run starts after the stop. - The ordering caveat from the triage applies here. The child root
node.updatedand the prompt are emitted before the movedsubagent.updated. Taking over the child thread only when that event arrives would still drop both, and the bar would still say "Cancelled".
- added 12 commits that reference this issue
on Oct 8, 2026
Before submitting
Area
apps/server
Steps to reproduce
The subagent must complete during the run of step 2. That run then resumes it. The 35 s sleep makes sure of this.
Expected behavior
After the resume, the child thread shows the SendMessage prompt, the 5 Bash calls and DONE_C. They appear live while a parent turn is active, or at the latest when the subagent completes.
Actual behavior
The child thread stops at READY_C, the end of the first run. It never gets the resume prompt, the Bash calls or DONE_C, also after completion. The subagent card in the parent thread does update, and shows
completedwith result DONE_C.The items are not stored, so a reload does not show them. In real use this looks like a hung subagent: an implementer that I resumed with a decision did 40 minutes of work, and its thread showed nothing after its earlier report.
Cause, as far as I can tell from the source at 7812230:
routeProviderEvent(RunExecutionService.ts L349) keeps a child-threadturn_item.updatedonly if the thread is in the run'sownedThreadIds.app_thread.createdfor it (it launched the subagent, L376). Or,canRouteRelatedSubagent(L179, used in ProviderTurnStartService.ts L944) passes it at run start, but only for acompletedsubagent.runIdto run B (ClaudeAdapterV2.ts L4111). This keeps B's fiber alive and routes the card. But thesubagent.updatedcase (RunExecutionService.ts L419) does not addchildThreadIdtoownedThreadIds. After this, no fiber owns the child thread.running, so they do not take it either.A possible fix: when
routeProviderEventaccepts asubagent.updatedbecause the run ownsevent.subagent.runId, addevent.subagent.childThreadIdtoownedThreadIds. Check the case where the launch run's fiber is still alive for other background work, because both fibers would then own the thread.The existing
claude_background_subagent_lifecyclefixture resumes a subagent from a run that started after its completion, so it does not catch this.Impact
Major degradation or frequent failure
Version or commit
0.0.46-nightly.20261005.2676 (7812230)
Environment
Linux server (systemd user service), web client. Claude Code 2.1.289. Parent model claude-opus-5-5 (original case) and claude-sonnet-5-5 (repro). Subagent model haiku (repro) and inherited from the parent (original case).
Logs or stack traces
Screenshots, recordings, or supporting files
No response
Workaround
Use
delegate_taskfor long subagent work, or read the subagent's transcript JSONL under~/.claude/projects/.../subagents/. Do not resume a native subagent from a run that started before the subagent's last completion. This is hard to control, because the parent usually answers the completion report in the same turn.