Skip to content

fix(server): a Claude subagent resumed after a server restart stays in its thread - #13735

Merged
juliusmarminge merged 1 commit into
t3code/codex-turn-mappingfrom
v2/claude-subagent-resume-restart
Sep 26, 2026
Merged

juliusmarminge merged 1 commit into
t3code/codex-turn-mappingfrom
v2/claude-subagent-resume-restart

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

After a server restart, sending a message to a Claude subagent from before the restart put the reply in the wrong place. The adapter's in-memory subagent registry was empty, so the resume's task_started looked like a new launch. The adapter re-created the child thread, overwrote the original task prompt with the SendMessage text, and reset startedAt. The resumed run's frames carry the original Agent call as parent_tool_use_id, which nothing mapped anymore, so its text and tool output never reached the child thread. This is what happened live in thread 42e302f0 (task a44c0ed810da9e4f3).

What changed

The task id derives every id a subagent needs: node, child thread, child root, and the parent's subagent row. The one thing T3 never stored is the tool call that launched the subagent, and the Claude CLI still has it. Every message in a subagent's transcript is stamped with it, and the SDK's getSubagentMessages(sessionId, agentId) returns it as parent_tool_use_id.

  • ClaudeAgentSdkQueryRunner.subagentLaunchToolUseId reads the first message of that transcript. It is logged to the provider event log as subagent.lookup / subagent.found.
  • On a task_started for a task the registry doesn't know, under a tracked tool call, the adapter looks up the launch id and registers the subagent as a finished run. An Agent launch never becomes a tracked tool call, so a new launch never takes this path. The existing resume path then does the rest: same child thread, launch prompt untouched, the SendMessage text appended as a new prompt, and the resumed frames routed by their launch parent_tool_use_id.
  • If the lookup fails or finds nothing, the adapter logs a warning and falls back to today's behavior.

No contract, schema, or table changes.

Verification

  • New replay test in OrchestratorReplayRecovery.integration.test.ts: the math.ts subagent from thread 42e302f0, hand-ported from the live provider log (paths and user names scrubbed). It runs two orchestrator runtimes against one SQLite file and one transcript cursor, as the Codex and Cursor recovery tests do. It asserts one subagent and one child thread, with messages in order [launch prompt, reply 1, SendMessage prompt, reply 2]. It also checks the launch prompt is the original task, that the resumed run's 3 commands land in the child thread, and that its reply does not leak into the parent thread.
    • Without the recovery call it fails: expected [ 'user', 'assistant', 'assistant' ] to deeply equal [ 'user', 'assistant', 'user', 'assistant' ].
    • With it: passes.
  • The subagent.found value in the transcript is what getSubagentMessages returned against the real session: toolu_015Qai4vD7RwbU7ctXxoebJ6. An unknown agent id returns [].
  • vp test run on OrchestratorReplayRecovery, ClaudeReplayFixtures, ClaudeAdapterV2.test, ProviderRuntimeRecoveryService (test + regression), and OrchestratorReplayFixtures.contract: 6 files, 139 tests pass.
  • OrchestratorReplayFixtures.integration.test.ts -t claude: 26 pass.
  • vp exec tsc --noEmit -p . in apps/server: no errors. vp run knip:check: clean. vp lint on touched files: nothing new. The two unused-declaration warnings were already on the base.

Not run: a fresh live restart-and-resume against the real CLI with this build, or repo-wide checks.

Left open

  • The transcript orders the subagent's frames before each root result, as claude_background_subagent_lifecycle recorded. Live, both runs finished after the root turn settled, so they drained through wake continuation runs. A restart-then-resume through the wake drain is untested.
  • A recovered subagent's parent-thread row comes back with a blank prompt and a startedAt taken from the resume, not the launch. The adapter can't read the persisted row, and the resume overwrites both fields anyway, which matches the in-session behavior.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code


Devin Review

…n its thread

The Claude adapter tracks subagents in memory. After a server restart,
SendMessage to a subagent from before the restart re-emitted task_started
for a task the adapter no longer knew. The adapter treated it as a new
launch: it re-created the child thread, overwrote the launch prompt with
the SendMessage text, and reset startedAt. The resumed run's frames carry
the original Agent call as parent_tool_use_id, which nothing mapped
anymore, so its output never reached the child thread.

The task id derives every projection id. The one piece T3 never stored is
the launch tool_use_id, and the Claude CLI keeps it in the subagent's
session transcript. On a task_started for an unknown task under a tracked
tool call (an Agent launch never becomes a tracked tool call), the adapter
reads it with the SDK's getSubagentMessages and registers the subagent as
a finished run. The resume then takes the in-session path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Sep 26, 2026
new Map(current).set(resume.taskId, subagent),
);
yield* Ref.update(sessionSubagentTaskIdsByToolUseId, (current) =>
new Map(current).set(launchToolUseId, resume.taskId),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 High Adapters/ClaudeAdapterV2.ts:3566

Recovered subagent frames remain stuck in pendingSubagentFramesByToolUseId and never appear in the child thread. Recovery registers only launchToolUseId at sessionSubagentTaskIdsByToolUseId, but the subsequent takeReleasableSubagentFrames call uses the resumed task_started's SendMessage tool id, so it never releases frames buffered under the original launch id. Replay or release the buffered frames using the launch id as well.

🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.ts around line 3566:

Recovered subagent frames remain stuck in `pendingSubagentFramesByToolUseId` and never appear in the child thread. Recovery registers only `launchToolUseId` at `sessionSubagentTaskIdsByToolUseId`, but the subsequent `takeReleasableSubagentFrames` call uses the resumed `task_started`'s SendMessage tool id, so it never releases frames buffered under the original launch id. Replay or release the buffered frames using the launch id as well.

@github-actions

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.8 KiB — 29.3 KiB ✅
Claude Live turn messages — 2 — 8 ✅

Baseline: unavailable · PR result: d74eb8f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@macroscopeapp

macroscopeapp Bot commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Would Approve

Macroscope's review found this PR approvable — This is a focused Claude restart-recovery fix with dedicated replay coverage and no schema, deployment, security, billing, or default changes. An unresolved High-severity finding reports that some recovered frames may remain buffered instead of reaching the child thread, representing a correctness blocker that must be addressed separately.

Not approved because:

  • 1 blocking correctness issue found at or above your repo's Minimum Blocking Severity

Adjust the Minimum Blocking Severity for this repo — including turning it Off — in Settings. You can add or adjust custom eligibility rules. Learn more.

@juliusmarminge
juliusmarminge merged commit 70629cd into t3code/codex-turn-mapping Sep 26, 2026
24 of 25 checks passed
@juliusmarminge
juliusmarminge deleted the v2/claude-subagent-resume-restart branch September 26, 2026 01:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L 100-499 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant