Repository navigation
delegate_task mode:"wait" dies at the client ~5 min ceiling and returns no taskId, causing duplicate child dispatch #11168
Description
Activity
Triage
Confirmed as a real Orchestrator V2 MCP bug. The report matches the current V2 branch (
t3code/codex-turn-mapping/ PR #2829). This code is not onmain.What breaks
delegate_taskwithmode: "wait"creates the child, then holds the same MCPtools/callopen whilewaitForTaskpolls every 50ms until the child is terminal or the server wait budget elapses:- default wait: 10 minutes
- max wait: 60 minutes
Nothing is written on the HTTP/SSE stream during that poll. A Node
fetch/ undici client dies first at the default 300s header/body timeout and surfaces a barefetch failedwith no tool result.taskIdandchildThreadIdexist on the server before the wait starts, but they are only returned after the wait resolves. The caller therefore has no handle fortask_statusand cannot tell "never launched" from "still running". Re-dispatch is the natural next step and creates a second child in the same cwd.The two ~5m45s retries in the report match that ceiling plus one agent turn. Short waits still return a complete envelope, so wait-mode itself works; only long waits are unsurvivable.
Why the documented recovery cannot fire
The tool description and schema (from PR #7427) tell agents that
timeoutMsis only the parent's wait budget, thatwaitTimedOut: truedoes not cancel the child, and that they should keeptaskIdand polltask_status. That envelope is only produced when the server timer wins. Against a 5-minute client ceiling and a 10-minute default, it cannot.clientRequestIdis idempotent when the caller supplies the same value (integration tests replaydelegate-claude-1and get the sametaskId). If it is omitted, the server generates a random UUID, so a retry cannot dedupe.t3_thread_listcan list subagent children (includeSubagentsdefaults to true). The tool text does not tell agents to use that before retrying, and there is notask_list.Wait-mode also sets
completionWake: "settled_only"and only upgrades to"always"on the server timeout path. A client disconnect leaves the original child running without that upgrade, so a later completion may not wake a still-active parent.The first child in the report ran to
completedafter the wait died, so the fetch timeout itself does not interrupt the child. Parent handoff/interrupt coupling looks like a separate question, not this failure mode.t3_thread_waitreuses the same 10/60-minute silent budget and has the same keep-alive gap.Related
- Not a duplicate of [Bug]: Session hung forever after spawning subagents #2778 (OpenCode permission hang).
- PR fix(orchestrator): Stop treating a wait timeout as a dead child #7427 documented timeout ≠ dead child; it does not fix the client dying first.
- PR fix(orchestrator): report terminal runs after wait timeout #11156 is a different
t3_thread_waittimedOutflag issue. - PR feat(orchestrator): introduce new orchestrator #2829 is the V2 tracking branch and does not claim this fix.
Suggested fix (on the V2 branch)
- Emit MCP progress notifications (or SSE keep-alives) during
mode: "wait"so Node client timers reset. - Publish
taskId/childThreadIdin an early progress notification so a dead wait can still be reconciled withtask_status. - Treat client disconnect like server timeout for the
completionWakeupgrade.
Optional safety net: clamp default/max wait below typical client ceilings so the documented
waitTimedOut: trueenvelope can actually return. Also document: on anydelegate_tasktransport error, callt3_thread_listbefore re-dispatch, and pass a caller-ownedclientRequestId.- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 11, 2026 Additional confirmed occurrence on Windows x64, T3 Code
0.0.46-nightly.20261008.2849, using the Codex provider (codex-proxy,gpt-6.1-sol). This is the same 300-second failure, with nested delegation explaining why a short triage run held the outer wait open.Observed timeline (2026-10-09, UTC)
The scheduled parent called:
{ "target": { "providerInstanceId": "codex-proxy", "model": "gpt-6.1-sol", "options": { "reasoningEffort": "medium" } }, "role": "review", "mode": "wait", "timeoutMs": 600000, "runtimeMode": "full-access", "clientRequestId": "fobbitmc-triage-20261009T0305", "task": "Read the monitor instructions and state, perform triage, and return a one-line result." }The task text above summarizes the original file-based prompt; the remaining fields are exact.
- 03:05:14.727: triage child A created.
- 03:05:44.823: A created review child B through
delegate_task(mode: "async"), usingclaude-proxy/claude-opus-5-5. - 03:06:04.809: A's original run completed with its one-line escalation result.
- 03:08:07.480: B created implementation child C through
delegate_task(mode: "async"), usingcodex-proxy/gpt-6.1-sol. - 03:08:48.977: B's original run completed; C continued running.
- 03:10:14.690: the outer
delegate_tasktool result was a transport error, with no task handle:
tool call error: tool call failed for `t3-code/delegate_task` Caused by: timed out awaiting tools/call after 300s- 03:10:18.120: the parent concluded,
Triage child failed: delegate_task tool call timed out after 300 seconds. - A subsequent read found C still running, with activity updated at 03:14:27.358. The transport timeout had not stopped the descendant work.
These facts were read from the durable parent, child, grandchild, and implementation-thread timelines. No isolated reproduction was run, and no changes were made to those threads or schedules.
Source check
The local upstream checkout inspected was clean
mainat3143335fc3a568cbbb5174764961889272135cb9. That checkout is not asserted to be the exact installed nightly commit.OrchestratorMcpService.ts:99still sets a 10-minute default / 60-minute maximum wait budget.waitForTaskwaits for terminal task status, not merely the original child run ending.delegatedTaskProgressdeliberately reportswaiting_for_childrenwhile nested tasks or their completion deliveries remain outstanding. Thus A's short original run does not imply the outer wait can finish.CodexAdapterV2.ts:1338supplies the MCP endpoint and authorization header without an explicit tool timeout override.delegateTaskreturns the handle only after the blocking wait, and changescompletionWakefromsettled_onlytoalwayson the server timeout path.
This adds a Codex-client occurrence to the original Pi/undici report; the exact Codex transport error establishes the 300-second client limit here, without attributing its implementation to undici. The bug is the client/server wait-budget mismatch and lost recovery envelope. Waiting for nested work itself is intentional. No duplicate dispatch or permanent result loss was established in this occurrence.
Reproduction recipe and workaround
Use a Codex parent to delegate A with
mode: "wait", timeoutMs: 600000; have A start B asynchronously and return, and have B leave a descendant running beyond five minutes. In this observed run, the parent lost its outer call at 300 seconds while descendant work continued.For these scheduled checks, using outer
mode: "async"would return the handle immediately and avoid the long blocking call. This is a source-supported workaround, not a change made to the user's schedules or separately tested here. The graceful wait path needs to returntaskIdandwaitTimedOutbefore the client deadline, or otherwise reconcile the transport budget with the advertised server wait.
Summary
delegate_taskwithmode: "wait"is unsurvivable for any child that runs longer than ~5 minutes when the MCP client is Node-based (Pi provider here, but this is true of anyfetch/undici client). The wait is a single synchronous MCP request held open for the child's entire run; undici's defaultheadersTimeoutis 300 s, so the HTTP request dies first and the caller gets a barefetch failed— with notaskIdand nochildThreadId.The consequence is worse than a lost result: the parent agent has no handle to reconcile with, cannot call
task_status, and cannot tell "the child never launched" from "the child is running fine". The only apparent recourse is re-dispatch — which spawns a second child writing into the same cwd.Evidence (two independent occurrences, same day, same fingerprint)
completedinterrupted21:27:02Zinterrupted01:18:17ZBoth retries landed at ≈5m45s after dispatch — the undici 300 s ceiling plus one agent turn. Control: a trivial
mode:"wait"child (seconds long) returns a clean, complete envelope (taskId,childThreadId,status: completed,summary), so wait-mode itself is fine; only its duration tolerance is broken.In the second case the duplicate child was smart enough to notice the first child's commit and refuse to redo the work. That was luck, not a guarantee — two concurrent writers in one working tree is the real hazard.
Where the mismatch lives
apps/server/src/mcp/OrchestratorMcpService.tsclamps the wait budget to a default of 10 minutes and a maximum of 60 minutes. Both sit well above the ~5-minute ceiling a Nodefetchclient can hold a response open, and nothing is emitted on the wire during the wait to keep the connection alive. The tool description also advertiseswaitTimedOut+ "keep that taskId and read status on latertask_status" — but that graceful path is only reachable when the server's own timer fires first, which by construction it cannot for the default budget.Related: when the parent's turn was handed off/interrupted while a wait was in flight, the delegated child was interrupted too, so a child's lifetime appears coupled to the parent's in-flight tool call.
Suggested fixes (any one helps; 1+3 would close it)
mode: "wait". Traffic on the stream resets the client's header/body timers, which makes long waits viable and is the behavior most MCP clients expect for long-running tools.waitTimedOut: trueenvelope, which at least preserves thetaskId.taskIdobservable before the wait resolves — e.g. send it in an early progress notification, so a parent whose wait died can reconcile viatask_statusinstead of re-dispatching. Documenting "on anydelegate_tasktransport error, callt3_thread_listbefore retrying" would also help agents, since a server-generatedclientRequestIdcannot dedupe a retry (the caller never saw it).Environment
macOS 15, Pi provider (
anthropic/claude-opus-5), full-access runtime. Build is a local fork tracking the Orchestrator V2 branch (PR #2829) at8f44bec/ 0.0.40, with two unrelated local patches (Pi in the Usage dashboard; delegated-wake cap removal) — neither touches the MCP layer. The code paths cited are upstream-unmodified.