Skip to content

[Bug]: Retrying an MCP mutation with the same clientRequestId creates duplicate work after a server restart #16184

Description

@coygeek

Summary

Orchestrator MCP tools promise that reusing a clientRequestId makes a retry idempotent, and the injected agent instructions tell agents to rely on that. The derived command, thread, and message ids include scope.requestNamespace, which is a random UUID minted for each MCP credential and held only in memory. Any credential re-issue (a server restart, a provider session release and re-attach, a capability change, or liveness expiry) therefore changes the namespace, and a retry of the same call with the same key misses the durable command receipt and creates new work. When the opt-in "continue threads after server update" setting is on, a restart resumes the parent's interrupted turn, so the documented retry is a realistic path to a duplicate delegated child, duplicate threads, a duplicate message, or a duplicate scheduled task.

Steps to reproduce

Known case by static trace, not executed:

  1. In a thread on a provider with T3 MCP tools, call delegate_task with {"task":"...","mode":"wait","clientRequestId":"review-round-1"}.
  2. With the "continue threads after server update" setting turned on (it is off by default), restart the T3 server while the call is waiting (default wait budget is ten minutes), for example by applying an update.
  3. After restart, the parent's interrupted run is reconciled and a continuation run is dispatched on the same native provider thread. Other adapters receive the prompt "Continue where you left off."; on Codex a turn cut mid-way with no cancelled-work note resumes natively without that prompt text.
  4. In that continuation, call delegate_task again with the identical input, including clientRequestId: "review-round-1".
  5. Call t3_thread_list with includeSubagents: true, or inspect the parent timeline.

The missing runtime observation is whether a resumed provider agent re-issues the identical call on its own; step 4 performs that retry explicitly.

Expected behavior

The shared clientRequestId schema is described to agents as a "Stable idempotency key to reuse when retrying this mutation." The delegate_task tool description says to use "a distinct clientRequestId per round, stable across retries of that round", t3_thread_update, t3_thread_send, and t3_thread_interrupt say "clientRequestId makes retries idempotent", and the injected instructions say to "use stable clientRequestId values when retrying tools that accept them". None of the agent-facing text limits this to the lifetime of an in-memory credential, which the agent cannot observe. Step 4 should return the task created in step 1 rather than create a second child.

Actual behavior

By static trace: McpSessionRegistry.issue mints providerSessionId with crypto.randomUUIDv4 and sets requestNamespace: providerSessionId; registry state and the thread-to-credential map are in memory, so a restarted server issues a new UUID for the same thread. stableCommandId, stableThreadId, and stableMessageId all include stablePart(input.scope.requestNamespace), so step 4 derives a different command id. The orchestrator's durable dedupe is commandReceipts.getByCommandId(command.commandId), which finds nothing for the new id, and dispatchDelegatedTaskRequest derives the child thread and task node only from that command id. The parent-run check passes because the continuation run is active on the same provider instance. The result is a second child thread running the same task in the parent's shared workspace while the first child, kept open across the restart, continues. Within one running server the namespace can also change through session release (clearMcpSession), a browser or device access change that forces a new credential, or the 24-hour liveness expiry. Runtime reproduction was not executed.

Evidence

Runtime reproduction was not run. Reachability: the continuation is dispatched on the same native provider thread as a new active run, only when the continueThreadsAfterServerUpdate setting is enabled (default false, packages/contracts/src/settings.ts:1223-1225); ProviderRuntimeRecoveryService keeps delegated task rows open across restart; ProviderSessionManager reuses a credential only while readMcpProviderSession(threadId) exists in memory and resolves, and otherwise issues a new one (L461-488). No test covers retries across two credentials for the same thread; existing tests hard-code a single requestNamespace in their scope.

The only scoped wording is in the developer design document docs/orchestration-v2/orchestrator-mcp-server.md L444-445: "clientRequestId derives stable command, thread, and message IDs within the provider session." For adapters that share one provider session per instance (Codex), the provider session id is fixed per instance and survives restart, so the retry breaks even that scoped wording; for adapters with a session per thread the wording is ambiguous, and the agent-facing text above is the contract agents act on. The namespace here is the credential's random UUID; the code comment on the field (apps/server/src/mcp/McpInvocationContext.ts:47) states its purpose as keeping "two callers reusing a clientRequestId" from colliding, not limiting retries to one credential.

Restoration check

Issue two MCP credentials for the same thread (as a restart or re-attach does) and send the same delegate_task input with the same clientRequestId through each. Today the second call creates a second child task and thread; after the fix it returns the first task's taskId and childThreadId. Control: two different threads using the same clientRequestId string must still create independent work, and a new key from the same thread must still create a new round.

Additional context

Affected tools share the derivation: delegate_task, create_threads, schedule_task, t3_thread_send, t3_thread_update, t3_thread_interrupt, and task_cancel; the visible harm is largest for the creating tools. Related: #11168 (a mode:"wait" call that times out client-side without a taskId leads to a duplicate child; its triage recommends a caller-owned clientRequestId as the remedy, which this defect defeats across credential re-issue). Credential re-issue paths are also discussed in #15173, #16000, and #14076.

Area

apps/server

Impact

Minor bug or occasional failure. A documented retry duplicates delegated children, threads, messages, or schedules after any MCP credential re-issue; it needs a restart or session release to overlap an in-flight mutation and the agent to retry.

Version or commit

main at 6f9cea00ae967f38fa3cdc9c07f92031806f4264.

Environment

Source inspection of pingdotgg/t3code main at 6f9cea0. Not executed against a running T3 server.

Workaround

Before retrying, an agent could list the thread's subagents or threads with t3_thread_list and reuse what it finds, but agents are instructed to rely on clientRequestId instead and this check was not exercised. (Not verified.)

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions