Repository navigation
Codex subagents lose status and message tracking after server restart #15567
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Thanks for matching the native events to the stored projections, @amozeo. That made this quick to confirm. It reproduces on current
main(1302ccacbd) and is the Codex counterpart of #13735, whose Claude fix doesn't touch this adapter. It isn't a duplicate of #15358, #5529, or #14630.What I found
CodexAdapterV2keeps child registrations insubagentThreads, an in-memory map that starts empty for each provider session (around line 1668).resumeThreadcallsthread/resumewithexcludeTurns: trueand keeps onlythread.idandupdatedAt(decodeCodexResumeMetadata), so it never rebuilds existing children.- After restart, a child
turn/startedfrom an unregistered native thread gets parked inpendingSubagentTurns(rememberSubagentTurnStarted). That queue is flushed only byregisterSubagentThread, which runs for a newspawnAgentor asubAgentActivityof kindstarted. LatercollabAgentToolCallstatus updates and othersubAgentActivitykinds do nothing for unregistered children. - Item and
turn/completedevents then fail inresolveItemEventContext/activeTurnsand return without projecting, so the stored subagent keeps its pre-restart status and messages. That matches your table. A child spawned after restart registers normally. - "Continue threads after restarts" isn't the cause, since the parent was resumed explicitly and native events are arriving. Startup recovery also isn't involved: it cancels unfinished work whose process died and correctly leaves already-completed children alone.
- The existing test "preserves a subagent result across a trailing empty final and resume" stays within one adapter session, so it doesn't cover a process restart. The only
CodexAdapterV2change sincef391794is fix(codex): resume archived native sessions #15389 (resume archived native sessions), which doesn't rebuild this map.
Likely fix area
- One option: when a native thread with an existing child projection starts a turn in a new session, reattach it to the same child thread and subagent row, then flush the parked turn the way
registerSubagentThreaddoes. - Child and provider thread ids come from the native thread id, but the subagent node id comes from the original tool-call id. Treating it as a fresh
spawnAgentwould likely open a second Lineage row instead of reopening the old one, so the lookup would need to recover that original id.
A maintainer will decide on the fix direction.
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.
on Oct 4, 2026 Confirmed the same failure on macOS with T3 desktop
0.0.46-nightly.20261005.2676and native Codex0.160.0.After desktop server startup recovery, the parent resumed and revived 12 existing native children through
followup_task. Codex reported all 12 running, but T3's child status and conversation projections remained stale.Read-only correlation of the provider logs, Codex rollout and SQLite events gives this timeline. All timestamps are UTC on 2026-10-05:
20:53:58.753 startup reconciliation cancels the 12 saved child records 20:58:20.299 fresh Codex initialize for the parent 20:58:20.345 existing parent thread/resume with excludeTurns=true 20:59:29.367 through 21:01:49.488 new native turn/started for each of the 12 existing children 21:01:53.303 native list_agents reports all 12 children running T3 durable state reconstructed at that same instant: 12 child provider threads idle 11 latest subagent records cancelled, 1 completed 0 stored provider turns matching the 12 new native turn IDsAll 12 child identities and their parent lineage remain present, with none archived or deleted. The retained provider log also contains new completed assistant-message items for every child, while each child's stored message timestamp predates its resumed turn. This confirms lost tracking independently of the parent's running-agent claim.
All 12 children's saved
forkedFrom.providerThreadIdmatch the parent's current provider-thread ID. There was no replacement of the parent provider thread in this incident.I reviewed #15727 at
ed6c95b. Its saved-child recovery onturn/started, preserved launch identities and explicit reopening appear to cover this exact path. The separate concern about recovery after a parent provider-thread replacement does not explain this case. This is a source review, not a live verification of the fix.One diagnostic caveat: active-only
t3_thread_listis not an authoritative native-agent census. These children have no T3 run records, and its filter uses run-derived status. The evidence above instead matches native turn IDs against durable projections.No app state or agents were changed during the investigation. No controlled restart reproduction or live test of the proposed patch was performed. Conversation bodies, credentials and local paths are omitted.
Prepared with Codex, GPT-6.1 Sol, through
npx t3 triageand follow-up source review.
What happened
After restarting T3 Code, sending another message to the parent thread resumes its work, but existing Codex native subagents remain marked completed or stopped. Their new messages do not appear in their T3 conversations.
The parent reports four running subagents, while T3’s Lineage panel shows only one. Refreshing the browser does not resolve the discrepancy.
Diagnosis
Provider logs confirm that affected children emit new
turn/startedevents and completed assistant messages after restart. Their SQLite projections retain old completed statuses and message timestamps. This confirms that the missing activity is not merely an inaccurate claim by the parent agent.The installed version’s
CodexAdapterV2.tsexplains the behavior:subagentThreadsis an in-memory map initialized empty.resumeThreadcallsthread/resumewithexcludeTurns: truewithout rebuilding existing child registrations.rememberSubagentTurnStartedholds turns from unregistered children pending registration.A subagent created after restart is registered and tracked correctly. The evidence supports failure to reconstruct existing child mappings after restart as the cause.
“Continue threads after restarts” is unset and defaults to off, but that does not explain this failure: the parent was explicitly resumed, and native child events are arriving.
Steps to reproduce
Observed sequence; a separate controlled reproduction was not performed:
Expected behavior: existing children resume status and message tracking when they start new turns.
Actual behavior: existing children remain stale, while newly created children are tracked normally.
Version
v0.0.46-nightly.20261003.2632, Commit:f391794a35c604d57e166a3ab48d56fc6e4e469aEnvironment
NixOS, Linux x64, kernel
7.2.6, Node.js24.21.0, Codex CLI0.159.3, T3 server manually launched withserve --host <LAN address>inside tmux, Browser UI connected to that serverEvidence
Related issues
#13735 fixed a closely related restart-recovery failure for Claude: the adapter lost its in-memory subagent registry, preventing resumed child messages from reaching their original thread. That fix is present in the installed version but does not modify the Codex adapter. This report concerns the Codex counterpart. #15358 concerns threads stranded after restart without automatic continuation. Here, the parent is explicitly resumed and native child events arrive, but T3 does not project their status or messages. #5529 describes a stale resumed-subagent display in an older implementation. This incident also loses projected child messages and points to missing runtime registrations in the current Codex V2 adapter. No confirmed duplicate was found.
Fix applied or workaround
A local read-only viewer was created and verified to display native child status and recent assistant messages from retained provider logs alongside T3’s stored status. It restores terminal visibility but does not repair the UI or backfill conversations.
Creating replacement subagents after safely stopping and handing off existing work is a possible workaround for UI tracking; it has not been performed as part of this triage.
No T3 source patches or database writes were applied.
Filed by
Investigated and drafted by Codex (GPT-6), via
t3 triage.