Skip to content

Codex subagents lose status and message tracking after server restart #15567

Description

@amozeo

What happened

After restarting T3 Code, sending another message to the parent thread resumes its work, but existing Codex native subagents remain marked completed or stopped. Their new messages do not appear in their T3 conversations.

The parent reports four running subagents, while T3’s Lineage panel shows only one. Refreshing the browser does not resolve the discrepancy.

Diagnosis

Provider logs confirm that affected children emit new turn/started events and completed assistant messages after restart. Their SQLite projections retain old completed statuses and message timestamps. This confirms that the missing activity is not merely an inaccurate claim by the parent agent.

The installed version’s CodexAdapterV2.ts explains the behavior:

A subagent created after restart is registered and tracked correctly. The evidence supports failure to reconstruct existing child mappings after restart as the cause.

“Continue threads after restarts” is unset and defaults to off, but that does not explain this failure: the parent was explicitly resumed, and native child events are arriving.

Steps to reproduce

Observed sequence; a separate controlled reproduction was not performed:

  1. Run a Codex parent thread with native subagents.
  2. Restart the T3 server.
  3. Send another message to the parent that causes existing subagents to resume.
  4. Observe that the existing children remain completed or stopped in Lineage and their conversations receive no new messages.
  5. Compare the provider log with the stored projections: native child activity continues, while projected status and messages remain stale.
  6. Refresh the browser; the discrepancy persists.

Expected behavior: existing children resume status and message tracking when they start new turns.

Actual behavior: existing children remain stale, while newly created children are tracked normally.

Version

v0.0.46-nightly.20261003.2632, Commit: f391794a35c604d57e166a3ab48d56fc6e4e469a

Environment

NixOS, Linux x64, kernel 7.2.6, Node.js 24.21.0, Codex CLI 0.159.3, T3 server manually launched with serve --host <LAN address> inside tmux, Browser UI connected to that server

Evidence

All timestamps below are UTC on 2026-10-04. The current server process started at approximately `07:08:17`. Its startup orchestration recovery phase completed successfully.

Provider events and database projections observed during investigation:

| Child | Native `turn/started` | Native completed assistant message | Stored subagent status and update time |
|---|---|---|---|
| `task17_impl` | `07:53:41.924` | `08:02:53.055` | `completed`, `06:29:07.711` |
| `task19_impl` | `07:54:02.431` | `08:07:02.229` | `completed`, `06:32:45.133` |
| `task19_review` | `07:54:13.861` | `08:01:39.695` | `completed`, `06:48:58.384` |

The native events were matched to existing child projections using their native thread IDs. Stored child message timestamps also remained older than the post-restart native messages.

By comparison, `task19_visual_r4`, created after restart, received projected running status and new messages.

Related issues

#13735 fixed a closely related restart-recovery failure for Claude: the adapter lost its in-memory subagent registry, preventing resumed child messages from reaching their original thread. That fix is present in the installed version but does not modify the Codex adapter. This report concerns the Codex counterpart. #15358 concerns threads stranded after restart without automatic continuation. Here, the parent is explicitly resumed and native child events arrive, but T3 does not project their status or messages. #5529 describes a stale resumed-subagent display in an older implementation. This incident also loses projected child messages and points to missing runtime registrations in the current Codex V2 adapter. No confirmed duplicate was found.

Fix applied or workaround

A local read-only viewer was created and verified to display native child status and recent assistant messages from retained provider logs alongside T3’s stored status. It restores terminal visibility but does not repair the UI or backfill conversations.

Creating replacement subagents after safely stopping and handing off existing work is a possible workaround for UI tracking; it has not been performed as part of this triage.

No T3 source patches or database writes were applied.

Filed by

Investigated and drafted by Codex (GPT-6), via t3 triage.

Activity

  1. juliusmarminge commented on Oct 4, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for matching the native events to the stored projections, @amozeo. That made this quick to confirm. It reproduces on current main (1302ccacbd) and is the Codex counterpart of #13735, whose Claude fix doesn't touch this adapter. It isn't a duplicate of #15358, #5529, or #14630.

    What I found

    • CodexAdapterV2 keeps child registrations in subagentThreads, an in-memory map that starts empty for each provider session (around line 1668). resumeThread calls thread/resume with excludeTurns: true and keeps only thread.id and updatedAt (decodeCodexResumeMetadata), so it never rebuilds existing children.
    • After restart, a child turn/started from an unregistered native thread gets parked in pendingSubagentTurns (rememberSubagentTurnStarted). That queue is flushed only by registerSubagentThread, which runs for a new spawnAgent or a subAgentActivity of kind started. Later collabAgentToolCall status updates and other subAgentActivity kinds do nothing for unregistered children.
    • Item and turn/completed events then fail in resolveItemEventContext / activeTurns and return without projecting, so the stored subagent keeps its pre-restart status and messages. That matches your table. A child spawned after restart registers normally.
    • "Continue threads after restarts" isn't the cause, since the parent was resumed explicitly and native events are arriving. Startup recovery also isn't involved: it cancels unfinished work whose process died and correctly leaves already-completed children alone.
    • The existing test "preserves a subagent result across a trailing empty final and resume" stays within one adapter session, so it doesn't cover a process restart. The only CodexAdapterV2 change since f391794 is fix(codex): resume archived native sessions #15389 (resume archived native sessions), which doesn't rebuild this map.

    Likely fix area

    • One option: when a native thread with an existing child projection starts a turn in a new session, reattach it to the same child thread and subagent row, then flush the parked turn the way registerSubagentThread does.
    • Child and provider thread ids come from the native thread id, but the subagent node id comes from the original tool-call id. Treating it as a fresh spawnAgent would likely open a second Lineage row instead of reopening the old one, so the lookup would need to recover that original id.

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    on Oct 4, 2026
  3. xHeaven commented on Oct 5, 2026

    @xHeaven

    Confirmed the same failure on macOS with T3 desktop 0.0.46-nightly.20261005.2676 and native Codex 0.160.0.

    After desktop server startup recovery, the parent resumed and revived 12 existing native children through followup_task. Codex reported all 12 running, but T3's child status and conversation projections remained stale.

    Read-only correlation of the provider logs, Codex rollout and SQLite events gives this timeline. All timestamps are UTC on 2026-10-05:

    20:53:58.753  startup reconciliation cancels the 12 saved child records
    20:58:20.299  fresh Codex initialize for the parent
    20:58:20.345  existing parent thread/resume with excludeTurns=true
    20:59:29.367 through 21:01:49.488
                   new native turn/started for each of the 12 existing children
    21:01:53.303  native list_agents reports all 12 children running
    
    T3 durable state reconstructed at that same instant:
      12 child provider threads idle
      11 latest subagent records cancelled, 1 completed
      0 stored provider turns matching the 12 new native turn IDs
    

    All 12 child identities and their parent lineage remain present, with none archived or deleted. The retained provider log also contains new completed assistant-message items for every child, while each child's stored message timestamp predates its resumed turn. This confirms lost tracking independently of the parent's running-agent claim.

    All 12 children's saved forkedFrom.providerThreadId match the parent's current provider-thread ID. There was no replacement of the parent provider thread in this incident.

    I reviewed #15727 at ed6c95b. Its saved-child recovery on turn/started, preserved launch identities and explicit reopening appear to cover this exact path. The separate concern about recovery after a parent provider-thread replacement does not explain this case. This is a source review, not a live verification of the fix.

    One diagnostic caveat: active-only t3_thread_list is not an authoritative native-agent census. These children have no T3 run records, and its filter uses run-derived status. The evidence above instead matches native turn IDs against durable projections.

    No app state or agents were changed during the investigation. No controlled restart reproduction or live test of the proposed patch was performed. Conversation bodies, credentials and local paths are omitted.

    Prepared with Codex, GPT-6.1 Sol, through npx t3 triage and follow-up source review.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions