Repository navigation
Five bugs found while running the orchestrator v2 branch (#2829) day to day #13331
Description
Activity
Checked all five against
t3code/codex-turn-mapping@3e4ca4c532(#2829) and currentmain@d4cd7d5c33(ahead of theeffaab94e3you cited). All five are still real.Where they live: 1, 2, 4, and 5 are orchestrator-v2 only. Those files are not on
main(and 5 is a v2 regression of behaviormainalready got right). 3 is onmain. The preview host, desktoploadURLpath, and broker are identical on this branch andmain, so 3 should targetmain. The other four should target this v2 branch. Focused PRs as you offered are welcome; I have not reviewed the patches themselves.1. Restart continuation skips live non-
runningsessions (v2) — confirmedrestartContinuationRunonly keeps an in-flight run whensession.status === "running"(apps/server/src/orchestration-v2/RestartContinuation.ts). Recovery still cancels that run either way (ProviderRuntimeRecoveryService.ts); the"running"check is what decides whether aprovider-runtime.continueeffect is queued. With "Continue threads after restarts" on, a Codex/Claude/Cursor/ACP thread mid-turn is cancelled and left for a manual continue.One correction: OpenCode is not the only writer of
"running".PiAdapterV2also emitsprovider_session.updatedwith"running"on turn start. Codex, Claude, Cursor, and ACP create the session as"ready"and never promote it.ProviderSessionManager's busy/idle tracking is in-memory only; the statuses it persists are"stopped"and"error"on release. Those adapters do mark the provider thread"active"and the provider turn"running", so the session-status check is the gate that fails.Treating a session as live unless it is
"stopped"or"error"matchesOrchestrator,ProjectionStore,ProviderSwitchService, and recovery itself. Note that OpenCode also sets the session to"waiting"while a permission or question is pending; that status would start qualifying too. Prepared continuations already bypass this check, so this is the first crash of a live turn, not the second crash of an already-admitted continuation.2. Steered subagent stays completed in Lineage (v2) — confirmed
deriveThreadRelationshipGraphfirst builds the child edge fromthread.activityRunStatus, then overwrites it. Subagent edges fromprojection.subagentsuse the same source/target/kind key and replace that status withsubagent.status(packages/client-runtime/src/state/threadRelationships.ts). On the parent,ThreadRelationshipsControlthen groups "Previous agents" from that edge and, whenever a projection is loaded, counts "N running" only from records withstatus === "running".A new run on the child updates the child shell's
activityRunStatus. Nothing writes that back onto the parent's settled subagent row. Provider-native resume during the spawning turn can reopen the row (Codex adapter tests); a follow-up after the record has settled does not. Prefer the child shell's liveactivityRunStatusfor the edge and the count, and fall back to the record when the shell is absent. The hover card and the screen-reader status still readagent.statusoff the record, so those need the same preference or the tooltip stays "completed".#7314 is a different bug (stuck red after a usage-limit recovery, and status frozen for the whole run). Not a duplicate.
3. Hung
preview_navigate/ reusedpreview_openoutlives the broker deadline and drops the host (main) — confirmedThis is not a v2 bug. Same code on
main@d4cd7d5c33.In
PreviewAutomationHosts.tsx, theopenpath that reuses a tab and thenavigatepath bothawait previewBridge.navigate(...)with no deadline, then pass the fullrequest.timeoutMs(orinput.timeoutMs) intowaitForNavigationReadinessinstead of the time left beforehostDeadlineMs. On desktop,PreviewManager.navigateawaitswebContents.loadURLuntil that load settles (apps/desktop/src/preview/Manager.ts), so a page that never finishes loading never answers. The broker treats an unanswered request as a dead connection, disconnects that client, andremoveConnectionFromStatedeletes its assignments, so the current tab is forgotten and later calls open new tabs (PreviewAutomationBroker.ts).Overlay registration already races
hostDeadlineMs(the budget from #4685). Racingnavigateagainst that deadline and giving readiness only the remainder matches that. #12407 is the same eviction mechanism forpreview_wait_for; it is closed and does not cover this path. Resize is a lesser cousin:waitForRenderedViewportalso starts a fresh full timeout, but it is bounded, unlikeloadURL.4. Codex detach leaves the native thread and its MCP servers loaded (v2) — confirmed
Codex is the only driver with
supportsMultipleProviderThreadsPerSession: true. Detach on that path interrupts in-flight turns, then drops the app thread fromattachedThreadIdsandloadedProviderThreadKeyByThread. It never tells the runtime to unload the native thread, andProviderAdapterV2SessionRuntimehas no unload operation. The shared app-server stays up while any other thread is still attached (ProviderSessionManager.ts).CodexAdapterV2sendsthread/startandthread/resumeand neverthread/unsubscribe. The generated client already has that method; the response status isnotLoaded|notSubscribed|unsubscribed, which is the unload signal. An optionalunloadThread, implemented for Codex withthread/unsubscribeand called when detaching from a shared runtime, matches the leak. I did not reproduce the 80-process / 9 GB measurement or the 11-to-0 MCP check; those are yours.5. Failed turn stays "waiting" while background work is pending (v2) — confirmed
shellRuntimeforcesruntime.statusto"idle"wheneverpendingBackgroundTasksis nonempty, including when the shell status is"failed"(packages/client-runtime/src/state/models.ts).resolveSidebarThreadStatusreturns"waiting"for"idle"before it looks at"failed"(Sidebar.logic.ts).ThreadNotificationCoordinatoronly promotes a"ready"status whenlatestRun.status === "failed", so an idle/waiting failed thread never becomes the failure toast either.On
main, the resolver checks a failed session before background liveness ("A failed session outranks lingering background liveness"). This branch inverted that. The existing test expects idle plus a leftoverlastErrorto stay waiting, so the check should be "idle, and the latest run failed", not "any persisted lastError".resolveSidebarThreadStatusdoes not currently takelatestRun; both the sidebar and the coordinator pass a shell that has it, so adding it to the pick fixes both. Idle with background work and a non-failed latest run should stay waiting.- addedacceptedfeature request acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 24, 2026 Items 1, 4 and 5 have now been addressed
letrandat commented
on Oct 3, 2026 More actionsAdding a current observation of item 2: Lineage / “Previous agents” still shows “Done” while the same delegated worker is running follow-up work.
Environment: T3 Nightly 0.0.46-nightly.20261003.2610 on macOS; installed bundle embeds commit 8ed276c. A native Claude coordinator delegates to a Codex worker.
Observed sequence:
- The coordinator starts a worker thread. Its initial run completes, and the Lineage entry shows “Done”.
- The coordinator sends follow-up work to that same worker thread.
- During the follow-up,
t3_thread_waitreports the target run asrunning, with no pending approval, and local output files continue updating. - The coordinator's Lineage / “Previous agents” entry still shows “Done”.
Reusing the existing worker thread is expected. The misleading part is the status: it appears to retain the initial run's completion while later work is active. The UI should reflect the latest run/current activity, or clearly distinguish historical completion from current activity.
This was observed in an existing session. I have not established a separate clean reproduction or independently verified the source-level cause or a fix. No sensitive logs or screenshots are attached.
Confirmed item 2 with a rate-limit recovery trigger on T3 Desktop Nightly 0.0.46-nightly.20261003.2610, commit 8ed276c, macOS arm64, Claude Agent SDK.
After Claude hit its rate limit, the user waited for reset and sent a new message to resume. A read-only database check showed:
- A3 and G5: second runs running, parent subagent records still failed.
- G4: second run completed, parent subagent record still failed.
The parent's task queries returned the earlier rate-limit failure alongside hasPendingChildRuns: true.
At this installed commit, packages/client-runtime/src/state/threadRelationships.ts:99 uses subagent.status for the edge, and apps/web/src/components/chat/ThreadRelationshipsControl.tsx:243 counts running agents from those records.
Expected: Lineage reflects the child's current run while retaining the original failure as history.
Related proposed fix: #12977.
Diagnosed by GPT-6 via Codex.
I’m seeing another trigger for item 2: manually stopping and resuming a delegated subagent.
Steps to reproduce
- Have a parent thread delegate work to a subagent.
- Check the parent thread’s Lineage: the subagent shows as running.
- Stop the subagent. Its Lineage status changes to Stopped.
- Resume the same subagent and confirm it is actively working again.
- Return to the parent thread’s Lineage.
Expected behavior
The subagent’s Lineage status returns to Running when it resumes.
Actual behavior
The subagent resumes work, but the parent thread’s Lineage continues to show Stopped.
Environment
Ubuntu 26.04
I run a personal fork on top of the orchestrator v2 PR (#2829) and hit five bugs. Four are in #2829's code; one (3) is on
main. I checked each against #2829's current head (3e4ca4c532) andmain(effaab94e3), and they're still there. I have small, tested fixes for each and am happy to open focused PRs (against #2829's branch for 1, 2, 4, 5 and againstmainfor 3) if you'd like them. Just say which, so you don't have to re-fix them yourselves.1. Interrupted threads never resume after a restart (v2)
With "Continue threads after restarts" on, a thread that was mid-turn when the server restarted is cancelled and waits for a manual "continue".
restartContinuationRun()inapps/server/src/orchestration-v2/RestartContinuation.tsskips any run whose provider session status isn't"running". But only the OpenCode adapter ever writes"running"; the sharedProviderSessionManagerwritesready/stopped. Everywhere else v2 treats a session as live when it's notstopped/error(Orchestrator.ts,ProjectionStore.ts). Fix: checksession.status === "stopped" || session.status === "error"instead (3 lines plus a test).2. A steered subagent shows as finished in Lineage while it runs again (v2)
After a subagent finishes and is sent new work, Lineage still shows it as completed and leaves it out of the "N running" count.
The parent's subagent record settles when the spawning run ends, and a new run on the child thread doesn't reopen it.
deriveThreadRelationshipGraph()inpackages/client-runtime/src/state/threadRelationships.tsusessubagent.statusfor the edge, andThreadRelationshipsControl.tsxcounts running agents from the records, while thread nodes already usethread.activityRunStatus. Fix: prefer the child thread's liveactivityRunStatusfor subagent edges and the count, falling back to the record.Related but different trigger: #7314.
3. A hung page makes
preview_navigateoutlive its deadline and disconnects the browser host (main)When a page never finishes loading,
preview_navigate(orpreview_openreusing a tab) times out. The broker then disconnects the whole client's automation connection, and the agent's current tab is forgotten, so agents fall back to opening new tabs.In
apps/web/src/components/preview/PreviewAutomationHosts.tsx, theopen(reuse) andnavigatecases awaitbridge.navigate(...)with no limit (on desktop it settles only whenwebContents.loadURLdoes), then givewaitForNavigationReadinessthe full request timeout instead of what's left beforehostDeadlineMs. Every other operation respects the host budget from #4685.PreviewAutomationBroker.tstreats the unanswered request as a dead connection. Fix: race the navigate call against the host deadline and pass the remaining budget to the readiness wait.Related: #12407 (same eviction mechanism,
preview_wait_for).4. Disconnecting a Codex session leaves the thread loaded, with its MCP servers running (v2)
All Codex threads share one
codex app-server. Disconnecting a thread detaches it in T3, but the native thread stays loaded in the app-server along with its MCP servers (e.g. one language server per thread or subagent). Over a day I measured about 80 idle processes and 9 GB.ProviderSessionManager.ts's detach path for multi-thread runtimes never tells the runtime to unload the native thread; the adapter interface has no unload operation, andCodexAdapterV2.tsnever sendsthread/unsubscribe, although the generated client supports it. Fix: an optionalunloadThreadon the session runtime, implemented for Codex withthread/unsubscribeand called after detaching from a shared runtime. Verified live: 11 MCP processes went to 0.5. A failed turn shows as "waiting" while background work is still pending (v2)
If a turn fails while subagents or background tasks are still pending, the sidebar and the notification coordinator show the thread as waiting, not failed, until the background work drains.
resolveSidebarThreadStatusinapps/web/src/components/Sidebar.logic.tsreturns"waiting"forruntime.status === "idle"before checking for failure, and the runtime parks atidlewhile background work is pending. Fix: treatidlewith a failed latest run as failed, checked before the idle branch.