Repository navigation
fix(server): reopen Codex sessions whose app-server exited - #990
Merged
Merged
Conversation
A Codex app-server that exits between turns left the v2 provider session `ready` with a dead client, so every later message on the thread failed with "The provider could not start this turn" until idle release or a restart. The Codex client now exposes `awaitTermination`, and the v2 adapter fails its event stream once queued events drain, so the session manager releases the session as runtime_error and the next message opens a fresh app-server that resumes the native thread. Adds replay regression coverage for the pre-v2 incident (a non-retryable model-at-capacity failure followed by a message on the same thread), for an in-process session error, and for a session persisted as error across a restart. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
This was referenced Oct 3, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
On the pre-v2 orchestrator, a Codex runtime error ("Selected model is at capacity. Please try a different model.") put the provider session into
error. After that, every user message on the thread was rejected with "Provider session '' is error; it cannot accept a new turn.", so the thread stayed stuck for good.I added regression tests to check whether orchestrator v2 recovers from this. Two of the three cases already recovered. The third did not:
readywith a dead client, and every later message failed with "The provider could not start this turn". That lasted until idle release or a server restart. The session manager'sruntime_errorrecovery never ran, because it only runs when the event stream dies.Fix
effect-codex-app-server: the client now exposesawaitTermination, which resolves with the termination error once the transport ends.CodexAdapterV2: when the client terminates, the session event queue fails after already-queued events drain, so a turn's last events (such asturn/completed) still arrive. The session manager then releases the session asruntime_error. The next message opens a fresh app-server, which sendsthread/resumeon the same native thread.Tests
Added to
apps/server/src/orchestration-v2/testkit/OrchestratorReplayRecovery.integration.test.ts. They use the existing Codex replay harness and wait on receipts (the stored-event stream, the replay driver cursor), not sleeps.multi_turntranscript with the first turn rewritten to a non-retryableerrornotification followed by a failedturn/completed. Asserts:failedthencompleted.thread/startorthread/resumewas sent.provider_thread_resumerecording, where the first app-server exits after turn 1. The test waits for theprovider-session.updatedevent with statuserror, sends the next message, and asserts that the session reopened,thread/resumeran, and both runs completed on one native thread. Before the fix it timed out waiting for the session error, because the session stayedready.erroracross a restart. Same as case 2, but the second message runs in a new runtime on the same SQLite database with startup recovery enabled. Asserts the session is stillerrorafter recovery, then reopens and resumes it. Before the fix it timed out the same way.Also added a package test showing
awaitTerminationreports the app-server exit.The
serverOverloadedcode in the synthetic capacity frame is an assumption about Codex's wire code; the adapter classifies any string code the same way (provider_error).Validation: focused
vp test run(the recovery suite; the neighbouring Codex adapter, replay fixture, session manager, fork, merge-back, provider-switch and selection-restart suites; and theeffect-codex-app-servertests), scopedvp lint/vp fmt --check, andvp run -F t3 typecheck/vp run -F effect-codex-app-server typecheck. One test fails both with and without this change:claude_result_is_error/claudeAgentreads the localCLAUDE_CONFIG_DIR.🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.