Skip to content

fix(server): reopen Codex sessions whose app-server exited - #990

Merged
rynfar merged 1 commit into
pylonfrom
test/v2-codex-turn-error-recovery
Oct 3, 2026
Merged

rynfar merged 1 commit into
pylonfrom
test/v2-codex-turn-error-recovery

Conversation

@rynfar

@rynfar rynfar commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

On the pre-v2 orchestrator, a Codex runtime error ("Selected model is at capacity. Please try a different model.") put the provider session into error. After that, every user message on the thread was rejected with "Provider session '' is error; it cannot accept a new turn.", so the thread stayed stuck for good.

I added regression tests to check whether orchestrator v2 recovers from this. Two of the three cases already recovered. The third did not:

  • Turn fails with model-at-capacity: already fine. The run fails, the session stays usable, and the next message starts a turn on the same native Codex thread.
  • Codex app-server exits between turns: was stuck. The Codex adapter's event stream never ended when its app-server exited. The session stayed ready with a dead client, and every later message failed with "The provider could not start this turn". That lasted until idle release or a server restart. The session manager's runtime_error recovery never ran, because it only runs when the event stream dies.

Fix

  • effect-codex-app-server: the client now exposes awaitTermination, which resolves with the termination error once the transport ends.
  • CodexAdapterV2: when the client terminates, the session event queue fails after already-queued events drain, so a turn's last events (such as turn/completed) still arrive. The session manager then releases the session as runtime_error. The next message opens a fresh app-server, which sends thread/resume on the same native thread.

Tests

Added to apps/server/src/orchestration-v2/testkit/OrchestratorReplayRecovery.integration.test.ts. They use the existing Codex replay harness and wait on receipts (the stored-event stream, the replay driver cursor), not sleeps.

  1. Model at capacity. Uses the recorded multi_turn transcript with the first turn rewritten to a non-retryable error notification followed by a failed turn/completed. Asserts:
    • The runs end failed then completed.
    • The error item shows the capacity message.
    • Both runs share one provider thread.
    • The session was attached once.
    • The transcript was consumed exactly, so no extra thread/start or thread/resume was sent.
    • This test passes without the fix.
  2. In-process session error. Uses the provider_thread_resume recording, where the first app-server exits after turn 1. The test waits for the provider-session.updated event with status error, sends the next message, and asserts that the session reopened, thread/resume ran, and both runs completed on one native thread. Before the fix it timed out waiting for the session error, because the session stayed ready.
  3. Persisted error across a restart. Same as case 2, but the second message runs in a new runtime on the same SQLite database with startup recovery enabled. Asserts the session is still error after recovery, then reopens and resumes it. Before the fix it timed out the same way.

Also added a package test showing awaitTermination reports the app-server exit.

The serverOverloaded code in the synthetic capacity frame is an assumption about Codex's wire code; the adapter classifies any string code the same way (provider_error).

Validation: focused vp test run (the recovery suite; the neighbouring Codex adapter, replay fixture, session manager, fork, merge-back, provider-switch and selection-restart suites; and the effect-codex-app-server tests), scoped vp lint / vp fmt --check, and vp run -F t3 typecheck / vp run -F effect-codex-app-server typecheck. One test fails both with and without this change: claude_result_is_error/claudeAgent reads the local CLAUDE_CONFIG_DIR.

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

A Codex app-server that exits between turns left the v2 provider session
`ready` with a dead client, so every later message on the thread failed
with "The provider could not start this turn" until idle release or a
restart. The Codex client now exposes `awaitTermination`, and the v2
adapter fails its event stream once queued events drain, so the session
manager releases the session as runtime_error and the next message opens
a fresh app-server that resumes the native thread.

Adds replay regression coverage for the pre-v2 incident (a non-retryable
model-at-capacity failure followed by a message on the same thread), for
an in-process session error, and for a session persisted as error across
a restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Oct 3, 2026
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 4.9 KiB 4.9 KiB 0 B (0.0%) 6.8 KiB ✅
Codex Thread snapshot wire 3.7 KiB 3.7 KiB 0 B (0.0%) 4.9 KiB ✅
Codex Live turn WebSocket wire 1.2 KiB 1.2 KiB 0 B (0.0%) 2.0 KiB ✅
Codex Live turn WebSocket decoded 20.4 KiB 20.4 KiB 0 B (0.0%) 29.3 KiB ✅
Codex Live turn messages 2 2 0 (0.0%) 8 ✅
Claude Total thread wire 4.9 KiB 4.9 KiB 0 B (0.0%) 6.8 KiB ✅
Claude Thread snapshot wire 3.7 KiB 3.7 KiB 0 B (0.0%) 4.9 KiB ✅
Claude Live turn WebSocket wire 1.2 KiB 1.2 KiB 0 B (0.0%) 2.0 KiB ✅
Claude Live turn WebSocket decoded 20.8 KiB 20.8 KiB 0 B (0.0%) 29.3 KiB ✅
Claude Live turn messages 2 2 0 (0.0%) 8 ✅

Baseline: 5f1af34 · PR result: 497be7e · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@rynfar
rynfar merged commit b4c210e into pylon Oct 3, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant