Skip to content

[Bug]: no SSE fallback after an established Codex WebSocket dies mid-turn (prelude-timeout half resolved) #4191

Description

@mehmeemow

Client or integration

Codex App (also reproduced from Codex CLI when resuming the same problematic thread).

Area

Streaming

Summary

A long-running Codex conversation became effectively unusable only while traffic was routed through OpenCodex. The same account, model, machine, repository and thread work normally again as soon as the OpenCodex proxy is disabled/bypassed.

This does not look like a generic network outage because the failure is both thread-specific and proxy-path-specific.

With OpenCodex enabled, the problematic parent thread repeatedly failed with either:

stream disconnected before completion:
Upstream stream terminated unexpectedly:
codex websocket closed before a Responses terminal event
(close 1006 Connection ended)

or:

stream disconnected before completion:
Upstream stream terminated unexpectedly:
codex websocket response prelude timed out

After OpenCodex was disabled/bypassed, the same existing thread resumed and worked normally immediately, without fork, rollback, or compaction recovery.

Expected behavior: OpenCodex should transparently proxy the same Codex Responses request, or safely fall back to HTTP/SSE when the upstream WebSocket path cannot complete.

This appears related to one or more WebSocket relay edge cases such as large replay payloads, loss of upstream close code/reason, the response-prelude timeout, or lack of safe HTTP/SSE fallback after an already-open WebSocket terminates.

Related issues:

This report adds a real-world A/B observation: the exact same previously failing thread becomes healthy immediately when OpenCodex is bypassed.

Reproduction

  1. Run Codex Desktop on Windows with Codex traffic routed through OpenCodex to the normal ChatGPT/Codex upstream.
  2. Open/resume an existing long-running conversation.
  3. Observe repeated stream disconnected before completion failures for many hours.
  4. Switch the host network to a mobile hotspot. Observe the same failure.
  5. On the same machine, start Codex CLI and create a fresh conversation. The fresh conversation replies normally (Hello!).
  6. In Codex CLI, resume the same problematic existing conversation. The failure reproduces.
  7. While the parent conversation is failing, previously spawned sub-agents can continue executing normally.
  8. Observe one or both of the errors shown in the Logs section below.
  9. Disable/bypass OpenCodex without otherwise changing the account, model, repository, machine, or conversation.
  10. Resume the same existing problematic thread again.
  11. The thread now works normally.

A/B matrix:

Test Result
Codex Desktop + OpenCodex + existing long thread ❌ stream disconnected
Same path on mobile hotspot ❌ same failure
Codex CLI + OpenCodex + fresh thread ✅ normal reply
Codex CLI + OpenCodex + same existing long thread ❌ stream disconnected
Sub-agents under the problematic conversation ✅ continued executing
Same existing long thread with OpenCodex bypassed ✅ works normally

Suggested diagnostics for reproducing this in OpenCodex:

  • log serialized first-message/request size in bytes;
  • preserve/log the actual upstream WebSocket close code and reason;
  • log elapsed time from send() to quota/metadata control events and to the first non-control Responses event such as response.created;
  • record whether any downstream Responses bytes/events were emitted before failure;
  • record whether the exchange was eligible for a safe HTTP/SSE fallback.

Potential mitigations:

  1. Preserve and expose the real upstream WS close code/reason instead of collapsing it into a generic downstream 1006-style failure.
  2. Preflight large initial WS payloads and route to HTTP/SSE before opening WS when near the upstream size limit.
  3. If WS opens but closes/times out before any real Responses event is delivered downstream, use HTTP/SSE fallback where it is safe and cannot duplicate inference.
  4. Revisit or make configurable the response-prelude deadline for large/slow requests.

Version

OpenCodex v2.49.0

Codex Desktop 26.903.61454

Operating system

Windows 11

Provider and model

Canonical ChatGPT/Codex authenticated upstream. Exact model identifier was not captured as part of the initial incident.

Logs or error output

stream disconnected before completion:
Upstream stream terminated unexpectedly:
codex websocket closed before a Responses terminal event
(close 1006 Connection ended)

and later:

stream disconnected before completion:
Upstream stream terminated unexpectedly:
codex websocket response prelude timed out

Screenshots and supporting files

No private conversation content, repository content, credentials, or account identifiers are attached.

The strongest supporting evidence is the A/B result: the same existing long thread fails through OpenCodex and immediately succeeds when OpenCodex is bypassed.

Redacted configuration

{
  "client": "Codex Desktop / Codex CLI",
  "proxy_path": "Codex -> local OpenCodex -> canonical ChatGPT/Codex upstream",
  "control_test": "Codex -> canonical ChatGPT/Codex upstream (OpenCodex bypassed)",
  "result_with_proxy": "stream disconnected",
  "result_without_proxy": "same thread works normally"
}

No tokens, account identifiers, request credentials, private conversation text, or repository contents are included.

Checks

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.

Activity

  1. github-actions commented on Sep 10, 2026

    @github-actions
    Contributor

    Issue reopened

    The report now contains the information required by the automated check. Thanks for updating it.

  2. changed the title [-][Bug] Long Codex thread fails only through OpenCodex proxy (WS 1006 / response prelude timeout); bypass works immediately[/-] [+][Bug]: Long Codex thread fails only through OpenCodex proxy (WS 1006 / response prelude timeout); bypass works immediately[/+] on Sep 10, 2026
  3. lidge-jun commented on Sep 10, 2026

    @lidge-jun
    Owner

    리뷰 · 우선순위 74 / 80

    설명

    이 이슈는 “긴 Codex 대화가 OpenCodex 프록시를 거칠 때만 깨지고, 같은 계정·모델·머신·저장소·스레드에서 프록시를 끄면 바로 정상”이라는 A/B 관측입니다. 데스크톱이든 CLI든, 문제 있는 기존 스레드만 실패하고 새 스레드는 됩니다. 에러 문구도 리포에 있는 실제 문자열과 맞습니다. 하나는 codex websocket closed before a Responses terminal event (close 1006 Connection ended)이고, 다른 하나는 codex websocket response prelude timed out입니다.

    지금 dev HEAD(12c248f52, package 2.50.0)의 Codex 업스트림 WebSocket 경로는 이렇게 동작합니다. src/server/responses/fetch-helpers.ts의 providerFetch가 ChatGPT-family 스트리밍을 codexWsUpstreamFetch로 태우고, 업그레이드 전·전송 실패 전에는 HTTP SSE로 넘어갈 수 있습니다. 하지만 이미 ws.send(frameText)가 성공한 뒤에는 src/server/responses/codex-ws-exchange.ts의 failStream이 바디 실패로 끝내며 SSE로 다시 보내지 않습니다. 이유는 주석에 적혀 있습니다. 업스트림에서 이미 턴이 시작됐을 수 있어서, 재전송하면 추론이 두 번 돌 수 있기 때문입니다.

    관련으로 닫힌 이슈도 있습니다. #2471은 16 MiB create-frame 경계와 SSE 폴백을 다뤘고, src/server/responses/codex-ws-wire.ts에 MAX_CODEX_WS_CREATE_FRAME_BYTES와 codexWsCreateFrameExceedsLimit가 남아 있습니다. #4083/#3976은 response-prelude 타임아웃을 다루었고, 현재 상수는 CODEX_WS_RESPONSE_PRELUDE_TIMEOUT_MS = 90_000입니다(예전 30초보다 깁니다). close code도 closedBeforeTerminalMessage가 메시지에 붙입니다. 그런데도 이 리포트는 “프록시 ON이면 장시간 실패, OFF면 같은 스레드가 즉시 회복”이라서, 이미 머지된 완화만으로는 설명이 안 끝나는 실사용 회귀/잔여 구멍으로 보입니다. 긴 스레드는 create-frame이 커지고, 프록시 경로의 WS만 그 크기를 타고, 업스트림이 조용히 끊거나(1006) prelude 안에 의미 있는 Responses 이벤트가 안 오면 클라이언트가 반복 실패하는 그림과 잘 맞습니다.

    라인 레벨로 보면 지금 코드는 “진단·프리플라이트”는 있고 “이미 연 WS가 prelude/조기 close로 죽으면 안전한 SSE 폴백”은 아직 없습니다. 리포터 제안(첫 메시지 바이트 로그, close code/reason 보존, send→첫 Responses 이벤트 경과 시간, 다운스트림 바이트 유무, SSE 폴백 자격)은 재현 없이도 다음 패치의 측정 축이 됩니다.

    src/server/responses/codex-ws-exchange.ts 라인 214 - prelude 타이머가 만료되면 failStream("codex websocket response prelude timed out")만 하고 SSE로 내려가지 않는다. 리포트의 두 번째 에러와 정확히 같은 문자열이다.
    src/server/responses/codex-ws-exchange.ts 라인 138-146 - failStream은 이미 sent면 200 바디 실패로 고정한다. 이중 추론을 막으려는 설계지만, “Responses 이벤트가 한 번도 안 나간 조기 실패”와 “이미 일부 이벤트를 보낸 실패”를 같은 길로 묶는다.
    src/server/responses/codex-ws-wire.ts 라인 5 - prelude 한도는 지금 90초 고정이다. #3976은 닫혔지만 설정으로 늘리는 스위치는 이 HEAD 스냅샷만으로는 보이지 않는다.
    src/server/responses/codex-ws-wire.ts 라인 115-125 - close code/reason은 메시지 문자열에는 붙지만, 주석대로 request log의 타입 코드로는 안 남는다. 장시간 현장 디버깅이 어렵다.
    src/server/responses/codex-ws-wire.ts 라인 137-144 - create-frame이 한도 이상이면 SSE로 보내도록 프리플라이트가 있다. 한도 바로 아래의 긴 스레드(리플레이·스크린샷 누적)는 여전히 WS를 타고, 그때의 조기 close/prelude 실패는 폴백이 없다.

    메인테이너의 판단이 필요한 지점

    • “아직 Responses 이벤트가 하나도 다운스트림으로 안 나간 상태”에서만 SSE 폴백을 허용할지, 아니면 이중 추론 위험 때문에 WS 실패는 항상 바디 실패로 둘지.
    • prelude 90초를 더 늘리거나 설정화할지, 아니면 긴 스레드는 create-frame 크기/예상 TTFT로 WS 자격을 더 좁힐지.
    • 이 이슈를 #2471/#4083 후속 측정 이슈로 두고 재현 로그(프레임 바이트, close code, prelude 경과)를 필수 재현 조건으로 요구할지.
    • Windows Codex Desktop 2.49.0 사용자 경로와 dev 2.50.0 HEAD의 동작 차이를 릴리즈 노트에 따로 적을지.

    너의 추천
    열어 두고 streaming 우선순위로 둔다. 첫 패치는 관측부터다. create-frame 바이트, close code/reason, send부터 첫 non-control Responses 이벤트까지 시간, 다운스트림 바이트 유무를 로그에 남긴 뒤, “sent지만 아직 responseCommitted/다운스트림 이벤트가 없는 조기 실패”만 안전한 HTTP SSE 폴백 후보로 설계한다. 설정 없는 90초 고정만 더 늘리는 것은 재현 없이 하지 말자. 중복 이슈로 닫지 말고 #2471/#4083/#3976을 Related로 링크한 채 잔여 구멍으로 추적한다.

    이 댓글은 grok-bot이 작성했습니다

  4. Ingwannu commented on Sep 10, 2026

    @Ingwannu
    Owner

    One important constraint on the proposed fallback: at dev 6d3ad12, codex-ws-exchange.ts::failStream deliberately treats a successfully sent create frame as possibly executing upstream. No downstream Responses event yet is not proof the upstream did not accept or execute it. Please do not enable an automatic HTTP resend based only on responseCommitted === false or zero downstream bytes; that can duplicate the turn.

    The A/B report is worth keeping open. The next useful evidence is content-free stage information: create-frame byte count, whether send completed, close code, and elapsed time to the first meaningful event, plus exact OCX/client versions. Do not attach conversation bodies, authorization headers, cookies, or raw close reasons containing request content. Preserve the current no-replay-after-send contract unless the upstream provides affirmative non-execution/idempotency evidence.

  5. Rizek000 commented on Sep 12, 2026

    @Rizek000

    I hit a similar disconnect on macOS with OpenCodex 2.52.0-preview.20260911 and Codex CLI 0.154.0-alpha.6.2.

    The first frame arrived after 1.1 seconds, then the connection closed with 1006 after 56.1 seconds. The error reported after-response-started, sent=yes, and 134 relayed events. That points to a mid-response disconnect rather than the prelude timeout described above.

    The matching proxy log shows HTTP 502, one send, and no recovery attempt. Eight subsequent requests completed successfully. I haven't tested bypassing the proxy on the same failing thread, so I can't confirm it's the same underlying issue.

    One potentially useful detail: websockets: false was already set. In this version, that disables the client-to-proxy WebSocket path, but the proxy still uses WebSockets upstream to ChatGPT.

    Recording whether the socket was reused, the outgoing frame size, and the gap between the last frame and the close would help narrow this down. An explicit HTTP/SSE upstream option would also make a controlled comparison easier. I'd keep the no-replay-after-send behavior intact, since this request had already started producing output.

  6. 17 remaining items

  7. lidge-jun commented on Sep 21, 2026

    @lidge-jun
    Owner

    The per-request resend budget landed on dev in 3a0718c, and it does not reach this report.

    That change bounds how many times one logical request may replace a send once the composed recovery legs are in play — reset, prelude EOF, credential refresh, combo candidate. What it deliberately does not do is give a dead relay a second transport: an established WebSocket that stops mid-turn is still a failed leg, not a fallback to SSE, and the prelude timeout is a separate decision about when a turn is declared unanswered. Both remain as reported.

    Remaining scope: the mid-turn transport death needs a fallback path and a decision about what a partially delivered turn owes the client, and the prelude timeout needs to say whether it is the same failure class. Neither is covered by anything landed today.

  8. FredAmartey commented on Sep 23, 2026

    @FredAmartey
    Contributor

    Stage records from failing turns, as you asked. Canonical ChatGPT backend, macOS, direct connection with no proxy, two ChatGPT accounts, mostly gpt-6-astra. OpenCodex 2.57.0 to 2.63.0, September 17 to 22.

    12,587 Codex WS turns carry a stage record. 53 of them died the same way: the create frame went out, then close 1006 with zero upstream frames and zero relayed events, 174 ms to 4.3 s after the send. Every one was on a fresh socket (reused: false) and settled as a 502 upstream_server_error. They split evenly across the two accounts (26 and 27) and fell in 8 threads, 38 of them in the two largest.

    Size seems to matter, but it is not a hard limit. requestBytes is null on successful turns, so I used input tokens as the size measure. For a failed turn that is the input tokens of the retry that followed it.

    input tokens WS turns that completed early 1006 rate
    under 100k 3,061 5 0.2%
    100k to 200k 6,741 26 0.4%
    200k to 300k 2,652 19 0.7%

    The failed create frames were 0.2 MB to 14.5 MB, median 6.6 MB, all under the 16 MiB limit from #2471. The same threads completed thousands of WS turns in between.

    The part that might help this issue: I run a small local patch that leaves the failed request alone. After a transport failure like this, it sends the next requests for that account and route over HTTP/SSE for 60 seconds, then lets one request probe the WebSocket again before switching back. The next request in each thread, Codex retrying the failed turn one to eleven seconds later, went over HTTP in 52 of 53 cases. 49 got a 200 and 3 got a 429. The one that went back over WS also completed. Without the patch, shouldUseCodexWsUpstream sends that retry straight back over WS with the same frame. That fits the original report: a long thread that only recovers once the proxy is bypassed.

    The proxy never sends anything twice here, so this sidesteps the double-inference question. The client decides to retry, and the proxy only picks the transport for that new request. Would you take a PR for it against dev? The same-request fallback can stay its own discussion. The patch also has an opt-in same-request HTTP retry, but it never fired in these records, so I'm not proposing it.

  9. q485707079-source commented on Sep 23, 2026

    @q485707079-source

    Independent corroboration (Windows, 2.63.0) + a verified way to reach the existing sseFallback today

    Adding a second environment to @FredAmartey's data, plus something I could not find in this thread: a deterministic way to force the upstream leg onto SSE while the proper fallback is still being designed.

    Environment

    Windows, Codex Desktop, opencodex 2.63.0, ChatGPT backend, canonical Responses requests in the 2.2–2.7 MB range, model_reasoning_effort = "max".

    Codex is pointed at the local proxy over plain HTTP (openai_base_url = http://127.0.0.1:10100/backend-api/codex) and the client-facing WebSocket is already disabled (websockets: false → 426 → clean HTTP fallback, working as designed). The close 1006 failures still occur, which independently confirms they originate on the upstream lane, not on the client-facing upgrade.

    Stage records from failing turns

    request=2188716B  frames=193  control=2  relayed=191  first-frame=662ms  elapsed=21530ms  close=1006
    request=2247292B  frames=13   control=2  relayed=11   first-frame=609ms  elapsed=66684ms  close=1006
    request=2614871B  frames=5    control=2  relayed=3    first-frame=499ms  elapsed=4663ms   close=1006
    request=2712862B  frames=21   control=2  relayed=19   first-frame=458ms  elapsed=21483ms  close=1006
    

    pings=0 pongs=0 on all four. Each settled as 502 upstream_server_error with resendPermission: "refused-committed".

    Two things worth flagging:

    • The elapsed=4663ms sample dies far too fast to be an idle/keepalive timeout — consistent with your "174 ms to 4.3 s" set, so ping tuning cannot address this.
    • Because the turn is refused for resend, Codex's own retry loop then reports exceeded retry limit, last status: 429 Too Many Requests. A lot of the "429" reports around this are probably downstream of this bug rather than a quota problem.

    A verified way to force the SSE path today

    Setting the global proxy to a SOCKS5 URL makes resolveProxyRoute return fallback, which routes the turn through the existing sseFallback. Three lines in the shipped source make this deterministic:

    • src/config/proxy-env.ts:171-175 — a socks5:// value deletes HTTP_PROXY/HTTPS_PROXY and sets only ALL_PROXY
    • src/lib/proxy-env.ts:66-73 — resolveProxyRoute returns { kind: "fallback" } when the proxy scheme is not http(s)
    • src/server/responses/ws-upstream.ts:160 — if (proxyRoute.kind === "fallback") return sseFallback(url, init);
    // ~/.opencodex/config.json
    "proxy": "socks5://127.0.0.1:7897"   // any SOCKS5 endpoint reachable by the client

    then ocx stop && ocx service start.

    Verified by calling resolveProxyRoute(new URL("wss://chatgpt.com/backend-api/codex")) against the shipped source:

    HTTPS_PROXY=http://127.0.0.1:7897  ->  {"kind":"proxy","proxy":"http://127.0.0.1:7897"}
    ALL_PROXY=socks5://127.0.0.1:7897  ->  {"kind":"fallback"}
    ALL_PROXY=http://127.0.0.1:7897    ->  {"kind":"proxy","proxy":"http://127.0.0.1:7897"}   // control
    

    The third row is the control: the fallback comes from the socks5 scheme, not merely from the value arriving via ALL_PROXY.

    End-to-end check: after restarting with the SOCKS5 value, the next real turn recorded no codexWsStage field at all and completed with 200, and service.log confirms the mode at startup (outbound proxy: socks5://...).

    The cost is the TTFT the WS lane exists to win back (~1.0 s vs ~3.9 s p50 per the note in ws-upstream.ts), which seems like a fair trade for anyone whose turns currently die mid-stream.

    Suggestion

    It would help if this escape hatch were a supported, documented knob instead of a side effect of the proxy scheme — e.g. transport: "auto" | "sse", or having websockets: false cover the upstream lane as well. Today the only route to the already-implemented sseFallback is changing the proxy scheme, which is not discoverable and silently costs TTFT for every provider, not just the Codex one.

    Thanks for the detailed triage in this thread — the "failed leg is not a fallback" framing is exactly what made this reproducible.

  10. added
    priority: P1High: reproducible failure in a core path (routing, failover, account pool, streaming, usage, auth,
    on Sep 24, 2026
  11. mehmeemow commented on Sep 24, 2026

    @mehmeemow
    Author

    Today's overtime message: unexpected status 502 Bad Gateway: codex websocket closed before a Responses terminal event (close 1011 keepalive ping timeout) [cause=no-upstream-frame request=5666621B sent=yes frames=0 control=0 relayed=0 first-frame=n/a elapsed=82924ms pings=5 pongs=0], url: http://127.0.0.1:10100/v1/responses

  12. devin-ai-integration commented on Sep 25, 2026

    @devin-ai-integration
    Contributor

    New matched pull request: #5825 (would resolve this issue) — fix: P0 #5369 spill growth + P1 bug batch [priority: P0]

  13. added a commit that references this issue on Sep 25, 2026
  14. lidge-jun commented on Sep 25, 2026

    @lidge-jun
    Owner

    #5825 (22b22ae) adds only a diagnostic: WS stage records and failure messages now include first-response= (time to the first non-control Responses event). It does not add the SSE fallback after a WebSocket dies mid-turn that this issue asks for, so the issue stays open.

  15. lidge-jun commented on Sep 27, 2026

    @lidge-jun
    Owner

    Closing for the case this issue's data shows: a Codex WebSocket that dies after the create frame and before the first Responses event now falls back to one HTTP/SSE replacement under retryOnReset. That behavior landed in #5675 (aed3bb8f42), and #6011, carried in #6059 (merge 8923ad9835), pins it with tests and ADR-4191. A socket that dies after output has started stays a failed leg on purpose, because replaying it could run a second inference. If you still see failures after the first event, please open a new issue with the stage record.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingpriority: P1High: reproducible failure in a core path (routing, failover, account pool, streaming, usage, auth,streamingSSE, WebSocket, terminal stream frames

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions