Repository navigation
[Bug]: no SSE fallback after an established Codex WebSocket dies mid-turn (prelude-timeout half resolved) #4191
Description
Activity
- addedstreamingSSE, WebSocket, terminal stream framesSSE, WebSocket, terminal stream frames
on Sep 10, 2026 github-actions commented
on Sep 10, 2026 on Sep 10, 2026 – with GitHub ActionsContributorMore actionsIssue reopened
The report now contains the information required by the automated check. Thanks for updating it.
- changed the title
[-][Bug] Long Codex thread fails only through OpenCodex proxy (WS 1006 / response prelude timeout); bypass works immediately[/-][+][Bug]: Long Codex thread fails only through OpenCodex proxy (WS 1006 / response prelude timeout); bypass works immediately[/+]on Sep 10, 2026 리뷰 · 우선순위 74 / 80
설명
이 이슈는 “긴 Codex 대화가 OpenCodex 프록시를 거칠 때만 깨지고, 같은 계정·모델·머신·저장소·스레드에서 프록시를 끄면 바로 정상”이라는 A/B 관측입니다. 데스크톱이든 CLI든, 문제 있는 기존 스레드만 실패하고 새 스레드는 됩니다. 에러 문구도 리포에 있는 실제 문자열과 맞습니다. 하나는
codex websocket closed before a Responses terminal event (close 1006 Connection ended)이고, 다른 하나는codex websocket response prelude timed out입니다.지금
devHEAD(12c248f52, package2.50.0)의 Codex 업스트림 WebSocket 경로는 이렇게 동작합니다.src/server/responses/fetch-helpers.ts의providerFetch가 ChatGPT-family 스트리밍을codexWsUpstreamFetch로 태우고, 업그레이드 전·전송 실패 전에는 HTTP SSE로 넘어갈 수 있습니다. 하지만 이미ws.send(frameText)가 성공한 뒤에는src/server/responses/codex-ws-exchange.ts의failStream이 바디 실패로 끝내며 SSE로 다시 보내지 않습니다. 이유는 주석에 적혀 있습니다. 업스트림에서 이미 턴이 시작됐을 수 있어서, 재전송하면 추론이 두 번 돌 수 있기 때문입니다.관련으로 닫힌 이슈도 있습니다.
#2471은 16 MiB create-frame 경계와 SSE 폴백을 다뤘고,src/server/responses/codex-ws-wire.ts에MAX_CODEX_WS_CREATE_FRAME_BYTES와codexWsCreateFrameExceedsLimit가 남아 있습니다.#4083/#3976은 response-prelude 타임아웃을 다루었고, 현재 상수는CODEX_WS_RESPONSE_PRELUDE_TIMEOUT_MS = 90_000입니다(예전 30초보다 깁니다). close code도closedBeforeTerminalMessage가 메시지에 붙입니다. 그런데도 이 리포트는 “프록시 ON이면 장시간 실패, OFF면 같은 스레드가 즉시 회복”이라서, 이미 머지된 완화만으로는 설명이 안 끝나는 실사용 회귀/잔여 구멍으로 보입니다. 긴 스레드는 create-frame이 커지고, 프록시 경로의 WS만 그 크기를 타고, 업스트림이 조용히 끊거나(1006) prelude 안에 의미 있는 Responses 이벤트가 안 오면 클라이언트가 반복 실패하는 그림과 잘 맞습니다.라인 레벨로 보면 지금 코드는 “진단·프리플라이트”는 있고 “이미 연 WS가 prelude/조기 close로 죽으면 안전한 SSE 폴백”은 아직 없습니다. 리포터 제안(첫 메시지 바이트 로그, close code/reason 보존, send→첫 Responses 이벤트 경과 시간, 다운스트림 바이트 유무, SSE 폴백 자격)은 재현 없이도 다음 패치의 측정 축이 됩니다.
src/server/responses/codex-ws-exchange.ts라인 214 - prelude 타이머가 만료되면failStream("codex websocket response prelude timed out")만 하고 SSE로 내려가지 않는다. 리포트의 두 번째 에러와 정확히 같은 문자열이다.
src/server/responses/codex-ws-exchange.ts라인 138-146 -failStream은 이미sent면 200 바디 실패로 고정한다. 이중 추론을 막으려는 설계지만, “Responses 이벤트가 한 번도 안 나간 조기 실패”와 “이미 일부 이벤트를 보낸 실패”를 같은 길로 묶는다.
src/server/responses/codex-ws-wire.ts라인 5 - prelude 한도는 지금 90초 고정이다.#3976은 닫혔지만 설정으로 늘리는 스위치는 이 HEAD 스냅샷만으로는 보이지 않는다.
src/server/responses/codex-ws-wire.ts라인 115-125 - close code/reason은 메시지 문자열에는 붙지만, 주석대로 request log의 타입 코드로는 안 남는다. 장시간 현장 디버깅이 어렵다.
src/server/responses/codex-ws-wire.ts라인 137-144 - create-frame이 한도 이상이면 SSE로 보내도록 프리플라이트가 있다. 한도 바로 아래의 긴 스레드(리플레이·스크린샷 누적)는 여전히 WS를 타고, 그때의 조기 close/prelude 실패는 폴백이 없다.메인테이너의 판단이 필요한 지점
- “아직 Responses 이벤트가 하나도 다운스트림으로 안 나간 상태”에서만 SSE 폴백을 허용할지, 아니면 이중 추론 위험 때문에 WS 실패는 항상 바디 실패로 둘지.
- prelude 90초를 더 늘리거나 설정화할지, 아니면 긴 스레드는 create-frame 크기/예상 TTFT로 WS 자격을 더 좁힐지.
- 이 이슈를
#2471/#4083후속 측정 이슈로 두고 재현 로그(프레임 바이트, close code, prelude 경과)를 필수 재현 조건으로 요구할지. - Windows Codex Desktop 2.49.0 사용자 경로와
dev2.50.0 HEAD의 동작 차이를 릴리즈 노트에 따로 적을지.
너의 추천
열어 두고 streaming 우선순위로 둔다. 첫 패치는 관측부터다. create-frame 바이트, close code/reason, send부터 첫 non-control Responses 이벤트까지 시간, 다운스트림 바이트 유무를 로그에 남긴 뒤, “sent지만 아직 responseCommitted/다운스트림 이벤트가 없는 조기 실패”만 안전한 HTTP SSE 폴백 후보로 설계한다. 설정 없는 90초 고정만 더 늘리는 것은 재현 없이 하지 말자. 중복 이슈로 닫지 말고#2471/#4083/#3976을 Related로 링크한 채 잔여 구멍으로 추적한다.이 댓글은 grok-bot이 작성했습니다
One important constraint on the proposed fallback: at dev 6d3ad12,
codex-ws-exchange.ts::failStreamdeliberately treats a successfully sent create frame as possibly executing upstream. No downstream Responses event yet is not proof the upstream did not accept or execute it. Please do not enable an automatic HTTP resend based only onresponseCommitted === falseor zero downstream bytes; that can duplicate the turn.The A/B report is worth keeping open. The next useful evidence is content-free stage information: create-frame byte count, whether send completed, close code, and elapsed time to the first meaningful event, plus exact OCX/client versions. Do not attach conversation bodies, authorization headers, cookies, or raw close reasons containing request content. Preserve the current no-replay-after-send contract unless the upstream provides affirmative non-execution/idempotency evidence.
- added a commit that references this issue
on Sep 11, 2026 - added a commit that references this issue
on Sep 12, 2026 I hit a similar disconnect on macOS with OpenCodex
2.52.0-preview.20260911and Codex CLI0.154.0-alpha.6.2.The first frame arrived after 1.1 seconds, then the connection closed with
1006after 56.1 seconds. The error reportedafter-response-started,sent=yes, and 134 relayed events. That points to a mid-response disconnect rather than the prelude timeout described above.The matching proxy log shows HTTP 502, one send, and no recovery attempt. Eight subsequent requests completed successfully. I haven't tested bypassing the proxy on the same failing thread, so I can't confirm it's the same underlying issue.
One potentially useful detail:
websockets: falsewas already set. In this version, that disables the client-to-proxy WebSocket path, but the proxy still uses WebSockets upstream to ChatGPT.Recording whether the socket was reused, the outgoing frame size, and the gap between the last frame and the close would help narrow this down. An explicit HTTP/SSE upstream option would also make a controlled comparison easier. I'd keep the no-replay-after-send behavior intact, since this request had already started producing output.
- added a commit that references this issue
on Sep 13, 2026 17 remaining items
The per-request resend budget landed on
devin 3a0718c, and it does not reach this report.That change bounds how many times one logical request may replace a send once the composed recovery legs are in play — reset, prelude EOF, credential refresh, combo candidate. What it deliberately does not do is give a dead relay a second transport: an established WebSocket that stops mid-turn is still a failed leg, not a fallback to SSE, and the prelude timeout is a separate decision about when a turn is declared unanswered. Both remain as reported.
Remaining scope: the mid-turn transport death needs a fallback path and a decision about what a partially delivered turn owes the client, and the prelude timeout needs to say whether it is the same failure class. Neither is covered by anything landed today.
Stage records from failing turns, as you asked. Canonical ChatGPT backend, macOS, direct connection with no proxy, two ChatGPT accounts, mostly gpt-6-astra. OpenCodex 2.57.0 to 2.63.0, September 17 to 22.
12,587 Codex WS turns carry a stage record. 53 of them died the same way: the create frame went out, then close 1006 with zero upstream frames and zero relayed events, 174 ms to 4.3 s after the send. Every one was on a fresh socket (
reused: false) and settled as a 502upstream_server_error. They split evenly across the two accounts (26 and 27) and fell in 8 threads, 38 of them in the two largest.Size seems to matter, but it is not a hard limit.
requestBytesis null on successful turns, so I used input tokens as the size measure. For a failed turn that is the input tokens of the retry that followed it.input tokens WS turns that completed early 1006 rate under 100k 3,061 5 0.2% 100k to 200k 6,741 26 0.4% 200k to 300k 2,652 19 0.7% The failed create frames were 0.2 MB to 14.5 MB, median 6.6 MB, all under the 16 MiB limit from #2471. The same threads completed thousands of WS turns in between.
The part that might help this issue: I run a small local patch that leaves the failed request alone. After a transport failure like this, it sends the next requests for that account and route over HTTP/SSE for 60 seconds, then lets one request probe the WebSocket again before switching back. The next request in each thread, Codex retrying the failed turn one to eleven seconds later, went over HTTP in 52 of 53 cases. 49 got a 200 and 3 got a 429. The one that went back over WS also completed. Without the patch,
shouldUseCodexWsUpstreamsends that retry straight back over WS with the same frame. That fits the original report: a long thread that only recovers once the proxy is bypassed.The proxy never sends anything twice here, so this sidesteps the double-inference question. The client decides to retry, and the proxy only picks the transport for that new request. Would you take a PR for it against
dev? The same-request fallback can stay its own discussion. The patch also has an opt-in same-request HTTP retry, but it never fired in these records, so I'm not proposing it.Independent corroboration (Windows, 2.63.0) + a verified way to reach the existing
sseFallbacktodayAdding a second environment to @FredAmartey's data, plus something I could not find in this thread: a deterministic way to force the upstream leg onto SSE while the proper fallback is still being designed.
Environment
Windows, Codex Desktop, opencodex 2.63.0, ChatGPT backend, canonical Responses requests in the 2.2–2.7 MB range,
model_reasoning_effort = "max".Codex is pointed at the local proxy over plain HTTP (
openai_base_url = http://127.0.0.1:10100/backend-api/codex) and the client-facing WebSocket is already disabled (websockets: false→ 426 → clean HTTP fallback, working as designed). Theclose 1006failures still occur, which independently confirms they originate on the upstream lane, not on the client-facing upgrade.Stage records from failing turns
request=2188716B frames=193 control=2 relayed=191 first-frame=662ms elapsed=21530ms close=1006 request=2247292B frames=13 control=2 relayed=11 first-frame=609ms elapsed=66684ms close=1006 request=2614871B frames=5 control=2 relayed=3 first-frame=499ms elapsed=4663ms close=1006 request=2712862B frames=21 control=2 relayed=19 first-frame=458ms elapsed=21483ms close=1006pings=0 pongs=0on all four. Each settled as502 upstream_server_errorwithresendPermission: "refused-committed".Two things worth flagging:
- The
elapsed=4663mssample dies far too fast to be an idle/keepalive timeout — consistent with your "174 ms to 4.3 s" set, so ping tuning cannot address this. - Because the turn is refused for resend, Codex's own retry loop then reports
exceeded retry limit, last status: 429 Too Many Requests. A lot of the "429" reports around this are probably downstream of this bug rather than a quota problem.
A verified way to force the SSE path today
Setting the global proxy to a SOCKS5 URL makes
resolveProxyRoutereturnfallback, which routes the turn through the existingsseFallback. Three lines in the shipped source make this deterministic:src/config/proxy-env.ts:171-175— asocks5://value deletesHTTP_PROXY/HTTPS_PROXYand sets onlyALL_PROXYsrc/lib/proxy-env.ts:66-73—resolveProxyRoutereturns{ kind: "fallback" }when the proxy scheme is nothttp(s)src/server/responses/ws-upstream.ts:160—if (proxyRoute.kind === "fallback") return sseFallback(url, init);
// ~/.opencodex/config.json "proxy": "socks5://127.0.0.1:7897" // any SOCKS5 endpoint reachable by the client
then
ocx stop && ocx service start.Verified by calling
resolveProxyRoute(new URL("wss://chatgpt.com/backend-api/codex"))against the shipped source:HTTPS_PROXY=http://127.0.0.1:7897 -> {"kind":"proxy","proxy":"http://127.0.0.1:7897"} ALL_PROXY=socks5://127.0.0.1:7897 -> {"kind":"fallback"} ALL_PROXY=http://127.0.0.1:7897 -> {"kind":"proxy","proxy":"http://127.0.0.1:7897"} // controlThe third row is the control: the fallback comes from the
socks5scheme, not merely from the value arriving viaALL_PROXY.End-to-end check: after restarting with the SOCKS5 value, the next real turn recorded no
codexWsStagefield at all and completed with200, andservice.logconfirms the mode at startup (outbound proxy: socks5://...).The cost is the TTFT the WS lane exists to win back (~1.0 s vs ~3.9 s p50 per the note in
ws-upstream.ts), which seems like a fair trade for anyone whose turns currently die mid-stream.Suggestion
It would help if this escape hatch were a supported, documented knob instead of a side effect of the proxy scheme — e.g.
transport: "auto" | "sse", or havingwebsockets: falsecover the upstream lane as well. Today the only route to the already-implementedsseFallbackis changing the proxy scheme, which is not discoverable and silently costs TTFT for every provider, not just the Codex one.Thanks for the detailed triage in this thread — the "failed leg is not a fallback" framing is exactly what made this reproducible.
- The
- addedpriority: P1High: reproducible failure in a core path (routing, failover, account pool, streaming, usage, auth,High: reproducible failure in a core path (routing, failover, account pool, streaming, usage, auth,
on Sep 24, 2026 Today's overtime message: unexpected status 502 Bad Gateway: codex websocket closed before a Responses terminal event (close 1011 keepalive ping timeout) [cause=no-upstream-frame request=5666621B sent=yes frames=0 control=0 relayed=0 first-frame=n/a elapsed=82924ms pings=5 pongs=0], url: http://127.0.0.1:10100/v1/responses
devin-ai-integration commented
on Sep 25, 2026 ContributorMore actionsClosing for the case this issue's data shows: a Codex WebSocket that dies after the create frame and before the first Responses event now falls back to one HTTP/SSE replacement under
retryOnReset. That behavior landed in #5675 (aed3bb8f42), and #6011, carried in #6059 (merge8923ad9835), pins it with tests and ADR-4191. A socket that dies after output has started stays a failed leg on purpose, because replaying it could run a second inference. If you still see failures after the first event, please open a new issue with the stage record.
Client or integration
Codex App (also reproduced from Codex CLI when resuming the same problematic thread).
Area
Streaming
Summary
A long-running Codex conversation became effectively unusable only while traffic was routed through OpenCodex. The same account, model, machine, repository and thread work normally again as soon as the OpenCodex proxy is disabled/bypassed.
This does not look like a generic network outage because the failure is both thread-specific and proxy-path-specific.
With OpenCodex enabled, the problematic parent thread repeatedly failed with either:
or:
After OpenCodex was disabled/bypassed, the same existing thread resumed and worked normally immediately, without fork, rollback, or compaction recovery.
Expected behavior: OpenCodex should transparently proxy the same Codex Responses request, or safely fall back to HTTP/SSE when the upstream WebSocket path cannot complete.
This appears related to one or more WebSocket relay edge cases such as large replay payloads, loss of upstream close code/reason, the response-prelude timeout, or lack of safe HTTP/SSE fallback after an already-open WebSocket terminates.
Related issues:
This report adds a real-world A/B observation: the exact same previously failing thread becomes healthy immediately when OpenCodex is bypassed.
Reproduction
stream disconnected before completionfailures for many hours.Hello!).A/B matrix:
Suggested diagnostics for reproducing this in OpenCodex:
send()to quota/metadata control events and to the first non-control Responses event such asresponse.created;Potential mitigations:
1006-style failure.Version
OpenCodex v2.49.0
Codex Desktop 26.903.61454
Operating system
Windows 11
Provider and model
Canonical ChatGPT/Codex authenticated upstream. Exact model identifier was not captured as part of the initial incident.
Logs or error output
and later:
Screenshots and supporting files
No private conversation content, repository content, credentials, or account identifiers are attached.
The strongest supporting evidence is the A/B result: the same existing long thread fails through OpenCodex and immediately succeeds when OpenCodex is bypassed.
Redacted configuration
{ "client": "Codex Desktop / Codex CLI", "proxy_path": "Codex -> local OpenCodex -> canonical ChatGPT/Codex upstream", "control_test": "Codex -> canonical ChatGPT/Codex upstream (OpenCodex bypassed)", "result_with_proxy": "stream disconnected", "result_without_proxy": "same thread works normally" }No tokens, account identifiers, request credentials, private conversation text, or repository contents are included.
Checks