Summary
When the upstream ChatGPT/Codex account returns a usage_limit_exceeded whose reset timestamp is days away, computeQuotaCooldownUntil() clamps the cooldown to CODEX_MAX_QUOTA_COOLDOWN_MS (24h) and pins the account for that entire window. The proxy then short-circuits every request locally with 429 Selected Codex account is cooling down, and keeps doing so long after the account has actually recovered. There is no way to clear the cooldown other than restarting the proxy, and none of the built-in diagnostics (health, account list) make the state legible.
In my case the account was demonstrably healthy again within ~2.5 hours, but the proxy would have kept blocking for ~24h.
This is distinct from #104 (cooldown leaking to non-forward providers). Here the cooldown applies to the forward provider itself, and the complaint is its lifetime and unclearability, not its scope.
Environment
- opencodex
2.7.33 (verified the same code path is unchanged on main / 2.7.39)
- codex-cli
0.144.6, Codex Desktop 0.146.0-alpha.3.1
- macOS (Darwin 25.5.0, arm64)
providers.openai: adapter: openai-responses, authMode: "forward", codexAccountMode: "pool", single account (__main__)
configuredServiceTier: "default" (not fast mode)
What happened
-
A burst of multi-agent work (~12.7M input tokens in 13 minutes) hit a genuine upstream usage limit. Upstream returned:
You've hit your usage limit. Visit .../settings/usage to purchase more credits
or try again at Jul 29th, 2026 8:01 AM.
codex_error_info: usage_limit_exceeded
That reset timestamp was ~4 days out.
-
computeQuotaCooldownUntil() derived the cooldown from that timestamp and clamped it to the 24h cap.
-
From then on, every request — any model, any effort — was rejected locally:
{"model":"gpt-5.5","status":429,"durationMs":5,
"errorCode":"rate_limit_exceeded",
"upstreamError":"Selected Codex account is cooling down"}
durationMs: 5 — the request never left the machine.
-
The account was actually fine. Same account, same model, ~1 minute apart:
| route |
result |
direct to https://chatgpt.com/backend-api/codex (bypassing the proxy) |
200, normal completion |
through 127.0.0.1:10100 |
429 cooling down |
The bypass test:
codex exec --sandbox read-only --skip-git-repo-check \
-c openai_base_url="https://chatgpt.com/backend-api/codex" \
-m gpt-5.5 "reply with the single word: ok"
Independently, ocx account refresh openai reported weekly 0% resets 2026-08-01, and ocx doctor showed WHAM reachability status=200, authenticated. Every available signal said the account was healthy — the proxy blocked anyway.
-
ocx restart cleared it (cooldown lives in the in-memory upstreamHealth map). Verified working through the proxy immediately after.
Why this is painful to diagnose
ocx health --json returns {"ok":true} throughout — it only probes the local port, never the upstream path it is blocking.
ocx account list shows STATUS: next session rather than something like cooling down until HH:MM. Easy to read as benign.
- The 429 surfaced to the Codex client is
exceeded retry limit, last status: 429 Too Many Requests, indistinguishable from a real upstream rate limit.
- Worse, when the cooldown was first set, the client surfaced upstream's original text verbatim — telling the user to buy credits or wait until Jul 29. Meanwhile chatgpt.com's usage page correctly showed 100% remaining on every visible bucket. Three sources, three contradictory stories, and the only correct one (
usage.jsonl) is not something a user would think to read.
I spent well over an hour chasing account quota, credit balance, service tiers, and per-model buckets before finding upstreamError in ~/.opencodex/usage.jsonl.
Suggestions
Roughly in order of value to me:
- Re-probe instead of trusting a stale cooldown. After some modest interval (say 5–10 min), let one canary request through per account. If it succeeds, clear the cooldown. Plan-level quota often frees up well before the advertised reset, and a 24h local block on a working account is far more costly than one wasted probe.
- Cap far-future resets much lower, or don't derive a hard local block from
resetAt at all when it is many hours out. Retry-After is a rate-limit hint and worth honoring literally; a weekly-quota resetAt is not the same kind of signal.
- Add an explicit escape hatch — e.g.
ocx account clear-cooldown <provider> [id]. Right now "restart the whole proxy" is the only remedy, which is disruptive when other providers are working fine.
- Make the state visible: include
cooldownUntil in ocx account list and in ocx health --json, and put the remaining duration into the 429 body (Selected Codex account is cooling down for another 21h 14m). That one string would have ended my debugging in seconds.
Happy to test a patch — the reproduction is straightforward to trigger with a resetAt far in the future.
Summary
When the upstream ChatGPT/Codex account returns a
usage_limit_exceededwhose reset timestamp is days away,computeQuotaCooldownUntil()clamps the cooldown toCODEX_MAX_QUOTA_COOLDOWN_MS(24h) and pins the account for that entire window. The proxy then short-circuits every request locally with429 Selected Codex account is cooling down, and keeps doing so long after the account has actually recovered. There is no way to clear the cooldown other than restarting the proxy, and none of the built-in diagnostics (health,account list) make the state legible.In my case the account was demonstrably healthy again within ~2.5 hours, but the proxy would have kept blocking for ~24h.
This is distinct from #104 (cooldown leaking to non-forward providers). Here the cooldown applies to the forward provider itself, and the complaint is its lifetime and unclearability, not its scope.
Environment
2.7.33(verified the same code path is unchanged onmain/2.7.39)0.144.6, Codex Desktop0.146.0-alpha.3.1providers.openai:adapter: openai-responses,authMode: "forward",codexAccountMode: "pool", single account (__main__)configuredServiceTier: "default"(not fast mode)What happened
A burst of multi-agent work (~12.7M input tokens in 13 minutes) hit a genuine upstream usage limit. Upstream returned:
That reset timestamp was ~4 days out.
computeQuotaCooldownUntil()derived the cooldown from that timestamp and clamped it to the 24h cap.From then on, every request — any model, any effort — was rejected locally:
{"model":"gpt-5.5","status":429,"durationMs":5, "errorCode":"rate_limit_exceeded", "upstreamError":"Selected Codex account is cooling down"}durationMs: 5— the request never left the machine.The account was actually fine. Same account, same model, ~1 minute apart:
https://chatgpt.com/backend-api/codex(bypassing the proxy)127.0.0.1:10100cooling downThe bypass test:
Independently,
ocx account refresh openaireportedweekly 0% resets 2026-08-01, andocx doctorshowed WHAM reachabilitystatus=200, authenticated. Every available signal said the account was healthy — the proxy blocked anyway.ocx restartcleared it (cooldown lives in the in-memoryupstreamHealthmap). Verified working through the proxy immediately after.Why this is painful to diagnose
ocx health --jsonreturns{"ok":true}throughout — it only probes the local port, never the upstream path it is blocking.ocx account listshowsSTATUS: next sessionrather than something likecooling down until HH:MM. Easy to read as benign.exceeded retry limit, last status: 429 Too Many Requests, indistinguishable from a real upstream rate limit.usage.jsonl) is not something a user would think to read.I spent well over an hour chasing account quota, credit balance, service tiers, and per-model buckets before finding
upstreamErrorin~/.opencodex/usage.jsonl.Suggestions
Roughly in order of value to me:
resetAtat all when it is many hours out.Retry-Afteris a rate-limit hint and worth honoring literally; a weekly-quotaresetAtis not the same kind of signal.ocx account clear-cooldown <provider> [id]. Right now "restart the whole proxy" is the only remedy, which is disruptive when other providers are working fine.cooldownUntilinocx account listand inocx health --json, and put the remaining duration into the 429 body (Selected Codex account is cooling down for another 21h 14m). That one string would have ended my debugging in seconds.Happy to test a patch — the reproduction is straightforward to trigger with a
resetAtfar in the future.