Skip to content

Quota cooldown from a far-future upstream reset pins the account for the full 24h cap, with no way to clear it but a restart #433

Description

@yzq2003-dotcom

Summary

When the upstream ChatGPT/Codex account returns a usage_limit_exceeded whose reset timestamp is days away, computeQuotaCooldownUntil() clamps the cooldown to CODEX_MAX_QUOTA_COOLDOWN_MS (24h) and pins the account for that entire window. The proxy then short-circuits every request locally with 429 Selected Codex account is cooling down, and keeps doing so long after the account has actually recovered. There is no way to clear the cooldown other than restarting the proxy, and none of the built-in diagnostics (health, account list) make the state legible.

In my case the account was demonstrably healthy again within ~2.5 hours, but the proxy would have kept blocking for ~24h.

This is distinct from #104 (cooldown leaking to non-forward providers). Here the cooldown applies to the forward provider itself, and the complaint is its lifetime and unclearability, not its scope.

Environment

  • opencodex 2.7.33 (verified the same code path is unchanged on main / 2.7.39)
  • codex-cli 0.144.6, Codex Desktop 0.146.0-alpha.3.1
  • macOS (Darwin 25.5.0, arm64)
  • providers.openai: adapter: openai-responses, authMode: "forward", codexAccountMode: "pool", single account (__main__)
  • configuredServiceTier: "default" (not fast mode)

What happened

  1. A burst of multi-agent work (~12.7M input tokens in 13 minutes) hit a genuine upstream usage limit. Upstream returned:

    You've hit your usage limit. Visit .../settings/usage to purchase more credits
    or try again at Jul 29th, 2026 8:01 AM.
    codex_error_info: usage_limit_exceeded
    

    That reset timestamp was ~4 days out.

  2. computeQuotaCooldownUntil() derived the cooldown from that timestamp and clamped it to the 24h cap.

  3. From then on, every request — any model, any effort — was rejected locally:

    {"model":"gpt-5.5","status":429,"durationMs":5,
     "errorCode":"rate_limit_exceeded",
     "upstreamError":"Selected Codex account is cooling down"}

    durationMs: 5 — the request never left the machine.

  4. The account was actually fine. Same account, same model, ~1 minute apart:

    route result
    direct to https://chatgpt.com/backend-api/codex (bypassing the proxy) 200, normal completion
    through 127.0.0.1:10100 429 cooling down

    The bypass test:

    codex exec --sandbox read-only --skip-git-repo-check \
      -c openai_base_url="https://chatgpt.com/backend-api/codex" \
      -m gpt-5.5 "reply with the single word: ok"

    Independently, ocx account refresh openai reported weekly 0% resets 2026-08-01, and ocx doctor showed WHAM reachability status=200, authenticated. Every available signal said the account was healthy — the proxy blocked anyway.

  5. ocx restart cleared it (cooldown lives in the in-memory upstreamHealth map). Verified working through the proxy immediately after.

Why this is painful to diagnose

  • ocx health --json returns {"ok":true} throughout — it only probes the local port, never the upstream path it is blocking.
  • ocx account list shows STATUS: next session rather than something like cooling down until HH:MM. Easy to read as benign.
  • The 429 surfaced to the Codex client is exceeded retry limit, last status: 429 Too Many Requests, indistinguishable from a real upstream rate limit.
  • Worse, when the cooldown was first set, the client surfaced upstream's original text verbatim — telling the user to buy credits or wait until Jul 29. Meanwhile chatgpt.com's usage page correctly showed 100% remaining on every visible bucket. Three sources, three contradictory stories, and the only correct one (usage.jsonl) is not something a user would think to read.

I spent well over an hour chasing account quota, credit balance, service tiers, and per-model buckets before finding upstreamError in ~/.opencodex/usage.jsonl.

Suggestions

Roughly in order of value to me:

  1. Re-probe instead of trusting a stale cooldown. After some modest interval (say 5–10 min), let one canary request through per account. If it succeeds, clear the cooldown. Plan-level quota often frees up well before the advertised reset, and a 24h local block on a working account is far more costly than one wasted probe.
  2. Cap far-future resets much lower, or don't derive a hard local block from resetAt at all when it is many hours out. Retry-After is a rate-limit hint and worth honoring literally; a weekly-quota resetAt is not the same kind of signal.
  3. Add an explicit escape hatch — e.g. ocx account clear-cooldown <provider> [id]. Right now "restart the whole proxy" is the only remedy, which is disruptive when other providers are working fine.
  4. Make the state visible: include cooldownUntil in ocx account list and in ocx health --json, and put the remaining duration into the 429 body (Selected Codex account is cooling down for another 21h 14m). That one string would have ended my debugging in seconds.

Happy to test a patch — the reproduction is straightforward to trigger with a resetAt far in the future.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions