Skip to content

Z.ai: quota probe returns nothing when baseUrl is the Anthropic-compatible coding endpoint (https://api.z.ai/api/anthropic) #4154

Description

@nordz0r

Client or integration

Other — OpenCodex provider UI / provider quota probe (zai provider)

Provider or upstream service

Z.ai — GLM Coding Plan

OpenCodex version

2.49.0

Endpoint or capability

Provider quota probe (GET https://api.z.ai/api/monitor/usage/quota/limit) vs. Anthropic-compatible messages endpoint POST https://api.z.ai/api/anthropic/v1/messages

Current behaviour

fetchZaiQuota() only runs when isCanonicalZaiBaseUrl() matches. The canonical list accepts https://api.z.ai and https://api.z.ai/api/coding/paas/v4, but not the documented Anthropic-compatible coding endpoint https://api.z.ai/api/anthropic.

Z.ai's GLM Coding Plan serves the Anthropic wire format only at https://api.z.ai/api/anthropic:

  • POST https://api.z.ai/api/anthropic/v1/messages → 200 (works)
  • POST https://api.z.ai/v1/messages → 404
  • POST https://api.z.ai/api/coding/paas/v4/v1/messages → 404
  • POST https://api.z.ai/api/coding/paas/v4/messages → 404

Since the Anthropic adapter builds the messages URL as {baseUrl}/v1/messages, users face a mutually exclusive choice:

  • baseUrl = https://api.z.ai/api/anthropic → messages work, provider quota view is empty
  • baseUrl = https://api.z.ai/api/coding/paas/v4 → quota view works, all message requests 404

Expected behaviour

The quota probe should also treat https://api.z.ai/api/anthropic as a canonical Z.ai Coding Plan base URL, so a provider configured for the Anthropic wire format still gets its quota/usage view populated.

Minimal redacted request or reproduction

# 1. Configure a provider with:
#    adapter  = anthropic
#    baseUrl  = https://api.z.ai/api/anthropic
#    authMode = key (x-api-key)
#
# 2. Verify messages work through the proxy — 200 OK.
# 3. Open the provider "Usage"/quota view — it stays empty.

# The probe's guard rejects the URL:
#   src/providers/quota.ts :: isCanonicalZaiBaseUrl("https://api.z.ai/api/anthropic") === false
#   → fetchZaiQuota() returns null before any request is made.

# Messages endpoint itself is healthy (key redacted):
curl -s -o /dev/null -w '%{http_code}\n' https://api.z.ai/api/anthropic/v1/messages \
  -H 'x-api-key: <REDACTED>' -H 'anthropic-version: 2023-06-01' \
  -H 'Content-Type: application/json' \
  -d '{"model":"glm-5.3-flash","max_tokens":8,"messages":[{"role":"user","content":"hi"}]}'
# → 200

Actual response or error

No error is surfaced — the quota view silently stays empty, because fetchZaiQuota bails out at the canonical-URL guard before requesting https://api.z.ai/api/monitor/usage/quota/limit.

For reference, the quota endpoint itself works fine for Coding Plan keys (verified with Authorization: Bearer <coding-plan-key> → 200 with a data.limits[] payload), it is only the canonical-URL gate that excludes the Anthropic base URL.

Upstream documentation

  • Z.ai GLM Coding Plan / DevPack docs: https://docs.z.ai/devpack/overview — the Anthropic-compatible base URL for coding tools is https://api.z.ai/api/anthropic (e.g. used as ANTHROPIC_BASE_URL for Claude Code).
  • Quota monitor endpoint used by the probe: GET https://api.z.ai/api/monitor/usage/quota/limit (no public spec; verified empirically, returns {code, data:{limits:[{type,unit,number,usage,currentValue,remaining,percentage,nextResetTime}]}}).

Suggested mapping or implementation notes

One-line addition in isCanonicalZaiBaseUrl (src/providers/quota.ts):

|| normalized === `${ZAI_BASE_URL}/api/anthropic`

fetchZaiQuota already derives the monitor host correctly for the international host and sends Authorization: Bearer <key>, which the monitor endpoint accepts for Coding Plan keys.

Additional data point from a controlled A/B test (30 identical requests ≈ 75k tokens per leg, GLM-5.3-Flash, inside the September 2026 Flash campaign window, neutral user-agent): the OpenAI-compatible path (/api/coding/paas/v4/chat/completions) reduced the plan's remaining budget by ~31 units per window vs ~6 units for the Anthropic path (/api/anthropic/v1/messages) — i.e. the two paths account quota noticeably differently, which makes a working quota view for the Anthropic base URL even more useful.

Additional context and attachments

Configured provider (redacted):

{
  "name": "zai",
  "adapter": "anthropic",
  "baseUrl": "https://api.z.ai/api/anthropic",
  "defaultModel": "glm-5.3-flash",
  "authMode": "key",
  "apiKeyTransport": "x-api-key"
}

Happy to test a fix against 2.49.x.

Checks

  • I searched existing provider and compatibility issues.
  • The request and response were redacted.
  • The expected behaviour is based on an upstream specification or a concrete client requirement.

Activity

  1. added
    account-poolOAuth, credentials, Codex pool, quota, failover, plans
    providerProvider adapters, OpenAI-compat presets, upstream API quirks
    on Sep 9, 2026
  2. lidge-jun commented on Sep 9, 2026

    @lidge-jun
    Owner

    리뷰 · 우선순위 58 / 80

    Z.ai GLM Coding Plan을 Anthropic 호환 베이스 URL(https://api.z.ai/api/anthropic)로 쓰면 메시지는 되는데 Usage/quota 화면이 비는 버그입니다. 지금 dev의 src/providers/quota.ts에서 fetchZaiQuota는 맨 앞에서 isCanonicalZaiBaseUrl(config.baseUrl)이 아니면 바로 null을 돌려줍니다(854행). 정규화 허용 목록(348-355행)에는 https://api.z.ai, .../api/coding/paas/v4, CN 호스트와 CN paas/v4, CN /api/v1만 있고 /api/anthropic은 없습니다. Anthropic 어댑터는 {baseUrl}/v1/messages로 붙이므로, 문서상 코딩용 Anthropic 베이스를 쓰면 메시지는 200이 나고, paas/v4를 쓰면 쿼터 프로브만 살아 남고 메시지는 404가 납니다. 둘 중 하나만 고를 수 있는 상태입니다.

    이슈가 제안한 한 줄(normalized === ZAI_BASE_URL + "/api/anthropic")은 방향이 맞습니다. 다만 같은 파일의 monitorHost 분기(858-860행)를 꼭 같이 봐야 합니다. 국제 호스트로 인정되는 값이 ZAI_BASE_URL과 .../api/coding/paas/v4뿐이라, /api/anthropic만 canonical에 넣고 여기를 안 고치면 monitorHost가 ZAI_CN_BASE_URL로 떨어지고 Authorization도 CN 형식(키 그대로)이 됩니다. 국제 Coding Plan 키로 CN 모니터를 치면 프로브가 또 실패하거나 엉뚱한 쪽으로 갑니다. 국제 Anthropic 베이스도 ZAI_BASE_URL 모니터 + Bearer로 가야 이슈 재현(모니터 엔드포인트는 Bearer로 200)과 맞습니다. CN 쪽에 Anthropic 호환 베이스가 따로 있으면 대칭으로 넣을지, 국제만 먼저 막을지도 정하면 됩니다.

    round-2 계획(#4155) 이선 목록에는 이 이슈가 없습니다. 그래도 프로브 가드 한곳 + 모니터 호스트 분기 한곳이라 범위가 작고, Lane A/B 스택과 파일 충돌도 거의 없어 독립 PR로 넣기 좋습니다. types/config 분할 캠페인과도 무관해서 close-don't-rebase 대상이 아닙니다.

    라인 348-355 - isCanonicalZaiBaseUrl이 api.z.ai/api/anthropic을 허용하지 않아 fetchZaiQuota가 요청 전에 null을 반환합니다
    라인 854 - canonical 가드 실패 시 조용히 null이라 Usage UI에 에러 없이 빈 화면만 남습니다
    라인 858-860 - monitorHost가 국제 paas/v4와 root만 국제로 보고, /api/anthropic을 넣지 않으면 CN 호스트+CN Authorization으로 떨어집니다
    경로 src/providers/quota.ts fetchZaiQuota - 모니터 URL 자체(GET /api/monitor/usage/quota/limit)는 Coding Plan 키로 동작한다고 제보됨. 가드/호스트만 맞추면 됩니다

    메인테이너의 판단이 필요한 지점

    • 국제 /api/anthropic만 넣을지, CN Anthropic 호환 베이스가 있으면 대칭 추가할지.
    • round-2 Lane 밖에 독립 PR로 바로 받을지, 다음 배치로 미룰지.
    • 쿼터 차감이 Anthropic 경로와 OpenAI paas 경로에서 다르다는 제보를 문서/가이드에 한 줄 남길지(코드 필수 아님).

    너의 추천
    독립 작은 PR로 isCanonicalZaiBaseUrl에 https://api.z.ai/api/anthropic을 추가하고, monitorHost 국제 분기에도 같은 URL을 넣어 ZAI_BASE_URL+Bearer로 모니터하게 고쳐라. 단위 테스트로 canonical true와 모니터 호스트가 국제인지 잠그면 충분하다. round-2 이선과 섞지 말고 따로 머지해도 된다.

    이 댓글은 grok-bot이 작성했습니다

  3. nordz0r commented on Sep 9, 2026

    @nordz0r
    ContributorAuthor

    Correction on the quota-consumption data point in the original post.

    A controlled re-test (with a drift-measurement leg) showed no meaningful per-endpoint difference after all. Design: quota snapshot from the monitor endpoint -> 6 min idle -> 30 identical requests (~91k tokens) via /api/anthropic/v1/messages -> snapshot -> 30 identical requests (~91k tokens) via /api/coding/paas/v4/chat/completions -> snapshot, GLM-5.3-Flash, inside the campaign window.

    • 6 min idle: 0 units (no background tick)
    • Anthropic leg: -7 units / 91k tokens
    • OpenAI-chat leg: -4 units / 91k tokens
    • OpenAI Responses leg (https://api.z.ai/api/v1/responses, the endpoint Codex uses per docs): -2 units / 19k tokens

    The ~5x discrepancy in my original measurements was contamination from concurrent plan usage, not a property of the endpoints. The feature request itself stands: isCanonicalZaiBaseUrl should include https://api.z.ai/api/anthropic so the quota view works for Anthropic-wire configurations (this is the documented base URL for Claude Code in the Z.ai DevPack docs: https://docs.z.ai/devpack/tool/claude). Note the docs also define a third wire endpoint for Codex — OpenAI Responses at https://api.z.ai/api/v1 — which is likewise absent from the canonical list for the international host (only the CN /api/v1 variant is listed).

  4. Ingwannu commented on Sep 10, 2026

    @Ingwannu
    Owner

    The remaining quota-display bug is confirmed in current dev. I will make a separate narrow PR covering the documented international Anthropic base (/api/anthropic) and Codex Responses base (/api/v1), using one exact-base-to-monitor mapping for both eligibility and dispatch. This avoids admitting the new international paths and accidentally falling through to the CN monitor/authentication branch.

    I will preserve the existing CN allowlist and redirect refusal, and add mocked tests for the international host/Bearer pairing, existing CN/raw-key pairing, and unsupported destinations causing no request. No inference endpoint, saved configuration, or quota-consumption policy needs to change.

    Thanks for correcting the original consumption comparison. I will not carry the withdrawn 5x claim into code or documentation; this fix is only for quota-probe selection.

  5. Ingwannu commented on Sep 10, 2026

    @Ingwannu
    Owner

    The proposed fix is now in Draft PR #4174, with a shared destination mapping, 14 added mocked regression cases, and documentation. The extracted-function base/candidate checks confirmed the missing international endpoints and preserved the CN host/auth pairing. Full Bun/typecheck/docs validation remains pending; this issue stays open until the fix is reviewed and landed. No live credentials or provider requests were used.

  6. nordz0r commented on Sep 10, 2026

    @nordz0r
    ContributorAuthor

    Reviewed the Draft PR #4174 diff — the shared zaiQuotaMonitorHost mapping resolves both the admission guard and the monitorHost dispatch caveat in one place. One confirmation from live behavior (Coding Plan key, redacted):

    • GET https://api.z.ai/api/monitor/usage/quota/limit with Authorization: Bearer <coding-plan-key> → 200, payload: {code: 200, data: {level: "lite", limits: [{type: "CREDIT_LIMIT", unit: 3, number: 5, usage, currentValue, remaining, percentage, nextResetTime}, {unit: 6, number: 1, ...}]}}
    • unit=3, number=5 is the 5-hour window and unit=6, number=1 the weekly window — worth locking as a mocked fixture so parseZaiQuotaLimits doesn't regress against the real shape.

    On the open documentation question (quota-deduction note per endpoint) — from the controlled re-test above, suggested wording for providers.md:

    "All GLM Coding Plan wire endpoints (/api/anthropic for Anthropic-wire clients, /api/coding/paas/v4 for OpenAI Chat, /api/v1 for OpenAI Responses) draw on the same plan quota; controlled tests found no per-endpoint accounting difference. Campaign rates depend on client identification, not on the wire format."

    Follow-up candidate, separate issue: custom/self-hosted providers whose baseUrl points at a cluster-local adapter currently have no quota probe at all (name list + canonical-URL gate). A per-provider quotaProbeUrl override would cover that class — can file it if useful.

  7. lidge-jun commented on Sep 11, 2026

    @lidge-jun
    Owner

    Landed via #4262 at f40e432

    dev now admits the Anthropic-compatible Coding Plan base (https://api.z.ai/api/anthropic) and related international Responses/Coding bases through one zaiQuotaMonitorHost mapping, so the Usage panel can probe again. Closing as completed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    account-poolOAuth, credentials, Codex pool, quota, failover, plansproviderProvider adapters, OpenAI-compat presets, upstream API quirksprovider-compatibilityProvider compatibility reports

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions