Repository navigation
Z.ai: quota probe returns nothing when baseUrl is the Anthropic-compatible coding endpoint (https://api.z.ai/api/anthropic) #4154
Description
Activity
- addedprovider-compatibilityProvider compatibility reportsProvider compatibility reportsaccount-poolOAuth, credentials, Codex pool, quota, failover, plansOAuth, credentials, Codex pool, quota, failover, plansproviderProvider adapters, OpenAI-compat presets, upstream API quirksProvider adapters, OpenAI-compat presets, upstream API quirks
on Sep 9, 2026 리뷰 · 우선순위 58 / 80
Z.ai GLM Coding Plan을 Anthropic 호환 베이스 URL(
https://api.z.ai/api/anthropic)로 쓰면 메시지는 되는데 Usage/quota 화면이 비는 버그입니다. 지금dev의src/providers/quota.ts에서fetchZaiQuota는 맨 앞에서isCanonicalZaiBaseUrl(config.baseUrl)이 아니면 바로null을 돌려줍니다(854행). 정규화 허용 목록(348-355행)에는https://api.z.ai,.../api/coding/paas/v4, CN 호스트와 CN paas/v4, CN/api/v1만 있고/api/anthropic은 없습니다. Anthropic 어댑터는{baseUrl}/v1/messages로 붙이므로, 문서상 코딩용 Anthropic 베이스를 쓰면 메시지는 200이 나고, paas/v4를 쓰면 쿼터 프로브만 살아 남고 메시지는 404가 납니다. 둘 중 하나만 고를 수 있는 상태입니다.이슈가 제안한 한 줄(
normalized === ZAI_BASE_URL + "/api/anthropic")은 방향이 맞습니다. 다만 같은 파일의monitorHost분기(858-860행)를 꼭 같이 봐야 합니다. 국제 호스트로 인정되는 값이ZAI_BASE_URL과.../api/coding/paas/v4뿐이라,/api/anthropic만 canonical에 넣고 여기를 안 고치면monitorHost가ZAI_CN_BASE_URL로 떨어지고 Authorization도 CN 형식(키 그대로)이 됩니다. 국제 Coding Plan 키로 CN 모니터를 치면 프로브가 또 실패하거나 엉뚱한 쪽으로 갑니다. 국제 Anthropic 베이스도ZAI_BASE_URL모니터 +Bearer로 가야 이슈 재현(모니터 엔드포인트는 Bearer로 200)과 맞습니다. CN 쪽에 Anthropic 호환 베이스가 따로 있으면 대칭으로 넣을지, 국제만 먼저 막을지도 정하면 됩니다.round-2 계획(#4155) 이선 목록에는 이 이슈가 없습니다. 그래도 프로브 가드 한곳 + 모니터 호스트 분기 한곳이라 범위가 작고, Lane A/B 스택과 파일 충돌도 거의 없어 독립 PR로 넣기 좋습니다. types/config 분할 캠페인과도 무관해서 close-don't-rebase 대상이 아닙니다.
라인 348-355 - isCanonicalZaiBaseUrl이 api.z.ai/api/anthropic을 허용하지 않아 fetchZaiQuota가 요청 전에 null을 반환합니다
라인 854 - canonical 가드 실패 시 조용히 null이라 Usage UI에 에러 없이 빈 화면만 남습니다
라인 858-860 - monitorHost가 국제 paas/v4와 root만 국제로 보고, /api/anthropic을 넣지 않으면 CN 호스트+CN Authorization으로 떨어집니다
경로 src/providers/quota.ts fetchZaiQuota - 모니터 URL 자체(GET /api/monitor/usage/quota/limit)는 Coding Plan 키로 동작한다고 제보됨. 가드/호스트만 맞추면 됩니다메인테이너의 판단이 필요한 지점
- 국제
/api/anthropic만 넣을지, CN Anthropic 호환 베이스가 있으면 대칭 추가할지. - round-2 Lane 밖에 독립 PR로 바로 받을지, 다음 배치로 미룰지.
- 쿼터 차감이 Anthropic 경로와 OpenAI paas 경로에서 다르다는 제보를 문서/가이드에 한 줄 남길지(코드 필수 아님).
너의 추천
독립 작은 PR로isCanonicalZaiBaseUrl에https://api.z.ai/api/anthropic을 추가하고,monitorHost국제 분기에도 같은 URL을 넣어ZAI_BASE_URL+Bearer로 모니터하게 고쳐라. 단위 테스트로 canonical true와 모니터 호스트가 국제인지 잠그면 충분하다. round-2 이선과 섞지 말고 따로 머지해도 된다.이 댓글은 grok-bot이 작성했습니다
- 국제
Correction on the quota-consumption data point in the original post.
A controlled re-test (with a drift-measurement leg) showed no meaningful per-endpoint difference after all. Design: quota snapshot from the monitor endpoint -> 6 min idle -> 30 identical requests (~91k tokens) via /api/anthropic/v1/messages -> snapshot -> 30 identical requests (~91k tokens) via /api/coding/paas/v4/chat/completions -> snapshot, GLM-5.3-Flash, inside the campaign window.
- 6 min idle: 0 units (no background tick)
- Anthropic leg: -7 units / 91k tokens
- OpenAI-chat leg: -4 units / 91k tokens
- OpenAI Responses leg (https://api.z.ai/api/v1/responses, the endpoint Codex uses per docs): -2 units / 19k tokens
The ~5x discrepancy in my original measurements was contamination from concurrent plan usage, not a property of the endpoints. The feature request itself stands:
isCanonicalZaiBaseUrlshould includehttps://api.z.ai/api/anthropicso the quota view works for Anthropic-wire configurations (this is the documented base URL for Claude Code in the Z.ai DevPack docs: https://docs.z.ai/devpack/tool/claude). Note the docs also define a third wire endpoint for Codex — OpenAI Responses athttps://api.z.ai/api/v1— which is likewise absent from the canonical list for the international host (only the CN/api/v1variant is listed).The remaining quota-display bug is confirmed in current dev. I will make a separate narrow PR covering the documented international Anthropic base (/api/anthropic) and Codex Responses base (/api/v1), using one exact-base-to-monitor mapping for both eligibility and dispatch. This avoids admitting the new international paths and accidentally falling through to the CN monitor/authentication branch.
I will preserve the existing CN allowlist and redirect refusal, and add mocked tests for the international host/Bearer pairing, existing CN/raw-key pairing, and unsupported destinations causing no request. No inference endpoint, saved configuration, or quota-consumption policy needs to change.
Thanks for correcting the original consumption comparison. I will not carry the withdrawn 5x claim into code or documentation; this fix is only for quota-probe selection.
The proposed fix is now in Draft PR #4174, with a shared destination mapping, 14 added mocked regression cases, and documentation. The extracted-function base/candidate checks confirmed the missing international endpoints and preserved the CN host/auth pairing. Full Bun/typecheck/docs validation remains pending; this issue stays open until the fix is reviewed and landed. No live credentials or provider requests were used.
Reviewed the Draft PR #4174 diff — the shared
zaiQuotaMonitorHostmapping resolves both the admission guard and themonitorHostdispatch caveat in one place. One confirmation from live behavior (Coding Plan key, redacted):GET https://api.z.ai/api/monitor/usage/quota/limitwithAuthorization: Bearer <coding-plan-key>→ 200, payload:{code: 200, data: {level: "lite", limits: [{type: "CREDIT_LIMIT", unit: 3, number: 5, usage, currentValue, remaining, percentage, nextResetTime}, {unit: 6, number: 1, ...}]}}unit=3, number=5is the 5-hour window andunit=6, number=1the weekly window — worth locking as a mocked fixture soparseZaiQuotaLimitsdoesn't regress against the real shape.
On the open documentation question (quota-deduction note per endpoint) — from the controlled re-test above, suggested wording for
providers.md:"All GLM Coding Plan wire endpoints (
/api/anthropicfor Anthropic-wire clients,/api/coding/paas/v4for OpenAI Chat,/api/v1for OpenAI Responses) draw on the same plan quota; controlled tests found no per-endpoint accounting difference. Campaign rates depend on client identification, not on the wire format."Follow-up candidate, separate issue: custom/self-hosted providers whose
baseUrlpoints at a cluster-local adapter currently have no quota probe at all (name list + canonical-URL gate). A per-providerquotaProbeUrloverride would cover that class — can file it if useful.
Client or integration
Other — OpenCodex provider UI / provider quota probe (zai provider)
Provider or upstream service
Z.ai — GLM Coding Plan
OpenCodex version
2.49.0
Endpoint or capability
Provider quota probe (
GET https://api.z.ai/api/monitor/usage/quota/limit) vs. Anthropic-compatible messages endpointPOST https://api.z.ai/api/anthropic/v1/messagesCurrent behaviour
fetchZaiQuota()only runs whenisCanonicalZaiBaseUrl()matches. The canonical list acceptshttps://api.z.aiandhttps://api.z.ai/api/coding/paas/v4, but not the documented Anthropic-compatible coding endpointhttps://api.z.ai/api/anthropic.Z.ai's GLM Coding Plan serves the Anthropic wire format only at
https://api.z.ai/api/anthropic:POST https://api.z.ai/api/anthropic/v1/messages→ 200 (works)POST https://api.z.ai/v1/messages→ 404POST https://api.z.ai/api/coding/paas/v4/v1/messages→ 404POST https://api.z.ai/api/coding/paas/v4/messages→ 404Since the Anthropic adapter builds the messages URL as
{baseUrl}/v1/messages, users face a mutually exclusive choice:baseUrl = https://api.z.ai/api/anthropic→ messages work, provider quota view is emptybaseUrl = https://api.z.ai/api/coding/paas/v4→ quota view works, all message requests 404Expected behaviour
The quota probe should also treat
https://api.z.ai/api/anthropicas a canonical Z.ai Coding Plan base URL, so a provider configured for the Anthropic wire format still gets its quota/usage view populated.Minimal redacted request or reproduction
Actual response or error
No error is surfaced — the quota view silently stays empty, because
fetchZaiQuotabails out at the canonical-URL guard before requestinghttps://api.z.ai/api/monitor/usage/quota/limit.For reference, the quota endpoint itself works fine for Coding Plan keys (verified with
Authorization: Bearer <coding-plan-key>→ 200 with adata.limits[]payload), it is only the canonical-URL gate that excludes the Anthropic base URL.Upstream documentation
https://api.z.ai/api/anthropic(e.g. used asANTHROPIC_BASE_URLfor Claude Code).GET https://api.z.ai/api/monitor/usage/quota/limit(no public spec; verified empirically, returns{code, data:{limits:[{type,unit,number,usage,currentValue,remaining,percentage,nextResetTime}]}}).Suggested mapping or implementation notes
One-line addition in
isCanonicalZaiBaseUrl(src/providers/quota.ts):fetchZaiQuotaalready derives the monitor host correctly for the international host and sendsAuthorization: Bearer <key>, which the monitor endpoint accepts for Coding Plan keys.Additional data point from a controlled A/B test (30 identical requests ≈ 75k tokens per leg, GLM-5.3-Flash, inside the September 2026 Flash campaign window, neutral user-agent): the OpenAI-compatible path (
/api/coding/paas/v4/chat/completions) reduced the plan'sremainingbudget by ~31 units per window vs ~6 units for the Anthropic path (/api/anthropic/v1/messages) — i.e. the two paths account quota noticeably differently, which makes a working quota view for the Anthropic base URL even more useful.Additional context and attachments
Configured provider (redacted):
{ "name": "zai", "adapter": "anthropic", "baseUrl": "https://api.z.ai/api/anthropic", "defaultModel": "glm-5.3-flash", "authMode": "key", "apiKeyTransport": "x-api-key" }Happy to test a fix against 2.49.x.
Checks