Summary
nemoclaw status can report healthy inference while the configured model cannot actually be invoked. The inference health probes check that the provider endpoint answers an HTTP request (or that the local daemon lists models), not that the configured model responds to a real completion call. This confirms the observation @mikemason left on #6635 ("At the time OpenClaw couldn't reach the model, nemoclaw status was reporting healthy inference, so whatever it's checking is slightly different than what OpenClaw actually uses to hit the model") — and it reproduces without the corporate-proxy / TLS specifics of that thread.
Mechanism (code)
For every provider except the NVIDIA Kimi K2.6 special case, the health probe is a reachability/listing check:
- Local (
ollama-local / vllm-local) — probeLocalProviderHealth(provider, options) (src/lib/inference/local.ts) does not take a model argument at all. It hits /api/tags and returns healthy when the body parses as { "models": [...] } — and an empty array is explicitly accepted ("just means no models pulled yet", isValidOllamaTagsResponseBody). Daemon-up ⇒ healthy, regardless of the configured model.
- Remote (
anthropic-prod, openai-api, gemini-api, non-Kimi nvidia-*) — getRemoteProviderHealthEndpoint returns the provider's /v1/models listing URL, and probeRemoteProviderHealth curls it with no credential (["-sS", "--connect-timeout", "3", "--max-time", "5", endpoint]), then sets ok = (curlStatus === 0) with the comment "Even a 401/403 means the endpoint is reachable" (src/lib/inference/health.ts). Because it is unauthenticated, any HTTP answer — including 401 Unauthorized — reads as healthy: the check confirms the host is network-reachable, nothing about credential validity or model invocability.
- Only
NVIDIA_MANAGED_PROVIDERS + isKimiK26Model routes to probeNvidiaKimiK26Health, which hits /chat/completions — an actual invocation. That is the one path that reflects model invocability.
The in-sandbox probe (probeSandboxInferenceGatewayHealth, inference-route-health.ts) is likewise a route-reachability check on https://inference.local/v1/models, not a completion call.
Reproduction (by execution, proxy-independent)
Driving the compiled probes on x86_64/WSL2, no proxy dependence:
Local — the probe is model-blind:
$ node -e 'console.log(require("./dist/lib/inference/local.js").probeLocalProviderHealth("ollama-local"))'
{
ok: true,
providerLabel: 'Local Ollama',
endpoint: 'http://127.0.0.1:11434/api/tags',
detail: 'Local Ollama is reachable on http://127.0.0.1:11434/api/tags.',
probeLabel: 'ollama backend',
subprobes: [ { probeLabel: 'auth proxy', ok: false, ... } ] // auth-proxy hop down in this lab; incidental — the point is the top-level model-blind `ok: true`
}
# but an actual call to a model that is not present fails:
$ curl -s http://127.0.0.1:11434/api/chat -d '{"model":"nonexistent-model:latest","messages":[{"role":"user","content":"hi"}],"stream":false}'
{"error":"model 'nonexistent-model:latest' not found"}
probeLocalProviderHealth takes no model parameter, so a sandbox configured for a model that is not pulled still probes ok: true while a real call errors model not found.
Remote — a 401 reads as healthy:
$ node -e 'console.log(require("./dist/lib/inference/health.js").probeProviderHealth("openai-api"))'
{
ok: true,
probed: true,
providerLabel: 'OpenAI',
endpoint: 'https://api.openai.com/v1/models',
detail: 'OpenAI endpoint is reachable at https://api.openai.com/v1/models.'
}
# the endpoint actually answered unauthenticated:
$ curl -s -o /dev/null -w 'HTTP %{http_code}\n' https://api.openai.com/v1/models
HTTP 401
The probe sends no credential, so /v1/models returns 401; curlStatus === 0 (curl received a response) makes it ok: true regardless. A real deployment gets the same unauthenticated 401 here — the probe never authenticates — so this reachability signal is what drives the healthy line, independent of whether the key or model actually works.
Impact
A green status inference line can coexist with a model that cannot be reached: missing/invalid credential, a configured model that is not pulled (local) or not enabled (remote), the model backend down, or — as in #6635 — a broken in-sandbox upstream hop while the local gateway still serves /v1/models. The user sees "healthy" and looks elsewhere.
Scope / honesty
Reproduced by execution on ollama-local and openai-api (host-side probes), x86_64/WSL2, NemoClaw main (3de1de6b1). The anthropic-prod path @mikemason hit uses the same /v1/models reachability dispatch (established from the code, not separately executed). This is a diagnostic-accuracy gap independent of the gateway-supervisor topic #6635 was closed on (#6677); filing separately rather than reopening.
Suggested direction
The repo already has the right primitive: probeNvidiaKimiK26Health exercises /chat/completions. Extending a lightweight, authenticated model-invocation hop (a minimal completion, or at least an authenticated /v1/models call that confirms the configured model is present) to the general provider health path would make status reflect what OpenClaw actually depends on. The current remote probe is deliberately unauthenticated, so the fix is to add authentication + a configured-model/invocation check rather than to reinterpret its HTTP status; for local providers, the probe should likewise confirm the configured model is in /api/tags rather than accepting any well-formed (even empty) list.
Offer
Happy to send a PR generalizing the invocation-style probe (guarded to stay cheap and offline-friendly for local providers) with tests. Confirm the direction — full completion probe vs. auth+configured-model validation — and I'll implement it.
Summary
nemoclaw statuscan report healthy inference while the configured model cannot actually be invoked. The inference health probes check that the provider endpoint answers an HTTP request (or that the local daemon lists models), not that the configured model responds to a real completion call. This confirms the observation @mikemason left on #6635 ("At the time OpenClaw couldn't reach the model,nemoclaw statuswas reporting healthy inference, so whatever it's checking is slightly different than what OpenClaw actually uses to hit the model") — and it reproduces without the corporate-proxy / TLS specifics of that thread.Mechanism (code)
For every provider except the NVIDIA Kimi K2.6 special case, the health probe is a reachability/listing check:
ollama-local/vllm-local) —probeLocalProviderHealth(provider, options)(src/lib/inference/local.ts) does not take a model argument at all. It hits/api/tagsand returns healthy when the body parses as{ "models": [...] }— and an empty array is explicitly accepted ("just means no models pulled yet",isValidOllamaTagsResponseBody). Daemon-up ⇒ healthy, regardless of the configured model.anthropic-prod,openai-api,gemini-api, non-Kiminvidia-*) —getRemoteProviderHealthEndpointreturns the provider's/v1/modelslisting URL, andprobeRemoteProviderHealthcurls it with no credential (["-sS", "--connect-timeout", "3", "--max-time", "5", endpoint]), then setsok = (curlStatus === 0)with the comment "Even a 401/403 means the endpoint is reachable" (src/lib/inference/health.ts). Because it is unauthenticated, any HTTP answer — including401 Unauthorized— reads as healthy: the check confirms the host is network-reachable, nothing about credential validity or model invocability.NVIDIA_MANAGED_PROVIDERS+isKimiK26Modelroutes toprobeNvidiaKimiK26Health, which hits/chat/completions— an actual invocation. That is the one path that reflects model invocability.The in-sandbox probe (
probeSandboxInferenceGatewayHealth,inference-route-health.ts) is likewise a route-reachability check onhttps://inference.local/v1/models, not a completion call.Reproduction (by execution, proxy-independent)
Driving the compiled probes on x86_64/WSL2, no proxy dependence:
Local — the probe is model-blind:
probeLocalProviderHealthtakes no model parameter, so a sandbox configured for a model that is not pulled still probesok: truewhile a real call errorsmodel not found.Remote — a 401 reads as healthy:
The probe sends no credential, so
/v1/modelsreturns401;curlStatus === 0(curl received a response) makes itok: trueregardless. A real deployment gets the same unauthenticated401here — the probe never authenticates — so this reachability signal is what drives the healthy line, independent of whether the key or model actually works.Impact
A green
statusinference line can coexist with a model that cannot be reached: missing/invalid credential, a configured model that is not pulled (local) or not enabled (remote), the model backend down, or — as in #6635 — a broken in-sandbox upstream hop while the local gateway still serves/v1/models. The user sees "healthy" and looks elsewhere.Scope / honesty
Reproduced by execution on
ollama-localandopenai-api(host-side probes), x86_64/WSL2, NemoClawmain(3de1de6b1). Theanthropic-prodpath @mikemason hit uses the same/v1/modelsreachability dispatch (established from the code, not separately executed). This is a diagnostic-accuracy gap independent of the gateway-supervisor topic #6635 was closed on (#6677); filing separately rather than reopening.Suggested direction
The repo already has the right primitive:
probeNvidiaKimiK26Healthexercises/chat/completions. Extending a lightweight, authenticated model-invocation hop (a minimal completion, or at least an authenticated/v1/modelscall that confirms the configured model is present) to the general provider health path would makestatusreflect what OpenClaw actually depends on. The current remote probe is deliberately unauthenticated, so the fix is to add authentication + a configured-model/invocation check rather than to reinterpret its HTTP status; for local providers, the probe should likewise confirm the configured model is in/api/tagsrather than accepting any well-formed (even empty) list.Offer
Happy to send a PR generalizing the invocation-style probe (guarded to stay cheap and offline-friendly for local providers) with tests. Confirm the direction — full completion probe vs. auth+configured-model validation — and I'll implement it.