Skip to content

nemoclaw status reports inference healthy on endpoint reachability, not model invocability — a green status can mask an unreachable model #6846

Description

@Hokonoken

Summary

nemoclaw status can report healthy inference while the configured model cannot actually be invoked. The inference health probes check that the provider endpoint answers an HTTP request (or that the local daemon lists models), not that the configured model responds to a real completion call. This confirms the observation @mikemason left on #6635 ("At the time OpenClaw couldn't reach the model, nemoclaw status was reporting healthy inference, so whatever it's checking is slightly different than what OpenClaw actually uses to hit the model") — and it reproduces without the corporate-proxy / TLS specifics of that thread.

Mechanism (code)

For every provider except the NVIDIA Kimi K2.6 special case, the health probe is a reachability/listing check:

  • Local (ollama-local / vllm-local) — probeLocalProviderHealth(provider, options) (src/lib/inference/local.ts) does not take a model argument at all. It hits /api/tags and returns healthy when the body parses as { "models": [...] } — and an empty array is explicitly accepted ("just means no models pulled yet", isValidOllamaTagsResponseBody). Daemon-up ⇒ healthy, regardless of the configured model.
  • Remote (anthropic-prod, openai-api, gemini-api, non-Kimi nvidia-*) — getRemoteProviderHealthEndpoint returns the provider's /v1/models listing URL, and probeRemoteProviderHealth curls it with no credential (["-sS", "--connect-timeout", "3", "--max-time", "5", endpoint]), then sets ok = (curlStatus === 0) with the comment "Even a 401/403 means the endpoint is reachable" (src/lib/inference/health.ts). Because it is unauthenticated, any HTTP answer — including 401 Unauthorized — reads as healthy: the check confirms the host is network-reachable, nothing about credential validity or model invocability.
  • Only NVIDIA_MANAGED_PROVIDERS + isKimiK26Model routes to probeNvidiaKimiK26Health, which hits /chat/completions — an actual invocation. That is the one path that reflects model invocability.

The in-sandbox probe (probeSandboxInferenceGatewayHealth, inference-route-health.ts) is likewise a route-reachability check on https://inference.local/v1/models, not a completion call.

Reproduction (by execution, proxy-independent)

Driving the compiled probes on x86_64/WSL2, no proxy dependence:

Local — the probe is model-blind:

$ node -e 'console.log(require("./dist/lib/inference/local.js").probeLocalProviderHealth("ollama-local"))'
{
  ok: true,
  providerLabel: 'Local Ollama',
  endpoint: 'http://127.0.0.1:11434/api/tags',
  detail: 'Local Ollama is reachable on http://127.0.0.1:11434/api/tags.',
  probeLabel: 'ollama backend',
  subprobes: [ { probeLabel: 'auth proxy', ok: false, ... } ]   // auth-proxy hop down in this lab; incidental — the point is the top-level model-blind `ok: true`
}

# but an actual call to a model that is not present fails:
$ curl -s http://127.0.0.1:11434/api/chat -d '{"model":"nonexistent-model:latest","messages":[{"role":"user","content":"hi"}],"stream":false}'
{"error":"model 'nonexistent-model:latest' not found"}

probeLocalProviderHealth takes no model parameter, so a sandbox configured for a model that is not pulled still probes ok: true while a real call errors model not found.

Remote — a 401 reads as healthy:

$ node -e 'console.log(require("./dist/lib/inference/health.js").probeProviderHealth("openai-api"))'
{
  ok: true,
  probed: true,
  providerLabel: 'OpenAI',
  endpoint: 'https://api.openai.com/v1/models',
  detail: 'OpenAI endpoint is reachable at https://api.openai.com/v1/models.'
}

# the endpoint actually answered unauthenticated:
$ curl -s -o /dev/null -w 'HTTP %{http_code}\n' https://api.openai.com/v1/models
HTTP 401

The probe sends no credential, so /v1/models returns 401; curlStatus === 0 (curl received a response) makes it ok: true regardless. A real deployment gets the same unauthenticated 401 here — the probe never authenticates — so this reachability signal is what drives the healthy line, independent of whether the key or model actually works.

Impact

A green status inference line can coexist with a model that cannot be reached: missing/invalid credential, a configured model that is not pulled (local) or not enabled (remote), the model backend down, or — as in #6635 — a broken in-sandbox upstream hop while the local gateway still serves /v1/models. The user sees "healthy" and looks elsewhere.

Scope / honesty

Reproduced by execution on ollama-local and openai-api (host-side probes), x86_64/WSL2, NemoClaw main (3de1de6b1). The anthropic-prod path @mikemason hit uses the same /v1/models reachability dispatch (established from the code, not separately executed). This is a diagnostic-accuracy gap independent of the gateway-supervisor topic #6635 was closed on (#6677); filing separately rather than reopening.

Suggested direction

The repo already has the right primitive: probeNvidiaKimiK26Health exercises /chat/completions. Extending a lightweight, authenticated model-invocation hop (a minimal completion, or at least an authenticated /v1/models call that confirms the configured model is present) to the general provider health path would make status reflect what OpenClaw actually depends on. The current remote probe is deliberately unauthenticated, so the fix is to add authentication + a configured-model/invocation check rather than to reinterpret its HTTP status; for local providers, the probe should likewise confirm the configured model is in /api/tags rather than accepting any well-formed (even empty) list.

Offer

Happy to send a PR generalizing the invocation-style probe (guarded to stay cheap and offline-friendly for local providers) with tests. Confirm the direction — full completion probe vs. auth+configured-model validation — and I'll implement it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area: cliCommand line interface, flags, terminal UX, or outputarea: inferenceInference routing, serving, model selection, or outputsarea: local-modelsLocal model providers, downloads, launch, or connectivityarea: providersInference provider integrations and provider behavior

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions