Skip to content

docs(inference): correct validation API paths - #4249

Merged
cv merged 3 commits into
NVIDIA:mainfrom
rluo8:docs/3691-inference-options-validation-table
May 27, 2026
Merged

cv merged 3 commits into
NVIDIA:mainfrom
rluo8:docs/3691-inference-options-validation-table

Conversation

@rluo8

@rluo8 rluo8 commented May 26, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Updates the inference options docs so provider validation behavior matches the current onboarding implementation. In particular, NVIDIA Endpoints, Google Gemini, and Local NVIDIA NIM no longer claim to try /v1/responses first.

Related Issue

Fixes #3691

Changes

  • Clarified that NVIDIA Endpoints validate via /v1/chat/completions and skip /v1/responses because NVIDIA Build does not expose it.
  • Clarified that Google Gemini uses its OpenAI-compatible chat-completions path and skips the Responses API probe.
  • Clarified that Local NVIDIA NIM skips the /v1/responses probe.
  • Distinguished compatible-endpoint probe behavior from the selected runtime API default.

Type of Change

  • Code change (feature, bug fix, or refactor)
  • Code change with doc updates
  • [√] Doc only (prose changes, no code sample modifications)
  • Doc only (includes code sample changes)

Verification

  • npx prek run --all-files passes
  • npm test passes
  • Tests added or updated for new or changed behavior
  • [√] No secrets, API keys, or credentials committed
  • [√] Docs updated for user-facing behavior changes
  • [√] make docs builds without warnings (doc changes only)
  • Doc pages follow the style guide (doc changes only)
  • New doc pages include SPDX header and frontmatter (new pages only)

Signed-off-by: rluo8 ruluo@nvidia.com

Summary by CodeRabbit

  • Documentation
    • Clarified API routing behavior and validation procedures across different inference providers, including OpenAI-compatible endpoints, Google Gemini, and NVIDIA Endpoints
    • Updated configuration guidance for preferred API endpoint selection

Review Change Stack

@coderabbitai

coderabbitai Bot commented May 26, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 4557f064-0e0b-449e-ace0-7d0f61027e1e

📥 Commits

Reviewing files that changed from the base of the PR and between 206737f and dad9163.

📒 Files selected for processing (1)
  • docs/inference/inference-options.mdx

📝 Walkthrough

Walkthrough

Documentation updates to the inference options reference guide clarify API routing defaults and validation methods. Provider Options and Validation tables are revised to reflect actual runtime behavior: /v1/chat/completions defaults for all providers, with /v1/responses available only via explicit environment variable configuration, and provider-specific probe/skip behavior during validation.

Changes

Inference API Documentation Updates

Layer / File(s) Summary
Provider Options API defaults and environment variable behavior
docs/inference/inference-options.mdx
"Other OpenAI-compatible endpoint" and "Google Gemini" entries clarify that runtime defaults to /v1/chat/completions, NEMOCLAW_PREFERRED_API=openai-responses enables /v1/responses only for supported proxies, Telegram onboarding smoke-checks against https://inference.local/v1/chat/completions, and Gemini explicitly skips /v1/responses probing.
Validation methods table provider API support
docs/inference/inference-options.mdx
NVIDIA Endpoints and Local NVIDIA NIM validate via /v1/chat/completions only (skipping /v1/responses); Google Gemini validates via chat-completions only (skipping responses); OpenAI-compatible/compatible endpoints validate by trying /v1/responses first with fallback to /v1/chat/completions, while runtime still defaults to /v1/chat/completions unless NEMOCLAW_PREFERRED_API=openai-responses is set after successful validation.

🎯 2 (Simple) | ⏱️ ~10 minutes

🐰 A rabbit hops through docs so fine,
API routes now clearly align,
Chat completions by default,
Responses opt-in—a gentle cult,
Validation paths now redesign! ✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and accurately describes the main change: correcting API path documentation for validation behavior across multiple providers.
Linked Issues check ✅ Passed The PR changes fully address issue #3691 requirements: NVIDIA Endpoints and Local NIM now document /v1/chat/completions only, Google Gemini clarified to skip /v1/responses, and Other OpenAI-compatible endpoints clarify the runtime default and explicit opt-in behavior via NEMOCLAW_PREFERRED_API.
Out of Scope Changes check ✅ Passed All changes are documentation updates directly related to correcting validation API paths as specified in issue #3691; no unrelated or out-of-scope modifications are present.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Warning

There were issues while running some tools. Please review the errors and either fix the tool's configuration or disable the tool if it's a critical failure.

🔧 ESLint

If the error stems from missing dependencies, add them to the package.json file. For unrecoverable errors (e.g., due to private dependencies), disable the tool in the CodeRabbit configuration.

ESLint skipped: no ESLint configuration detected in root package.json. To enable, add eslint to devDependencies.


Comment @coderabbitai help to get the list of available commands and usage tips.

@wscurran

Copy link
Copy Markdown
Contributor

@rluo8 rluo8 added the v0.0.52 label May 27, 2026
@cv cv added v0.0.53 and removed v0.0.52 labels May 27, 2026
@cv
cv merged commit 5f9a23c into NVIDIA:main May 27, 2026
25 checks passed
@rluo8
rluo8 deleted the docs/3691-inference-options-validation-table branch June 1, 2026 04:59
@wscurran wscurran added area: inference Inference routing, serving, model selection, or outputs bug-fix PR fixes a bug or regression feature PR adds or expands user-visible functionality area: docs Documentation, examples, guides, or docs build chore Build, CI, dependency, or tooling maintenance and removed fix bug-fix PR fixes a bug or regression feature PR adds or expands user-visible functionality labels Jun 3, 2026
@wscurran wscurran added the NV QA Bugs found by the NVIDIA QA Team label Jun 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: docs Documentation, examples, guides, or docs build area: inference Inference routing, serving, model selection, or outputs chore Build, CI, dependency, or tooling maintenance NV QA Bugs found by the NVIDIA QA Team

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[All Platforms][Docs] inference/inference-options.html validation method has wrong api for nvidia-prod, nvidia-nim, and gemini-api

3 participants