Skip to content

fix(onboard): preflight NEMOCLAW_VLLM_MODEL before side effects (#5207) - #5214

Merged
cv merged 3 commits into
mainfrom
fix/onboard-vllm-model-preflight-5207
Jun 13, 2026
Merged

cv merged 3 commits into
mainfrom
fix/onboard-vllm-model-preflight-5207

Conversation

@jason-ma-nv

@jason-ma-nv jason-ma-nv commented Jun 11, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Non-interactive nemoclaw onboard now validates NEMOCLAW_VLLM_MODEL up front, so an unrecognised (or gated, token-less) slug fails fast with a non-zero exit code and the canonical, slug-listing error message — before preflight, Docker, or any sandbox side effects.

Related Issue

Fixes #5207

Changes

  • src/lib/onboard.ts: call preflightVllmModelEnv() early in onboard(), alongside the existing early NEMOCLAW_PROVIDER validation (before acquireOnboardLock/preflight). On failure it prints the installer's error verbatim and exits 1. This mirrors the connect preflight added in fix(inference): preflight NEMOCLAW_VLLM_MODEL on sandbox connect #4567 (NVBug 6282062) for the analogous code path the issue calls out.
  • test/onboard-vllm-model-preflight.test.ts: regression test — a non-interactive onboard with NEMOCLAW_VLLM_MODEL=not-a-real-model exits 1, surfaces the slug error on stderr, and never reaches the [1/8] Preflight checks step (proving the fast-fail happens before any side effects).

Root cause

NEMOCLAW_VLLM_MODEL was only consulted deep inside the express-vLLM installer (resolveVllmInstallModel, the [3/8] provider step). That made the variable validated late and path-dependent: any onboard path that does not run the installer silently ignored an invalid slug, so the failure was not reliably surfaced as a non-zero exit. The fix adds one up-front, path-independent validation surface, exactly as connect got in #4567.

Type of Change

  • Code change (feature, bug fix, or refactor)
  • Code change with doc updates
  • Doc only (prose changes, no code sample modifications)
  • Doc only (includes code sample changes)

Verification

  • npx prek run --all-files passes
  • npm test passes
  • Tests added or updated for new or changed behavior
  • No secrets, API keys, or credentials committed
  • Docs updated for user-facing behavior changes
  • npm run docs builds without warnings (doc changes only)
  • Doc pages follow the style guide (doc changes only)
  • New doc pages include SPDX header and frontmatter (new pages only)

Signed-off-by: Jason Ma jama@nvidia.com

Summary by CodeRabbit

  • Bug Fixes

    • Onboarding now validates the configured vLLM model and provider hint early and exits immediately with a clear error if values are unrecognized, preventing any further onboarding steps from running.
  • Tests

    • Added a regression test confirming non-interactive onboarding fails fast for an invalid vLLM model and that subsequent onboarding steps are not executed.

A non-interactive `nemoclaw onboard` with an unrecognised
NEMOCLAW_VLLM_MODEL slug only validated the value deep inside the
express-vLLM installer (the [3/8] provider step). The variable was thus
checked late and ignored entirely on non-installer paths, so onboarding
could run past the failure instead of failing fast with a non-zero exit.

Validate the variable up front in onboard(), mirroring the connect
preflight added in #4567: run the installer's selectVllmModelFromEnv +
assertGatedModelAccess checks before preflight/Docker and exit non-zero
with the canonical, slug-listing error message.

Signed-off-by: Jason Ma <jama@nvidia.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@jason-ma-nv jason-ma-nv self-assigned this Jun 11, 2026
@coderabbitai

coderabbitai Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

The PR adds an early call to preflightVllmModelEnvOrExit() during onboarding to validate NEMOCLAW_VLLM_MODEL (and provider hint), exiting with status 1 on invalid slugs; it also adds a regression test asserting the fast-fail happens before any preflight steps run.

Changes

vLLM Model Validation in Onboarding

Layer / File(s) Summary
vLLM preflight module
src/lib/onboard/vllm-model-preflight.ts
Adds preflightVllmModelEnvOrExit(env?) that runs preflightVllmModelEnv and prints an error + exits(1) on failure.
Resume config helper
src/lib/onboard/resume-config.ts
Imports the vLLM preflight helper and adds preflightEarlyOnboardEnv(nonInteractive?) which runs the vLLM preflight and returns the requested provider hint.
Onboard entrypoint wiring
src/lib/onboard.ts
Calls resumeConfig.preflightEarlyOnboardEnv() early in onboard() to validate provider and NEMOCLAW_VLLM_MODEL before further onboarding steps.
Regression: fast-fail on invalid vLLM slug
test/onboard-vllm-model-preflight.test.ts
Adds a Vitest regression test that spawns the compiled onboarding entrypoint with a bad NEMOCLAW_VLLM_MODEL, asserts exit code 1, stderr contains the slug and error, and confirms preflight step output does not appear.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related issues

Suggested labels

bug-fix, area: onboarding, provider: vllm

Suggested reviewers

  • cv
  • prekshivyas

Poem

🐰 I sniffed the env with a twitch and a twirl,
A slug hopped in that wasn't my pearl.
"Not valid," I said with a thump and a spin,
Exit one — no Docker to begin.
The rabbit hops off, mission done, cheeky grin.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 60.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely summarizes the main change: prefighting NEMOCLAW_VLLM_MODEL validation before side effects occur in the onboard process.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/onboard-vllm-model-preflight-5207

Comment @coderabbitai help to get the list of available commands and usage tips.

…l-preflight-5207

# Conflicts:
#	src/lib/onboard.ts

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lib/onboard.ts`:
- Around line 5699-5711: The added preflight block in onboard.ts
(preflightVllmModelEnv + vllmModelPreflight check) increases file size; extract
this wiring into a small helper to keep onboard.ts within budget. Create a new
helper (e.g., ensureVllmModelPreflight or validateVllmModelEnv) under
src/lib/onboard/* that calls preflightVllmModelEnv and exits on failure, then
replace the inline block in onboard.ts with a single call to that helper; keep
references to preflightVllmModelEnv, selectVllmModelFromEnv and
assertGatedModelAccess unchanged so tests and behavior remain the same.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 9f24e718-017a-433d-88e1-8c0bb0ab08a4

📥 Commits

Reviewing files that changed from the base of the PR and between 3d5dab6 and bebf9c7.

📒 Files selected for processing (2)
  • src/lib/onboard.ts
  • test/onboard-vllm-model-preflight.test.ts

Comment thread src/lib/onboard.ts Outdated
@github-actions

github-actions Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

E2E Advisor Recommendation

Required E2E: onboard-negative-paths-e2e
Optional E2E: onboard-resume-e2e, cloud-onboard-e2e, onboard-negative-paths-vitest

Dispatch hint: onboard-negative-paths-e2e

Workflow run

Full advisor summary

E2E Recommendation Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required E2E

  • onboard-negative-paths-e2e (medium): Closest existing merge-blocking E2E for non-interactive onboard validation and negative-path CLI exit/output behavior. It should catch broad regressions in early onboard validation ordering, clean error handling, and normal recovery after negative cases.

Optional E2E

  • onboard-resume-e2e (medium): resume-config.ts was modified and the new early validation runs before lock/session continuation. This gives confidence that --resume still completes with recorded configuration and credential hydration.
  • cloud-onboard-e2e (high): Useful happy-path confidence that a normal non-interactive cloud onboarding run is not blocked when NEMOCLAW_VLLM_MODEL is unset and provider validation still behaves as expected.
  • onboard-negative-paths-vitest (low): Fast live Vitest coverage for the direct CLI/non-interactive onboard negative boundary. It is narrower than the bash E2E and does not cover the new vLLM slug specifically, but is a useful adjacent smoke check.

New E2E recommendations

  • onboarding / vLLM model env validation (high): No existing live E2E appears to set NEMOCLAW_PROVIDER=install-vllm with an invalid NEMOCLAW_VLLM_MODEL and assert a non-zero fail-fast before '[1/8] Preflight checks'. The PR adds unit coverage, but an E2E would protect the real packaged CLI/dispatch path and side-effect ordering.
    • Suggested test: Add a focused onboard vLLM model preflight negative-path E2E, ideally as a free-standing e2e-vitest-scenarios job or as a case in test/e2e/test-onboard-negative-paths.sh, asserting invalid NEMOCLAW_VLLM_MODEL exits non-zero before Docker/OpenShell preflight and does not create sandbox/session side effects.
  • resume / vLLM env conflict behavior (medium): The new validation runs before resume conflict handling; there is no apparent E2E covering onboard --resume with NEMOCLAW_VLLM_MODEL set or invalid. Clarifying expected behavior would prevent regressions where resume is blocked by unrelated env state.
    • Suggested test: Add a resume-focused E2E case that documents whether invalid NEMOCLAW_VLLM_MODEL should block --resume and verifies the session remains unchanged on failure.

Dispatch hint

  • Workflow: nightly-e2e.yaml
  • jobs input: onboard-negative-paths-e2e

@github-actions

github-actions Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

Vitest E2E Scenario Recommendation

Required Vitest E2E scenarios: ubuntu-repo-cloud-openclaw
Optional Vitest E2E scenarios: None

Dispatch required Vitest E2E scenarios:

  • gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field scenarios=ubuntu-repo-cloud-openclaw

Workflow run

Full Vitest E2E advisor summary

Vitest E2E Scenario Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required Vitest E2E scenarios

  • ubuntu-repo-cloud-openclaw: The PR changes the shared onboard entry path and early resume/onboard environment validation. This scenario is the smallest live-supported Vitest registry scenario that runs real Ubuntu repo Docker onboarding through the cloud OpenClaw path and verifies the normal no-op behavior when NEMOCLAW_VLLM_MODEL is unset.
    • Dispatch: gh workflow run e2e-vitest-scenarios.yaml --ref <pr-head-ref> --field scenarios=ubuntu-repo-cloud-openclaw

Optional Vitest E2E scenarios

  • None.

Relevant changed files

  • src/lib/onboard.ts
  • src/lib/onboard/resume-config.ts
  • src/lib/onboard/vllm-model-preflight.ts

@github-actions

github-actions Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

PR Review Advisor

Findings: 0 needs attention, 0 worth checking, 0 nice ideas
Since last review: 0 prior items resolved, 0 still apply, 0 new items found

Consider writing more tests for
  • **Runtime validation** — onboard with `NEMOCLAW_VLLM_MODEL=not-a-real-model` prints the recognised vLLM slug list on stderr. The added child-process regression exercises the actual onboard export and is appropriate for this narrow host-side validation change. Additional behavior-specific runtime assertions would increase confidence around the sandbox/Docker side-effect boundary and fully cover prior test-depth suggestions, but no missing blocker-level coverage was found.
  • **Runtime validation** — onboard with `NEMOCLAW_VLLM_MODEL=deepseek-r1-distill-70b` and no `HF_TOKEN` exits 1 before `[1/8] Preflight checks`. The added child-process regression exercises the actual onboard export and is appropriate for this narrow host-side validation change. Additional behavior-specific runtime assertions would increase confidence around the sandbox/Docker side-effect boundary and fully cover prior test-depth suggestions, but no missing blocker-level coverage was found.
  • **Runtime validation** — onboard invalid vLLM model does not invoke Docker/OpenShell runner calls or create a `nemoclaw-vllm` container. The added child-process regression exercises the actual onboard export and is appropriate for this narrow host-side validation change. Additional behavior-specific runtime assertions would increase confidence around the sandbox/Docker side-effect boundary and fully cover prior test-depth suggestions, but no missing blocker-level coverage was found.

Workflow run details

This is an automated advisory review. A human maintainer must make the final merge decision.

Keep src/lib/onboard.ts net-neutral per the codebase-growth guardrail: the
NEMOCLAW_VLLM_MODEL fast-fail check now lives in
src/lib/onboard/vllm-model-preflight.ts and is composed with the early
NEMOCLAW_PROVIDER hint check via resume-config's preflightEarlyOnboardEnv().
Behavior is unchanged from the previous commit.

Signed-off-by: Jason Ma <jama@nvidia.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lib/onboard.ts`:
- Around line 5060-5064: The call to resumeConfig.preflightEarlyOnboardEnv()
omits the nonInteractive flag so it defaults to false and skips provider-hint
validation during non-interactive runs; update the call to pass the current
non-interactive setting (e.g.,
resumeConfig.preflightEarlyOnboardEnv(nonInteractive) or the appropriate
options.nonInteractive value available in scope) so preflightEarlyOnboardEnv
receives the correct boolean and runs the non-interactive validation path.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: f70751aa-2811-432e-bbf0-aac4d81e9327

📥 Commits

Reviewing files that changed from the base of the PR and between 875cd46 and 4f074f1.

📒 Files selected for processing (3)
  • src/lib/onboard.ts
  • src/lib/onboard/resume-config.ts
  • src/lib/onboard/vllm-model-preflight.ts

Comment thread src/lib/onboard.ts
Comment on lines +5060 to +5064
// Validate NEMOCLAW_PROVIDER and NEMOCLAW_VLLM_MODEL early so invalid values
// fail before preflight (Docker/OpenShell checks). Without this, users see a
// misleading 'Docker is not reachable' error instead of the real
// problem: an unsupported provider value.
getRequestedProviderHint();
// problem: an unsupported provider value or unrecognised vLLM model slug.
resumeConfig.preflightEarlyOnboardEnv();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Pass the non-interactive mode into early env preflight.

preflightEarlyOnboardEnv() defaults nonInteractive to false, so this call skips the provider-hint validation path in non-interactive onboard runs.

🔧 Minimal fix
-  resumeConfig.preflightEarlyOnboardEnv();
+  resumeConfig.preflightEarlyOnboardEnv(isNonInteractive());
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
// Validate NEMOCLAW_PROVIDER and NEMOCLAW_VLLM_MODEL early so invalid values
// fail before preflight (Docker/OpenShell checks). Without this, users see a
// misleading 'Docker is not reachable' error instead of the real
// problem: an unsupported provider value.
getRequestedProviderHint();
// problem: an unsupported provider value or unrecognised vLLM model slug.
resumeConfig.preflightEarlyOnboardEnv();
// Validate NEMOCLAW_PROVIDER and NEMOCLAW_VLLM_MODEL early so invalid values
// fail before preflight (Docker/OpenShell checks). Without this, users see a
// misleading 'Docker is not reachable' error instead of the real
// problem: an unsupported provider value or unrecognised vLLM model slug.
resumeConfig.preflightEarlyOnboardEnv(isNonInteractive());
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/lib/onboard.ts` around lines 5060 - 5064, The call to
resumeConfig.preflightEarlyOnboardEnv() omits the nonInteractive flag so it
defaults to false and skips provider-hint validation during non-interactive
runs; update the call to pass the current non-interactive setting (e.g.,
resumeConfig.preflightEarlyOnboardEnv(nonInteractive) or the appropriate
options.nonInteractive value available in scope) so preflightEarlyOnboardEnv
receives the correct boolean and runs the non-interactive validation path.

@cv cv added v0.0.65 and removed v0.0.64 labels Jun 12, 2026
@wscurran wscurran added area: local-models Local model providers, downloads, launch, or connectivity area: onboarding Onboarding FSM, provider setup, sandbox launch, or first-run flow bug-fix PR fixes a bug or regression provider: vllm vLLM local or hosted provider behavior v0.0.64 labels Jun 12, 2026
@wscurran

Copy link
Copy Markdown
Contributor

@wscurran
wscurran requested a review from cv June 12, 2026 14:21
@wscurran wscurran removed the v0.0.64 label Jun 12, 2026
@cv
cv merged commit a203b72 into main Jun 13, 2026
43 checks passed
@cv
cv deleted the fix/onboard-vllm-model-preflight-5207 branch June 13, 2026 05:37
@miyoungc miyoungc mentioned this pull request Jun 16, 2026
5 of 13 tasks
cv pushed a commit that referenced this pull request Jun 17, 2026
## Summary
Refreshes release-prep documentation for NemoClaw v0.0.65.
Adds the v0.0.65 release-notes section and refreshes generated
`nemoclaw-user-*` skills from the Fern MDX source docs.

## Changes
- Added the v0.0.65 release notes to `docs/about/release-notes.mdx` with
links to the deeper docs pages for lifecycle, troubleshooting,
inference, CLI commands, messaging, credentials, network policy, Hermes,
and sub-agents.
- Regenerated the `nemoclaw-user-*` skills with
`scripts/docs-to-skills.py` so release-prep skill output matches the
merged source docs.
- Used the v0.0.65 announcement discussion as release context:
#5472.

## Source Summary
- #2492 -> `docs/about/release-notes.mdx`: Documents deadline-based
gateway wait reliability in the v0.0.65 recovery summary.
- #4958 -> `docs/about/release-notes.mdx`: Documents re-execed OpenClaw
gateway health check recovery in the sandbox recovery summary.
- #5163 -> `docs/about/release-notes.mdx`: Documents safer uninstall TTY
confirmation behavior in the day-two CLI summary.
- #5178 -> `docs/about/release-notes.mdx`: Documents fail-closed config
restore merge behavior in the rebuild and restore summary.
- #5179 -> `docs/about/release-notes.mdx`: Documents WeChat QR token
redaction in the messaging summary.
- #5182 -> `docs/about/release-notes.mdx`: Documents sustained gateway
serving checks in the recovery summary.
- #5194 -> `docs/about/release-notes.mdx`: Documents model-router
teardown during uninstall in the day-two CLI summary.
- #5195 -> `docs/about/release-notes.mdx`: Documents Shields
auto-restore lock reconfirmation in the rebuild and restore summary.
- #5198 -> `docs/about/release-notes.mdx`: Documents Docker Desktop WSL
CDI injection failure handling in the onboarding diagnostics summary.
- #5201 -> `docs/about/release-notes.mdx`: Documents sandbox
download/upload wrappers and sessions export in the day-two CLI summary.
- #5205 -> `docs/about/release-notes.mdx`: Documents reporter-owned
model metadata preservation in the rebuild and restore summary.
- #5214 -> `docs/about/release-notes.mdx`: Documents managed vLLM model
preflight before side effects in the inference setup summary.
- #5215 -> `docs/about/release-notes.mdx`: Documents managed vLLM extra
serve arguments in the inference setup summary.
- #5216 -> `docs/about/release-notes.mdx`: Documents silent OpenClaw
runtime fallback surfacing in the onboarding diagnostics summary.
- #5225 -> `docs/about/release-notes.mdx`: Documents persisted sandbox
gateway lookup in the gateway recovery summary.
- #5238 -> `docs/about/release-notes.mdx`: Documents sub-agent gateway
dial-back through the sandbox interface in the Hermes and sub-agent
summary.
- #5248 -> `docs/about/release-notes.mdx`: Documents Discord per-account
proxy resolution in the messaging summary.
- #5264 -> `docs/about/release-notes.mdx`: Documents reserved Hermes
port `8642` handling in the Hermes compatibility summary.
- #5267 -> `docs/about/release-notes.mdx`: Documents the narrower Hermes
baseline policy in the Hermes compatibility summary.
- #5321 -> `docs/about/release-notes.mdx`: Documents restored gateway
guard chains in the gateway recovery summary.
- #5328 -> `docs/about/release-notes.mdx`: Documents compact persisted
messaging plans in the messaging summary.
- #5338 -> `docs/about/release-notes.mdx`: Documents manifest channel
migration in the messaging summary.
- #5352 -> `docs/about/release-notes.mdx`: Documents persisted agent
preservation through registry recovery in the rebuild and restore
summary.
- #5371 ->
`.agents/skills/nemoclaw-user-reference/references/commands.md`:
Refreshes generated skill output for custom build cache and
layer-ordering source docs.
- #5379 -> `docs/about/release-notes.mdx`: Documents dashboard port
allocation across multiple NemoClaw gateways in the recovery summary.
- #5382 -> `docs/about/release-notes.mdx`: Documents recovery when an
active gateway has no sandbox spec in the recovery summary.
- #5389 ->
`.agents/skills/nemoclaw-user-reference/references/troubleshooting.md`:
Refreshes generated skill output for declared agent `forward_ports`
recovery source docs.
- #5400 -> `docs/about/release-notes.mdx`: Documents bounded compatible
endpoint probes in the inference setup summary.
- #5410 -> `docs/about/release-notes.mdx`: Documents provider credential
hash removal from sandbox registry entries in the messaging summary.
- #5418 -> `docs/about/release-notes.mdx`: Documents summarized
inference validation failures in the onboarding diagnostics summary.
- #5457 -> `docs/about/release-notes.mdx`: Documents context-window
recomputation after runtime model switches in the inference setup
summary.
- #5463 -> `docs/about/release-notes.mdx`: Documents cleanup of
hard-coded messaging channel stragglers in the messaging summary.

## Skipped
- #5366 matched `docs/.docs-skip` entries through skipped experimental
paths, so this PR does not add new release-note text for that commit.

## Type of Change
- [ ] Code change (feature, bug fix, or refactor)
- [ ] Code change with doc updates
- [ ] Doc only (prose changes, no code sample modifications)
- [x] Doc only (includes code sample changes)

## Verification
- [x] Git hooks passed during commit and push, or `npx prek run
--from-ref main --to-ref HEAD` passes
- [ ] Targeted tests pass for changed behavior
- [ ] Full `npm test` passes (broad runtime changes only)
- [ ] Tests added or updated for new or changed behavior
- [x] No secrets, API keys, or credentials committed
- [x] Docs updated for user-facing behavior changes
- [ ] `npm run docs` builds without warnings (doc changes only)
- [x] Doc pages follow the [style
guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md)
(doc changes only)
- [ ] New doc pages include SPDX header and frontmatter (new pages only)

Verification notes:
- `npm run docs` passed after rerunning outside the sandbox. Fern
reported 0 errors and 1 hidden warning.
- The first sandboxed `npm run docs` attempt failed before validation
because `tsx` could not create its local IPC pipe under sandbox
restrictions.
- `npm run build:cli` passed before push to refresh the local `dist/`
artifacts used by the CLI typecheck hook.
- `npm test` was not run because this is a docs-only release refresh.

---
Signed-off-by: Miyoung Choi <miyoungc@nvidia.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Released NemoClaw v0.0.65 with improved gateway/sandbox recovery,
safer day-two workflows, and enhanced Hermes compatibility.
* Added managed vLLM extra-arguments configuration via
`NEMOCLAW_VLLM_EXTRA_ARGS_JSON`.
* Added Hermes troubleshooting guidance for port forwarding and health
checks.

* **Documentation**
* Updated NVIDIA Endpoints/NIM setup and examples to use
`NVIDIA_INFERENCE_API_KEY`.
* Refined NVIDIA network policy and Model Router API base configuration.
* Expanded CLI/environment variable documentation (including sub-agent
gateway connectivity) and plugin build performance tips.

* **Tests**
  * Expanded Vitest-backed E2E release validation coverage.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
@wscurran wscurran added the NV QA Bugs found by the NVIDIA QA Team label Jun 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: local-models Local model providers, downloads, launch, or connectivity area: onboarding Onboarding FSM, provider setup, sandbox launch, or first-run flow bug-fix PR fixes a bug or regression NV QA Bugs found by the NVIDIA QA Team provider: vllm vLLM local or hosted provider behavior

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[DGX Spark][Onboard] nemoclaw onboard exits 0 when NEMOCLAW_VLLM_MODEL set to unrecognised slug

3 participants