Skip to content

feat(installer): add DGX Station express install - #6881

Closed
ericksoa wants to merge 11 commits into
mainfrom
feat/dgx-station-express-install
Closed

ericksoa wants to merge 11 commits into
mainfrom
feat/dgx-station-express-install

Conversation

@ericksoa

@ericksoa ericksoa commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

Summary

DGX Station's initial installer now offers one express-install confirmation and then completes onboarding without additional provider, model, policy, or sandbox-name choices. The express path translates the validated Nemotron playbook into a pinned managed-vLLM recipe for NVIDIA Nemotron 3 Ultra 550B while keeping direct managed-vLLM onboarding defaults unchanged.

Changes

  • Select nemotron-3-ultra-550b-a55b automatically when DGX Station express install is accepted.
  • Add a model-specific managed-vLLM runtime overlay for the pinned vLLM 0.22 index digest, Hugging Face revision, Station GPU selection, shared memory, CPU offload, parser, and compatibility settings while retaining NemoClaw's bridge-network boundary.
  • Preserve the canonical NemoClaw model route identity so existing Nemotron Ultra request adapters and agent configuration apply to the local server.
  • Add installer, model-registry, runtime, storage/download, provider-selection/FSM handoff, and documentation contract coverage.
  • Document the one-prompt flow and its approximately 352 GB model download while retaining DGX Station's Deferred status pending physical end-to-end validation.
  • Build on merged fix(install): restore DGX Station GB300 express setup #6875 for broader Station GB300 DMI detection and the managed-vLLM Docker-image storage-preflight contract.

Type of Change

  • Code change (feature, bug fix, or refactor)
  • Code change with doc updates
  • Doc only (prose changes, no code sample modifications)
  • Doc only (includes code sample changes)

Quality Gates

  • Tests added or updated for changed behavior
  • Existing tests cover changed behavior — justification:
  • Tests not applicable — justification:
  • Docs updated for user-facing behavior changes
  • Docs not applicable — justification:
  • Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging)
  • Sensitive-path review completed or maintainer-approved waiver recorded — reviewer/approval link/justification: Pending maintainer review of the installer and managed-inference path.
  • Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue:

Verification

  • PR description includes a Signed-off-by: line and every commit appears as Verified in GitHub
  • Normal pre-commit, commit-msg, and pre-push hooks passed, or npm run check:diff passed when hooks were skipped or unavailable
  • Targeted behavior tests pass for the current change set, or tests are marked not applicable above — vitest run src/lib/onboard/setup-nim-flow.test.ts src/lib/onboard/machine/handlers/provider-inference.test.ts src/lib/onboard/machine/transition-traces.test.ts src/lib/inference/vllm-models.test.ts src/lib/inference/vllm.test.ts test/inference-options-docs.test.ts test/install-express-prompt.test.ts --testTimeout=15000 (168 passed, 1 existing skip)
  • Applicable broad gate passed — npm test for broad runtime/test-harness changes; npm run check for repo-wide validation/coverage changes — not applicable; this bounded installer/managed-vLLM change is covered by targeted tests and normal hooks.
  • Quality Gates section completed with required justifications or waivers — sensitive-path maintainer review remains pending.
  • No secrets, API keys, or credentials committed
  • npm run docs builds without warnings (doc changes only) — npm run docs:strict passed with zero errors and two pre-existing Fern warnings.
  • Doc pages follow the style guide (doc changes only)
  • New doc pages include SPDX header and frontmatter (new pages only) — not applicable; no new pages.

Signed-off-by: Aaron Erickson aerickson@nvidia.com

Summary by CodeRabbit

  • New Features

    • Added support for DGX Station pinned Nemotron 3 Ultra 550B (NVFP4) managed vLLM, including immutable/digest-pinned runtime behavior and per-model runtime handling.
  • Documentation

    • Updated quickstart and inference/provider/platform docs with clearer DGX Station express-install outcomes, model slug/override rules (including blank/whitespace), and more precise image digests/download expectations.
  • Bug Fixes

    • Improved Docker image storage preflight behavior: it now checks image storage only and continues cleanly when capacity information is inconclusive.
  • Tests

    • Expanded DGX Station express-install, model selection, and runtime-profile mapping coverage (including override/whitespace cases and Station GB300 detection).

Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
@ericksoa ericksoa added area: docs Documentation, examples, guides, or docs build area: inference Inference routing, serving, model selection, or outputs platform: dgx-station Affects DGX Station hardware or workflows labels Jul 14, 2026
@ericksoa ericksoa self-assigned this Jul 14, 2026
@github-code-quality

github-code-quality Bot commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

Code Coverage Overview

Languages: TypeScript

TypeScript / code-coverage/plugin

The overall coverage remains at 96%, unchanged from the main branch.

TypeScript / code-coverage/cli

The overall coverage in the feat/dgx-station-exp... branch remains at 79%, unchanged from the main branch.

Show a code coverage summary of the most impacted files.
File main c181e9e feat/dgx-station-exp... f028538 +/-
src/lib/inference/vllm.ts 79% 80% +1%
src/lib/state/m...-acquisition.ts 89% 90% +1%
src/lib/inferen.../vllm-models.ts 66% 77% +11%
src/lib/actions...ge-preflight.ts 74% 89% +15%
src/lib/core/pr...mpt-activity.ts 67% 92% +25%

Updated July 14, 2026 19:54 UTC
Code Coverage is in Public Preview. Learn more and provide us with your feedback.

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

PR Review Advisor — Informational

Advisor assessment: Informational / medium confidence
Next action: Review the warnings below.
Findings: 0 blockers · 2 warnings · 0 suggestions
Status: Canonical ledger: 0 blocker(s), 2 warning(s), 0 suggestion(s).

Model lanes

  • GPT-5.6 Terra (primary): Completed · medium confidence · 0 blockers · 2 warnings · 0 suggestions
  • Nemotron 3 Ultra (second opinion): Failed after a partial review · low confidence · 0 blockers · 5 warnings · 1 suggestion

Nemotron output stays in workflow artifacts and does not change the assessment above.

E2E guidance

Advisory only. E2E / PR Gate selects and runs jobs independently.

Recommended E2E: cloud-onboard, inference-routing, network-policy

2 optional E2E recommendations
  • gpu-e2e
  • spark-install
2 warnings · 0 suggestions

Warnings

Warnings do not block.

PRA-1 Warning — Exercise the Station runtime through the local-inference policy boundary

  • Location: src/lib/inference/vllm.test.ts:163
  • Category: security
  • Problem: The new model-specific Station runtime has unit coverage for its image, Docker args, and bridge-network shape, but checked-in tests do not demonstrate that an agent can reach it through the advertised local-inference route while policy continues to deny unintended egress.
  • Impact: A runtime-specific networking or route integration regression could either leave the agent unable to use the selected provider or require an unsafe policy workaround to restore connectivity.
  • Recommendation: Add a focused integration or contract test for the Ultra runtime that verifies the normal sandbox-to-local-inference route succeeds and non-allowlisted egress remains denied.
  • Verification: Inspect the Station Ultra tests in vllm.test.ts and existing local-inference route/policy tests; confirm no test combines the new runtime selection with allow/deny policy behavior.
  • Test coverage: Add a regression test selecting nemotron-3-ultra-550b-a55b that proves the sandbox route reaches the managed service over the existing bridge topology and a non-preset external destination is rejected.
  • Evidence: src/lib/inference/vllm.test.ts:163-182 checks runtime image, image size, run flags, and docker arguments for Ultra. src/lib/inference/vllm-models.ts:234-236 explicitly preserves the bridge-networked local-inference boundary. The risk plan lists both inference-routing and network-policy invariants for the changed inference files.

PRA-2 Warning — Cover GB300 firmware detection in the installer test

  • Location: scripts/install.sh:2641
  • Category: tests
  • Problem: The new Station detector accepts *Station*GB300*, but the changed installer tests exercise Station express behavior only by replacing detect_express_platform. Their product-name helper does not assert that a real DMI product string containing Station GB300 selects DGX Station.
  • Impact: A future edit to the shell pattern can silently stop offering the intended Station express recipe on GB300 firmware while all current Station express tests continue to pass.
  • Recommendation: Add a product-name detection test using a representative Station GB300 DMI value, then assert it returns DGX Station and enters the existing Station express configuration path.
  • Verification: Inspect test/install-express-prompt.test.ts: detectExpressPlatformForProductName is available, while current detection assertions cover Spark/other product names but not the new GB300 pattern.
  • Test coverage: In test/install-express-prompt.test.ts, pass a representative e.g. 'NVIDIA DGX Station GB300' product name to detectExpressPlatformForProductName and assert stdout is exactly 'DGX Station'; optionally couple it to the existing accepted express-prompt Ultra override assertion.
  • Evidence: scripts/install.sh:2641 changes the Station case from *DGX*Station* to *DGX*Station* | *Station*GB300*. test/install-express-prompt.test.ts:95-134 provides a real shell product-name harness but Station express tests at :181-232 stub the detector directly. The static test inventory contains no named GB300 product-name detection test.

Workflow run details

This automated review informs maintainers. Warnings and suggestions do not require a response. A maintainer decides whether to merge.

Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
@coderabbitai

coderabbitai Bot commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

DGX Station express installation now selects a pinned Nemotron Ultra managed-vLLM recipe. Runtime handling, Docker image storage checks, installer detection, onboarding tests, and platform documentation were updated to support the new model and deferred hardware-validation status.

Changes

DGX Station Nemotron Ultra managed vLLM

Layer / File(s) Summary
Pinned model and runtime contract
src/lib/inference/vllm-models.ts, src/lib/inference/vllm-models.test.ts
Added pinned revision, served model ID, digest-pinned runtime metadata, NVFP4 configuration, serving flags, and registry tests.
Runtime installation and storage preflight
src/lib/inference/vllm.ts, src/lib/inference/vllm-storage.ts, src/lib/inference/*.test.ts
Applied model-specific runtime overrides through installation and refactored Docker image storage probing, warnings, and inconclusive-capacity behavior.
DGX Station express selection
scripts/install.sh, test/install-express-prompt.test.ts, src/lib/onboard/setup-nim-flow.test.ts
Added Station GB300 detection, pinned default model export, override handling, and non-interactive onboarding coverage.
Platform and vLLM documentation
ci/platform-matrix.json, docs/get-started/quickstart.mdx, docs/inference/*.mdx, docs/reference/*.mdx, test/inference-options-docs.test.ts
Documented defaults, image and revision pinning, download details, storage checks, authentication behavior, and deferred physical-hardware validation.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant User
  participant ExpressInstaller
  participant installVllm
  participant DockerStorage
  participant Docker
  participant HuggingFace
  User->>ExpressInstaller: accept DGX Station express install
  ExpressInstaller->>installVllm: select Nemotron Ultra recipe
  installVllm->>DockerStorage: check image storage
  installVllm->>Docker: pull pinned runtime image
  installVllm->>HuggingFace: download pinned model revision
  installVllm->>Docker: start configured vLLM container
  Docker-->>installVllm: report readiness
Loading

Possibly related issues

Possibly related PRs

Suggested labels: feature, provider: vllm, area: cli, area: onboarding

Suggested reviewers: sandl99, cv

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 27.27% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly matches the main change: adding a DGX Station express install flow in the installer.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/dgx-station-express-install

Comment @coderabbitai help to get the list of available commands.

Signed-off-by: Aaron Erickson <aerickson@nvidia.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
docs/get-started/quickstart.mdx (1)

128-128: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Use consistent second-person language across the new MDX guidance.

The changed documentation is internally consistent, but much of it describes product behavior instead of directly instructing the reader. Rewrite each affected passage as active, present-tense guidance.

  • docs/get-started/quickstart.mdx#L128-L128: Use “Accept the express prompt…” wording.
  • docs/get-started/quickstart.mdx#L195-L206: Rewrite platform defaults and warning text as direct instructions.
  • docs/inference/choose-inference-provider.mdx#L35-L35: Rewrite provider notes with “Use…” and “Ensure…” phrasing.
  • docs/inference/set-up-vllm.mdx#L82-L93: Rewrite image and duration guidance directly to the reader.
  • docs/inference/set-up-vllm.mdx#L134-L150: Rewrite profile, express-flow, and warning guidance directly to the reader.
  • docs/inference/set-up-vllm.mdx#L170-L170: Use conditional second-person wording for provider-only runs.
  • docs/inference/set-up-vllm.mdx#L182-L183: Make the model-table notes direct user guidance.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/get-started/quickstart.mdx` at line 128, Rewrite the affected MDX
guidance in active, present-tense second-person language: in
docs/get-started/quickstart.mdx lines 128 and 195-206, use direct “Accept…”
instructions and rewrite platform defaults and warnings; in
docs/inference/choose-inference-provider.mdx line 35, use “Use…” and “Ensure…”
phrasing; in docs/inference/set-up-vllm.mdx lines 82-93, 134-150, 170, and
182-183, directly address the reader for image, duration, profile, express-flow,
warning, provider-only, and model-table guidance. Preserve the documented
behavior and meaning while updating every listed passage.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@scripts/install.sh`:
- Around line 2686-2691: Update the inference_disclosure assignment alongside
the NEMOCLAW_VLLM_MODEL branch so overrides use generic configured-image
wording, while the default NVIDIA Nemotron 3 Ultra 550B path retains the pinned
Station image and approximately 352 GB disclosure.
- Around line 2793-2795: Update the NEMOCLAW_VLLM_MODEL initialization in the
install environment exports to treat whitespace-only values as unset or reject
them before export; preserve explicitly non-empty model values and ensure the
default nemotron-3-ultra-550b-a55b model is selected rather than allowing
downstream trimming to silently choose another profile.

---

Nitpick comments:
In `@docs/get-started/quickstart.mdx`:
- Line 128: Rewrite the affected MDX guidance in active, present-tense
second-person language: in docs/get-started/quickstart.mdx lines 128 and
195-206, use direct “Accept…” instructions and rewrite platform defaults and
warnings; in docs/inference/choose-inference-provider.mdx line 35, use “Use…”
and “Ensure…” phrasing; in docs/inference/set-up-vllm.mdx lines 82-93, 134-150,
170, and 182-183, directly address the reader for image, duration, profile,
express-flow, warning, provider-only, and model-table guidance. Preserve the
documented behavior and meaning while updating every listed passage.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: dee7cfda-6982-4ac8-ae0f-8e9f046c1d24

📥 Commits

Reviewing files that changed from the base of the PR and between 3461d71 and 64054d8.

📒 Files selected for processing (13)
  • ci/platform-matrix.json
  • docs/get-started/quickstart.mdx
  • docs/inference/choose-inference-provider.mdx
  • docs/inference/set-up-vllm.mdx
  • docs/reference/commands.mdx
  • docs/reference/platform-support.mdx
  • scripts/install.sh
  • src/lib/inference/vllm-models.test.ts
  • src/lib/inference/vllm-models.ts
  • src/lib/inference/vllm.test.ts
  • src/lib/inference/vllm.ts
  • test/inference-options-docs.test.ts
  • test/install-express-prompt.test.ts

Comment thread scripts/install.sh Outdated
Comment thread scripts/install.sh Outdated
sandl99 and others added 6 commits July 14, 2026 19:27
<!-- markdownlint-disable MD041 -->
## Summary

DGX Station GB300 OEM systems now enter the DGX Station express-install
path. Managed vLLM storage preflight is also narrowed to the Docker
image pull: it blocks only for a verified shortage, recognizes both
default Linux socket spellings, and no longer aborts express onboarding
when capacity is inconclusive.

## Proof of test
### Happy case
```
  [3/8] Configuring inference provider
  ──────────────────────────────────────────────────
  [non-interactive] Provider: install-vllm

  vLLM (DGX Station):
    Image: nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2
    Model: deepseek-ai/DeepSeek-V4-Flash
    Image download on first run, cached after
    Model download on first run, cached after


  Installing vLLM. Progress will print below.
  ==> Pulling vLLM image: nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2
  ==> nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2: Pulling from nvidia/vllm
```
### Shortage of space
```
  Installing vLLM. Progress will print below.

  Insufficient Docker storage for the managed vLLM image.

  Image:     nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2
  Available: 9.7 GiB
  Required:  approximately 29.8 GiB
  Storage:   Docker root directory (/mnt/nemoclaw-docker-10g/docker)

  Free or expand Docker storage before continuing.
  Useful diagnostics:
    docker system df
    docker info --format '{{.DockerRootDir}}'
  Non-interactive setup stops before the guarded download. Set NEMOCLAW_IGNORE_VLLM_DISK_SPACE=1 to override.
  [non-interactive] Aborting: vLLM install failed. See errors above.
```
## Related Issue

Closes #6757.
Closes #6858.

## Changes

- Detect product names containing both `Station` and `GB300` as DGX
Station for express install.
- Keep the managed image-size estimate and backend-aware
Docker/containerd capacity probe while removing Hugging Face model-cache
sizing and bind-identity probes.
- Prompt or stop only for a verified Docker image-storage shortage;
continue when capacity cannot be established, and retain the explicit
non-interactive override for known shortages.
- Recognize both `/run/docker.sock` and `/var/run/docker.sock`, and
honor Docker's documented `DOCKER_CONTEXT` precedence.
- Update focused installer/storage tests and user documentation.

## Type of Change

- [ ] Code change (feature, bug fix, or refactor)
- [x] Code change with doc updates
- [ ] Doc only (prose changes, no code sample modifications)
- [ ] Doc only (includes code sample changes)

## Quality Gates

- [x] Tests added or updated for changed behavior
- [ ] Existing tests cover changed behavior — justification:
- [ ] Tests not applicable — justification:
- [x] Docs updated for user-facing behavior changes
- [ ] Docs not applicable — justification:
- [x] Sensitive paths changed (security, policy, credentials, preflight,
onboarding, inference, runner, sandbox, or messaging)
- [x] Sensitive-path review completed or maintainer-approved waiver
recorded — reviewer/approval link/justification: Full combined-diff
review found no secret, dependency, injection, authentication,
cryptography, privilege, or sandbox-policy issues; Docker context
precedence is covered in both directions.
- [ ] Non-success, skipped, or missing CI check accepted by maintainer —
check name, approval link, and follow-up issue:

## Verification

- [x] PR description includes a `Signed-off-by:` line and every commit
appears as `Verified` in GitHub
- [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or
`npm run check:diff` passed when hooks were skipped or unavailable
- [x] Targeted behavior tests pass for the current change set, or tests
are marked not applicable above — installer integration: 7 passed, 1
skipped; focused CLI: 88 passed; focused integration: 38 passed; `npm
run typecheck:cli` passed.
- [ ] Applicable broad gate passed — `npm test` for broad
runtime/test-harness changes; `npm run check` for repo-wide
validation/coverage changes — command/result:
- [x] Quality Gates section completed with required justifications or
waivers
- [x] No secrets, API keys, or credentials committed
- [ ] `npm run docs` builds without warnings (doc changes only) — build
passed with two pre-existing Fern warnings and no errors.
- [x] Doc pages follow the [style
guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md)
(doc changes only)
- [ ] New doc pages include SPDX header and frontmatter (new pages only)

---

Signed-off-by: San Dang <sdang@nvidia.com>



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary by CodeRabbit

* **New Features**
* Added support for recognizing additional DGX Station hardware variants
during installation.
* Added image-focused managed vLLM disk preflight checks before pulling
vLLM images.

* **Bug Fixes**
* Improved Docker image-storage detection with clearer inconclusive
behavior across local configurations.
* Tightened non-interactive and `--yes` / disk-override handling so only
explicitly verified cases can proceed.

* **Documentation**
* Updated vLLM setup guidance and command reference to clarify that
Hugging Face model-cache space is not preflight-estimated.

* **Tests**
* Updated vLLM storage and capacity test coverage to match the new
probing scope.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: San Dang <sdang@nvidia.com>
Co-authored-by: Julie Yaunches <jyaunches@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
scripts/install.sh (1)

2677-2681: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

DGX Spark still lacks the whitespace-only normalization applied to DGX Station.

The Station branch now trims whitespace before deciding whether NEMOCLAW_VLLM_MODEL is "set" (Lines 2686, 2796), but the Spark branch still uses a plain [ -n "${NEMOCLAW_VLLM_MODEL:-}" ] check (Lines 2677, 2789). A whitespace-only override on Spark will still silently bypass the profile default, the same bug just fixed for Station.

♻️ Suggested fix to mirror the Station normalization
     "DGX Spark")
-      if [ -n "${NEMOCLAW_VLLM_MODEL:-}" ]; then
+      if [ -n "$(printf "%s" "${NEMOCLAW_VLLM_MODEL:-}" | tr -d '[:space:]')" ]; then
         inference_summary="managed local vLLM with model ${NEMOCLAW_VLLM_MODEL}"
         "DGX Spark")
           export NEMOCLAW_SANDBOX_NAME="${NEMOCLAW_SANDBOX_NAME:-my-assistant}"
           export NEMOCLAW_PROVIDER=install-vllm
-          if [ -n "${NEMOCLAW_VLLM_MODEL:-}" ]; then
+          if [ -n "$(printf "%s" "${NEMOCLAW_VLLM_MODEL:-}" | tr -d '[:space:]')" ]; then
             export NEMOCLAW_VLLM_MODEL
           fi

Also applies to: 2789-2792

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/install.sh` around lines 2677 - 2681, Update the DGX Spark
inference-summary branches around the relevant model checks to normalize
NEMOCLAW_VLLM_MODEL by trimming whitespace before testing whether it is set. Use
the same normalization already applied in the DGX Station branch, so
whitespace-only values select the DGX Spark profile default model while
non-blank values remain in the managed-model path.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@scripts/install.sh`:
- Around line 2677-2681: Update the DGX Spark inference-summary branches around
the relevant model checks to normalize NEMOCLAW_VLLM_MODEL by trimming
whitespace before testing whether it is set. Use the same normalization already
applied in the DGX Station branch, so whitespace-only values select the DGX
Spark profile default model while non-blank values remain in the managed-model
path.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 57f413bc-6a05-4270-8e91-67458c094891

📥 Commits

Reviewing files that changed from the base of the PR and between c2b1d7d and a3dc0ff.

📒 Files selected for processing (2)
  • scripts/install.sh
  • test/install-express-prompt.test.ts

ericksoa added 2 commits July 14, 2026 12:35
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/lib/inference/vllm-storage.ts (1)

272-288: 🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift

Restore a model-cache capacity preflight before the 352 GB download.

This probe now gates only the image store. A Station host can pass the roughly 35 GB image requirement and then fill the filesystem backing ~/.cache/huggingface during the approximately 352.38 GB model download. Reintroduce a cache-path check before downloadModel and cover insufficient-space behavior.

As per path instructions, “Trace every in-scope entrypoint and lifecycle path, including ... storage.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/lib/inference/vllm-storage.ts` around lines 272 - 288, The storage
preflight currently checks only Docker image locations; restore a capacity check
for the model cache before the 352 GB download. Update probeDockerStorage and
its callers, including the path leading to downloadModel, to resolve and
validate the Hugging Face cache path in addition to Docker storage, reject
insufficient cache capacity through the existing StorageProbeResult behavior,
and add coverage for the insufficient-space case.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lib/inference/vllm-storage.ts`:
- Around line 163-201: Update localDockerHostProblem and the surrounding
resolveDockerStorageLocations result model to distinguish remote Docker
endpoints and named non-local contexts as blocking outcomes, preserving
non-blocking behavior only for genuinely local filesystem inspection failures.
Trace every in-scope entrypoint and lifecycle path, including fresh execution,
so these outcomes stop before pull, download, and run. In
src/lib/inference/vllm-storage.ts lines 163-201, implement the distinct blocking
result; in src/lib/inference/vllm.test.ts lines 689-709, assert the
remote-endpoint case halts before those operations and replace its non-blocking
fixture with a genuinely local inspection failure.

---

Outside diff comments:
In `@src/lib/inference/vllm-storage.ts`:
- Around line 272-288: The storage preflight currently checks only Docker image
locations; restore a capacity check for the model cache before the 352 GB
download. Update probeDockerStorage and its callers, including the path leading
to downloadModel, to resolve and validate the Hugging Face cache path in
addition to Docker storage, reject insufficient cache capacity through the
existing StorageProbeResult behavior, and add coverage for the
insufficient-space case.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: db52cbfb-04c8-4bbe-8b00-83adcc84a2cc

📥 Commits

Reviewing files that changed from the base of the PR and between a3dc0ff and f028538.

📒 Files selected for processing (10)
  • docs/inference/set-up-vllm.mdx
  • docs/reference/commands.mdx
  • scripts/install.sh
  • src/lib/inference/vllm-models.test.ts
  • src/lib/inference/vllm-models.ts
  • src/lib/inference/vllm-storage.test.ts
  • src/lib/inference/vllm-storage.ts
  • src/lib/inference/vllm.test.ts
  • src/lib/inference/vllm.ts
  • test/install-express-prompt.test.ts
💤 Files with no reviewable changes (2)
  • src/lib/inference/vllm-models.test.ts
  • src/lib/inference/vllm-models.ts
🚧 Files skipped from review as they are similar to previous changes (3)
  • docs/reference/commands.mdx
  • scripts/install.sh
  • src/lib/inference/vllm.ts

Comment on lines +163 to 201
function localDockerHostProblem(info: DockerInfoShape, deps: StorageProbeDeps): string | null {
if (deps.platform !== "linux") return `Docker runs behind a ${deps.platform} host boundary`;
if (/microsoft|wsl/i.test(deps.osRelease)) return "Docker runs behind a WSL host boundary";
if (deps.clientContainerized) {
return "Docker client runs inside a container, so daemon bind-mount storage cannot be verified";
}
if (info.OSType !== "linux") return "Docker is not using a Linux engine";

const product = `${String(info.Name ?? "")} ${String(info.OperatingSystem ?? "")}`;
if (/docker desktop|colima|podman/i.test(product)) {
return "Docker runs inside a VM or compatibility layer";
}

const dockerHost = deps.dockerHost?.trim() ?? "";
// DOCKER_HOST takes precedence over DOCKER_CONTEXT in the Docker CLI.
const explicitContext = dockerHost ? "" : deps.dockerContext?.trim();
const reportedContext =
typeof info.ClientInfo?.Context === "string" ? info.ClientInfo.Context.trim() : "";
const context = explicitContext || reportedContext;
if (!context) return "docker info did not report the effective Docker context";
if (context !== "default") {
return `Docker uses a named context (${context}) whose host filesystem cannot be verified`;
// An explicit DOCKER_CONTEXT overrides DOCKER_HOST in the Docker CLI.
const explicitContext = deps.dockerContext?.trim() ?? "";
if (explicitContext) {
if (explicitContext !== "default") {
return `Docker uses a named context (${explicitContext}) whose host filesystem cannot be inspected`;
}
return null;
}

const endpoint = dockerHost;
if (endpoint) {
const localSocket = endpoint.startsWith("unix://") || path.isAbsolute(endpoint);
if (!localSocket) return `Docker uses a remote endpoint (${endpoint})`;
if (endpoint !== DEFAULT_DOCKER_SOCKET && endpoint !== `unix://${DEFAULT_DOCKER_SOCKET}`) {
return `Docker uses a non-default socket (${endpoint}) whose daemon host filesystem cannot be verified`;
if (dockerHost) {
if (isDefaultDockerSocket(dockerHost)) return null;
if (dockerHost.startsWith("unix://") || path.isAbsolute(dockerHost)) {
return `Docker uses a non-default socket (${dockerHost}) whose host filesystem cannot be inspected`;
}
return `Docker uses a remote endpoint (${dockerHost})`;
}

const sharesPeerMountNamespace = deps.dockerSocketPeerSharesMountNamespace();
if (sharesPeerMountNamespace === false) {
return "Docker client and socket peer use different mount namespaces, so daemon filesystem identity cannot be verified";
}
if (sharesPeerMountNamespace === null) {
return "Docker socket peer PID or mount namespace could not be verified";
const reportedContext =
typeof info.ClientInfo?.Context === "string" ? info.ClientInfo.Context.trim() : "";
if (reportedContext && reportedContext !== "default") {
return `Docker uses a named context (${reportedContext}) whose host filesystem cannot be inspected`;
}
return null;
}

export function probeDockerHostLocality(
overrides: Partial<StorageProbeDeps> = {},
): DockerHostLocalityResult {
const deps = { ...defaultStorageProbeDeps(), ...overrides };
const info = parseDockerInfo(deps.dockerInfo());
if (!info) return { ok: false, reason: "docker info did not return valid JSON" };
const hostProblem = nativeDockerHostProblem(info, deps);
return hostProblem ? { ok: false, reason: hostProblem } : { ok: true };
}

export function resolveDockerStorageLocations(
rawInfo: string,
overrides: Partial<StorageProbeDeps> = {},
): { ok: true; locations: DockerStorageLocation[] } | { ok: false; reason: string } {
const deps = { ...defaultStorageProbeDeps(), ...overrides };
const info = parseDockerInfo(rawInfo);
if (!info) return { ok: false, reason: "docker info did not return valid JSON" };
const hostProblem = nativeDockerHostProblem(info, deps);
const hostProblem = localDockerHostProblem(info, deps);
if (hostProblem) return { ok: false, reason: hostProblem };

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift

Keep remote Docker hosts out of the non-blocking capacity path.

The new result model conflates an unmeasurable local filesystem with an unsupported remote daemon, causing installation to continue across the host boundary.

  • src/lib/inference/vllm-storage.ts#L163-L201: return a distinct blocking outcome for remote endpoints and named non-local contexts.
  • src/lib/inference/vllm.test.ts#L689-L709: assert that the remote-endpoint case stops before pull, download, and run; use a genuinely local inspection failure for non-blocking coverage.

As per path instructions, “Trace every in-scope entrypoint and lifecycle path, including fresh execution.”

📍 Affects 2 files
  • src/lib/inference/vllm-storage.ts#L163-L201 (this comment)
  • src/lib/inference/vllm.test.ts#L689-L709
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/lib/inference/vllm-storage.ts` around lines 163 - 201, Update
localDockerHostProblem and the surrounding resolveDockerStorageLocations result
model to distinguish remote Docker endpoints and named non-local contexts as
blocking outcomes, preserving non-blocking behavior only for genuinely local
filesystem inspection failures. Trace every in-scope entrypoint and lifecycle
path, including fresh execution, so these outcomes stop before pull, download,
and run. In src/lib/inference/vllm-storage.ts lines 163-201, implement the
distinct blocking result; in src/lib/inference/vllm.test.ts lines 689-709,
assert the remote-endpoint case halts before those operations and replace its
non-blocking fixture with a genuinely local inspection failure.

Source: Path instructions

@ericksoa

Copy link
Copy Markdown
Contributor Author

Superseded by #6883. After #6875 merged, the no-force-push branch policy required a merge bridge that made GitHub attribute already-merged storage changes to this PR. #6883 carries the same validated Station express implementation on a clean current-main history.

@ericksoa ericksoa closed this Jul 14, 2026
ericksoa added a commit that referenced this pull request Jul 15, 2026
<!-- markdownlint-disable MD041 -->
## Summary

DGX Station now uses the existing express-install and onboarding FSM to
offer a one-confirmation managed-vLLM install. The express default is
the canonical pinned NVIDIA Nemotron 3 Ultra 550B recipe;
`--station-deepseek` selects the existing DeepSeek V4 Flash recipe for
demos.

This PR does not add a parallel launcher, a local-machine image
dependency, or a new network mode. The previously proposed
`experimental-single-user` profile has been removed because its
qualified Docker config ID was not published as a registry manifest.

Supersedes #6881 with a clean history after #6875 merged; repository
policy disables force-pushing the original PR branch.

## Changes

- Detect DGX Station in the existing installer path and offer express
setup with no follow-up model/configuration choices after confirmation.
- Select Nemotron 3 Ultra by default for Station express setup while
preserving `--station-deepseek` as the explicit DeepSeek V4 Flash
override.
- Keep managed vLLM on NemoClaw's existing Docker bridge topology:
`--ipc=host`, explicit `-p 8000:8000`, no `--network host` override.
- Pin Ultra to:
  - model `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4`
  - revision `183968f87ae4cedce3039313cac1fd43d112c578`
  - served identity `nvidia/nemotron-3-ultra-550b-a55b`
  - context length `262144`
- runtime
`vllm/vllm-openai@sha256:0fec7ec5f3e6bc168e54899935fb0557da908a4832a1dbc88e2debcf2f889416`
- 150 GiB CPU offload, 16 GiB shared memory, memlock/stack ulimits, MTP,
`nemotron_v3`, and `qwen3_coder`
- Add the approximately 352 GB Hugging Face cache preflight,
post-image-pull capacity recheck, a 3600-second Ultra startup timeout,
managed-container ownership protection, and expected-versus-detected
model handling when port 8000 is occupied. Inconclusive model-cache
probes now require explicit interactive confirmation and fail closed in
non-interactive setup unless the exact disk-space override is set.
- Normalize the canonical Ultra served alias back to the registered
installer slug before managed-vLLM selection. Validate explicit
Station-only and conflicting flags before license state, Docker setup,
OpenShell build dependencies, or any other host mutation.
- Require every effective managed-vLLM runtime to use a pullable
immutable `repository@sha256:<manifest>` reference. Bare Docker
image/config IDs and mutable tags fail before callbacks, prompts, pulls,
or container launch.
- Keep explicit pulls against the immutable digest even on cache hits;
download and long-lived containers use `--pull=never` afterward so
Docker cannot substitute another image.
- Update Station express, managed-vLLM, storage, security, and
Deferred-validation documentation.

## Distribution and Network Boundary

All four shipped managed-vLLM refs were resolved directly from their
registries without pulling layers. Each returned HTTP 200 and a
`Docker-Content-Digest` equal to the requested digest:

-
`vllm/vllm-openai@sha256:0fec7ec5f3e6bc168e54899935fb0557da908a4832a1dbc88e2debcf2f889416`
— multi-arch index containing Linux ARM64 and AMD64.
-
`nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2`
— Linux ARM64.
-
`nvcr.io/nvidia/vllm@sha256:447995cbb57e6c7cf792cab95e9852e5f62b5fb6d2f39e030fa4eda9a54eadb4`
— Linux ARM64.
-
`nvcr.io/nvidia/vllm@sha256:7be6c2f676c36059a494fe17254e69ae5c677535ba6191044e5fc8e42a91c773`
— Linux AMD64.

The Station Ultra runtime follows the same network boundary as standard
managed vLLM. `0.0.0.0` is inside the container network namespace and
Docker publishes only port 8000. Because Docker's default publication
can bind on all host interfaces, the existing default-deny firewall
guidance still applies; this PR introduces no additional host-network
exception.

## Runtime Selection

| Station path | Selection | Runtime image | Network |
|---|---|---|---|
| Express default | Nemotron 3 Ultra 550B | published immutable Docker
Hub digest above | existing bridge + `-p 8000:8000` |
| `--station-deepseek` | DeepSeek V4 Flash | published immutable NGC
digest above | existing bridge + `-p 8000:8000` |
| Interactive managed vLLM | existing Station registry/default behavior
| published immutable registry digest | existing bridge + `-p 8000:8000`
|

There is no installer-selectable experimental/local-only profile in this
PR. A future qualified single-user recipe can be proposed only after its
exact runtime is published as a pullable immutable manifest and
integrated through this same registry/FSM path.

## Type of Change

- [ ] Code change (feature, bug fix, or refactor)
- [x] Code change with doc updates
- [ ] Doc only (prose changes, no code sample modifications)
- [ ] Doc only (includes code sample changes)

## Quality Gates

- [x] Tests added or updated for changed behavior
- [x] Docs updated for user-facing behavior changes
- [x] Sensitive paths changed (preflight, onboarding, inference, and
container launch)
- [x] Product/design scope is being coordinated directly with the PM
team; code/security review remains requested on the exact head.
- [ ] Non-success, skipped, or missing CI check accepted by maintainer —
no waiver requested.

## Verification

Exact local head: `7624d02c7da6d96bb49058bd49474941740e9bd1`.

- [x] Commit and pre-push hooks passed, including repository checks,
formatting, lint, ShellCheck, secret scan, installer env-var
documentation, and CLI typecheck.
- [x] Cumulative focused changed-surface verification: 342 passed, 1
existing skip across installer, vLLM registry/runtime/storage,
onboarding FSM, Hermes config/dashboard, CLI dispatch, and docs-contract
suites; the latest storage/ordering remediation subset is 144 passed, 1
existing skip.
- [x] `npm run typecheck` passed.
- [x] `npm run docs:strict` passed with zero errors and two existing
Fern warnings.
- [x] `npm run check:installer-hash`, `bash -n install.sh
scripts/install.sh`, and `shellcheck install.sh scripts/install.sh`
passed.
- [x] Synthetic merge-tree comparison against the pre-experiment
boundary plus current merged-main state found only the intended alias
normalization, canonical command assertions, occupied-port assertion,
and registry-digest enforcement; canonical Ultra and
`--station-deepseek` behavior are unchanged by the cleanup.
- [x] All shipped managed-vLLM image manifests resolve remotely at their
exact pinned digests.
- [ ] Fresh physical DGX Station qualification remains tracked by the
existing Deferred platform status; this PR does not claim to advance
that status.
- [x] No secrets, API keys, credentials, local image IDs, or
host-network runtime overrides are committed.

---
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>

---------

Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: docs Documentation, examples, guides, or docs build area: inference Inference routing, serving, model selection, or outputs platform: dgx-station Affects DGX Station hardware or workflows

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants