Conversation
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall coverage remains at 96%, unchanged from the TypeScript / code-coverage/cliThe overall coverage in the Show a code coverage summary of the most impacted files.
Updated |
|
🌿 Preview your docs: https://nvidia-preview-pr-6881.docs.buildwithfern.com/nemoclaw |
PR Review Advisor — InformationalAdvisor assessment: Informational / medium confidence Model lanes
Nemotron output stays in workflow artifacts and does not change the assessment above. E2E guidanceAdvisory only. E2E / PR Gate selects and runs jobs independently. Recommended E2E: 2 optional E2E recommendations
2 warnings · 0 suggestionsWarningsWarnings do not block.
|
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
📝 WalkthroughWalkthroughDGX Station express installation now selects a pinned Nemotron Ultra managed-vLLM recipe. Runtime handling, Docker image storage checks, installer detection, onboarding tests, and platform documentation were updated to support the new model and deferred hardware-validation status. ChangesDGX Station Nemotron Ultra managed vLLM
Estimated code review effort: 4 (Complex) | ~45 minutes Sequence Diagram(s)sequenceDiagram
participant User
participant ExpressInstaller
participant installVllm
participant DockerStorage
participant Docker
participant HuggingFace
User->>ExpressInstaller: accept DGX Station express install
ExpressInstaller->>installVllm: select Nemotron Ultra recipe
installVllm->>DockerStorage: check image storage
installVllm->>Docker: pull pinned runtime image
installVllm->>HuggingFace: download pinned model revision
installVllm->>Docker: start configured vLLM container
Docker-->>installVllm: report readiness
Possibly related issues
Possibly related PRs
Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
There was a problem hiding this comment.
Actionable comments posted: 2
🧹 Nitpick comments (1)
docs/get-started/quickstart.mdx (1)
128-128: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winUse consistent second-person language across the new MDX guidance.
The changed documentation is internally consistent, but much of it describes product behavior instead of directly instructing the reader. Rewrite each affected passage as active, present-tense guidance.
docs/get-started/quickstart.mdx#L128-L128: Use “Accept the express prompt…” wording.docs/get-started/quickstart.mdx#L195-L206: Rewrite platform defaults and warning text as direct instructions.docs/inference/choose-inference-provider.mdx#L35-L35: Rewrite provider notes with “Use…” and “Ensure…” phrasing.docs/inference/set-up-vllm.mdx#L82-L93: Rewrite image and duration guidance directly to the reader.docs/inference/set-up-vllm.mdx#L134-L150: Rewrite profile, express-flow, and warning guidance directly to the reader.docs/inference/set-up-vllm.mdx#L170-L170: Use conditional second-person wording for provider-only runs.docs/inference/set-up-vllm.mdx#L182-L183: Make the model-table notes direct user guidance.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@docs/get-started/quickstart.mdx` at line 128, Rewrite the affected MDX guidance in active, present-tense second-person language: in docs/get-started/quickstart.mdx lines 128 and 195-206, use direct “Accept…” instructions and rewrite platform defaults and warnings; in docs/inference/choose-inference-provider.mdx line 35, use “Use…” and “Ensure…” phrasing; in docs/inference/set-up-vllm.mdx lines 82-93, 134-150, 170, and 182-183, directly address the reader for image, duration, profile, express-flow, warning, provider-only, and model-table guidance. Preserve the documented behavior and meaning while updating every listed passage.Source: Coding guidelines
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@scripts/install.sh`:
- Around line 2686-2691: Update the inference_disclosure assignment alongside
the NEMOCLAW_VLLM_MODEL branch so overrides use generic configured-image
wording, while the default NVIDIA Nemotron 3 Ultra 550B path retains the pinned
Station image and approximately 352 GB disclosure.
- Around line 2793-2795: Update the NEMOCLAW_VLLM_MODEL initialization in the
install environment exports to treat whitespace-only values as unset or reject
them before export; preserve explicitly non-empty model values and ensure the
default nemotron-3-ultra-550b-a55b model is selected rather than allowing
downstream trimming to silently choose another profile.
---
Nitpick comments:
In `@docs/get-started/quickstart.mdx`:
- Line 128: Rewrite the affected MDX guidance in active, present-tense
second-person language: in docs/get-started/quickstart.mdx lines 128 and
195-206, use direct “Accept…” instructions and rewrite platform defaults and
warnings; in docs/inference/choose-inference-provider.mdx line 35, use “Use…”
and “Ensure…” phrasing; in docs/inference/set-up-vllm.mdx lines 82-93, 134-150,
170, and 182-183, directly address the reader for image, duration, profile,
express-flow, warning, provider-only, and model-table guidance. Preserve the
documented behavior and meaning while updating every listed passage.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: dee7cfda-6982-4ac8-ae0f-8e9f046c1d24
📒 Files selected for processing (13)
ci/platform-matrix.jsondocs/get-started/quickstart.mdxdocs/inference/choose-inference-provider.mdxdocs/inference/set-up-vllm.mdxdocs/reference/commands.mdxdocs/reference/platform-support.mdxscripts/install.shsrc/lib/inference/vllm-models.test.tssrc/lib/inference/vllm-models.tssrc/lib/inference/vllm.test.tssrc/lib/inference/vllm.tstest/inference-options-docs.test.tstest/install-express-prompt.test.ts
<!-- markdownlint-disable MD041 -->
## Summary
DGX Station GB300 OEM systems now enter the DGX Station express-install
path. Managed vLLM storage preflight is also narrowed to the Docker
image pull: it blocks only for a verified shortage, recognizes both
default Linux socket spellings, and no longer aborts express onboarding
when capacity is inconclusive.
## Proof of test
### Happy case
```
[3/8] Configuring inference provider
──────────────────────────────────────────────────
[non-interactive] Provider: install-vllm
vLLM (DGX Station):
Image: nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2
Model: deepseek-ai/DeepSeek-V4-Flash
Image download on first run, cached after
Model download on first run, cached after
Installing vLLM. Progress will print below.
==> Pulling vLLM image: nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2
==> nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2: Pulling from nvidia/vllm
```
### Shortage of space
```
Installing vLLM. Progress will print below.
Insufficient Docker storage for the managed vLLM image.
Image: nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2
Available: 9.7 GiB
Required: approximately 29.8 GiB
Storage: Docker root directory (/mnt/nemoclaw-docker-10g/docker)
Free or expand Docker storage before continuing.
Useful diagnostics:
docker system df
docker info --format '{{.DockerRootDir}}'
Non-interactive setup stops before the guarded download. Set NEMOCLAW_IGNORE_VLLM_DISK_SPACE=1 to override.
[non-interactive] Aborting: vLLM install failed. See errors above.
```
## Related Issue
Closes #6757.
Closes #6858.
## Changes
- Detect product names containing both `Station` and `GB300` as DGX
Station for express install.
- Keep the managed image-size estimate and backend-aware
Docker/containerd capacity probe while removing Hugging Face model-cache
sizing and bind-identity probes.
- Prompt or stop only for a verified Docker image-storage shortage;
continue when capacity cannot be established, and retain the explicit
non-interactive override for known shortages.
- Recognize both `/run/docker.sock` and `/var/run/docker.sock`, and
honor Docker's documented `DOCKER_CONTEXT` precedence.
- Update focused installer/storage tests and user documentation.
## Type of Change
- [ ] Code change (feature, bug fix, or refactor)
- [x] Code change with doc updates
- [ ] Doc only (prose changes, no code sample modifications)
- [ ] Doc only (includes code sample changes)
## Quality Gates
- [x] Tests added or updated for changed behavior
- [ ] Existing tests cover changed behavior — justification:
- [ ] Tests not applicable — justification:
- [x] Docs updated for user-facing behavior changes
- [ ] Docs not applicable — justification:
- [x] Sensitive paths changed (security, policy, credentials, preflight,
onboarding, inference, runner, sandbox, or messaging)
- [x] Sensitive-path review completed or maintainer-approved waiver
recorded — reviewer/approval link/justification: Full combined-diff
review found no secret, dependency, injection, authentication,
cryptography, privilege, or sandbox-policy issues; Docker context
precedence is covered in both directions.
- [ ] Non-success, skipped, or missing CI check accepted by maintainer —
check name, approval link, and follow-up issue:
## Verification
- [x] PR description includes a `Signed-off-by:` line and every commit
appears as `Verified` in GitHub
- [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or
`npm run check:diff` passed when hooks were skipped or unavailable
- [x] Targeted behavior tests pass for the current change set, or tests
are marked not applicable above — installer integration: 7 passed, 1
skipped; focused CLI: 88 passed; focused integration: 38 passed; `npm
run typecheck:cli` passed.
- [ ] Applicable broad gate passed — `npm test` for broad
runtime/test-harness changes; `npm run check` for repo-wide
validation/coverage changes — command/result:
- [x] Quality Gates section completed with required justifications or
waivers
- [x] No secrets, API keys, or credentials committed
- [ ] `npm run docs` builds without warnings (doc changes only) — build
passed with two pre-existing Fern warnings and no errors.
- [x] Doc pages follow the [style
guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md)
(doc changes only)
- [ ] New doc pages include SPDX header and frontmatter (new pages only)
---
Signed-off-by: San Dang <sdang@nvidia.com>
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Summary by CodeRabbit
* **New Features**
* Added support for recognizing additional DGX Station hardware variants
during installation.
* Added image-focused managed vLLM disk preflight checks before pulling
vLLM images.
* **Bug Fixes**
* Improved Docker image-storage detection with clearer inconclusive
behavior across local configurations.
* Tightened non-interactive and `--yes` / disk-override handling so only
explicitly verified cases can proceed.
* **Documentation**
* Updated vLLM setup guidance and command reference to clarify that
Hugging Face model-cache space is not preflight-estimated.
* **Tests**
* Updated vLLM storage and capacity test coverage to match the new
probing scope.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: San Dang <sdang@nvidia.com>
Co-authored-by: Julie Yaunches <jyaunches@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
There was a problem hiding this comment.
🧹 Nitpick comments (1)
scripts/install.sh (1)
2677-2681: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick winDGX Spark still lacks the whitespace-only normalization applied to DGX Station.
The Station branch now trims whitespace before deciding whether
NEMOCLAW_VLLM_MODELis "set" (Lines 2686, 2796), but the Spark branch still uses a plain[ -n "${NEMOCLAW_VLLM_MODEL:-}" ]check (Lines 2677, 2789). A whitespace-only override on Spark will still silently bypass the profile default, the same bug just fixed for Station.♻️ Suggested fix to mirror the Station normalization
"DGX Spark") - if [ -n "${NEMOCLAW_VLLM_MODEL:-}" ]; then + if [ -n "$(printf "%s" "${NEMOCLAW_VLLM_MODEL:-}" | tr -d '[:space:]')" ]; then inference_summary="managed local vLLM with model ${NEMOCLAW_VLLM_MODEL}""DGX Spark") export NEMOCLAW_SANDBOX_NAME="${NEMOCLAW_SANDBOX_NAME:-my-assistant}" export NEMOCLAW_PROVIDER=install-vllm - if [ -n "${NEMOCLAW_VLLM_MODEL:-}" ]; then + if [ -n "$(printf "%s" "${NEMOCLAW_VLLM_MODEL:-}" | tr -d '[:space:]')" ]; then export NEMOCLAW_VLLM_MODEL fiAlso applies to: 2789-2792
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@scripts/install.sh` around lines 2677 - 2681, Update the DGX Spark inference-summary branches around the relevant model checks to normalize NEMOCLAW_VLLM_MODEL by trimming whitespace before testing whether it is set. Use the same normalization already applied in the DGX Station branch, so whitespace-only values select the DGX Spark profile default model while non-blank values remain in the managed-model path.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@scripts/install.sh`:
- Around line 2677-2681: Update the DGX Spark inference-summary branches around
the relevant model checks to normalize NEMOCLAW_VLLM_MODEL by trimming
whitespace before testing whether it is set. Use the same normalization already
applied in the DGX Station branch, so whitespace-only values select the DGX
Spark profile default model while non-blank values remain in the managed-model
path.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 57f413bc-6a05-4270-8e91-67458c094891
📒 Files selected for processing (2)
scripts/install.shtest/install-express-prompt.test.ts
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/lib/inference/vllm-storage.ts (1)
272-288: 🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy liftRestore a model-cache capacity preflight before the 352 GB download.
This probe now gates only the image store. A Station host can pass the roughly 35 GB image requirement and then fill the filesystem backing
~/.cache/huggingfaceduring the approximately 352.38 GB model download. Reintroduce a cache-path check beforedownloadModeland cover insufficient-space behavior.As per path instructions, “Trace every in-scope entrypoint and lifecycle path, including ... storage.”
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/lib/inference/vllm-storage.ts` around lines 272 - 288, The storage preflight currently checks only Docker image locations; restore a capacity check for the model cache before the 352 GB download. Update probeDockerStorage and its callers, including the path leading to downloadModel, to resolve and validate the Hugging Face cache path in addition to Docker storage, reject insufficient cache capacity through the existing StorageProbeResult behavior, and add coverage for the insufficient-space case.Source: Path instructions
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/lib/inference/vllm-storage.ts`:
- Around line 163-201: Update localDockerHostProblem and the surrounding
resolveDockerStorageLocations result model to distinguish remote Docker
endpoints and named non-local contexts as blocking outcomes, preserving
non-blocking behavior only for genuinely local filesystem inspection failures.
Trace every in-scope entrypoint and lifecycle path, including fresh execution,
so these outcomes stop before pull, download, and run. In
src/lib/inference/vllm-storage.ts lines 163-201, implement the distinct blocking
result; in src/lib/inference/vllm.test.ts lines 689-709, assert the
remote-endpoint case halts before those operations and replace its non-blocking
fixture with a genuinely local inspection failure.
---
Outside diff comments:
In `@src/lib/inference/vllm-storage.ts`:
- Around line 272-288: The storage preflight currently checks only Docker image
locations; restore a capacity check for the model cache before the 352 GB
download. Update probeDockerStorage and its callers, including the path leading
to downloadModel, to resolve and validate the Hugging Face cache path in
addition to Docker storage, reject insufficient cache capacity through the
existing StorageProbeResult behavior, and add coverage for the
insufficient-space case.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: db52cbfb-04c8-4bbe-8b00-83adcc84a2cc
📒 Files selected for processing (10)
docs/inference/set-up-vllm.mdxdocs/reference/commands.mdxscripts/install.shsrc/lib/inference/vllm-models.test.tssrc/lib/inference/vllm-models.tssrc/lib/inference/vllm-storage.test.tssrc/lib/inference/vllm-storage.tssrc/lib/inference/vllm.test.tssrc/lib/inference/vllm.tstest/install-express-prompt.test.ts
💤 Files with no reviewable changes (2)
- src/lib/inference/vllm-models.test.ts
- src/lib/inference/vllm-models.ts
🚧 Files skipped from review as they are similar to previous changes (3)
- docs/reference/commands.mdx
- scripts/install.sh
- src/lib/inference/vllm.ts
| function localDockerHostProblem(info: DockerInfoShape, deps: StorageProbeDeps): string | null { | ||
| if (deps.platform !== "linux") return `Docker runs behind a ${deps.platform} host boundary`; | ||
| if (/microsoft|wsl/i.test(deps.osRelease)) return "Docker runs behind a WSL host boundary"; | ||
| if (deps.clientContainerized) { | ||
| return "Docker client runs inside a container, so daemon bind-mount storage cannot be verified"; | ||
| } | ||
| if (info.OSType !== "linux") return "Docker is not using a Linux engine"; | ||
|
|
||
| const product = `${String(info.Name ?? "")} ${String(info.OperatingSystem ?? "")}`; | ||
| if (/docker desktop|colima|podman/i.test(product)) { | ||
| return "Docker runs inside a VM or compatibility layer"; | ||
| } | ||
|
|
||
| const dockerHost = deps.dockerHost?.trim() ?? ""; | ||
| // DOCKER_HOST takes precedence over DOCKER_CONTEXT in the Docker CLI. | ||
| const explicitContext = dockerHost ? "" : deps.dockerContext?.trim(); | ||
| const reportedContext = | ||
| typeof info.ClientInfo?.Context === "string" ? info.ClientInfo.Context.trim() : ""; | ||
| const context = explicitContext || reportedContext; | ||
| if (!context) return "docker info did not report the effective Docker context"; | ||
| if (context !== "default") { | ||
| return `Docker uses a named context (${context}) whose host filesystem cannot be verified`; | ||
| // An explicit DOCKER_CONTEXT overrides DOCKER_HOST in the Docker CLI. | ||
| const explicitContext = deps.dockerContext?.trim() ?? ""; | ||
| if (explicitContext) { | ||
| if (explicitContext !== "default") { | ||
| return `Docker uses a named context (${explicitContext}) whose host filesystem cannot be inspected`; | ||
| } | ||
| return null; | ||
| } | ||
|
|
||
| const endpoint = dockerHost; | ||
| if (endpoint) { | ||
| const localSocket = endpoint.startsWith("unix://") || path.isAbsolute(endpoint); | ||
| if (!localSocket) return `Docker uses a remote endpoint (${endpoint})`; | ||
| if (endpoint !== DEFAULT_DOCKER_SOCKET && endpoint !== `unix://${DEFAULT_DOCKER_SOCKET}`) { | ||
| return `Docker uses a non-default socket (${endpoint}) whose daemon host filesystem cannot be verified`; | ||
| if (dockerHost) { | ||
| if (isDefaultDockerSocket(dockerHost)) return null; | ||
| if (dockerHost.startsWith("unix://") || path.isAbsolute(dockerHost)) { | ||
| return `Docker uses a non-default socket (${dockerHost}) whose host filesystem cannot be inspected`; | ||
| } | ||
| return `Docker uses a remote endpoint (${dockerHost})`; | ||
| } | ||
|
|
||
| const sharesPeerMountNamespace = deps.dockerSocketPeerSharesMountNamespace(); | ||
| if (sharesPeerMountNamespace === false) { | ||
| return "Docker client and socket peer use different mount namespaces, so daemon filesystem identity cannot be verified"; | ||
| } | ||
| if (sharesPeerMountNamespace === null) { | ||
| return "Docker socket peer PID or mount namespace could not be verified"; | ||
| const reportedContext = | ||
| typeof info.ClientInfo?.Context === "string" ? info.ClientInfo.Context.trim() : ""; | ||
| if (reportedContext && reportedContext !== "default") { | ||
| return `Docker uses a named context (${reportedContext}) whose host filesystem cannot be inspected`; | ||
| } | ||
| return null; | ||
| } | ||
|
|
||
| export function probeDockerHostLocality( | ||
| overrides: Partial<StorageProbeDeps> = {}, | ||
| ): DockerHostLocalityResult { | ||
| const deps = { ...defaultStorageProbeDeps(), ...overrides }; | ||
| const info = parseDockerInfo(deps.dockerInfo()); | ||
| if (!info) return { ok: false, reason: "docker info did not return valid JSON" }; | ||
| const hostProblem = nativeDockerHostProblem(info, deps); | ||
| return hostProblem ? { ok: false, reason: hostProblem } : { ok: true }; | ||
| } | ||
|
|
||
| export function resolveDockerStorageLocations( | ||
| rawInfo: string, | ||
| overrides: Partial<StorageProbeDeps> = {}, | ||
| ): { ok: true; locations: DockerStorageLocation[] } | { ok: false; reason: string } { | ||
| const deps = { ...defaultStorageProbeDeps(), ...overrides }; | ||
| const info = parseDockerInfo(rawInfo); | ||
| if (!info) return { ok: false, reason: "docker info did not return valid JSON" }; | ||
| const hostProblem = nativeDockerHostProblem(info, deps); | ||
| const hostProblem = localDockerHostProblem(info, deps); | ||
| if (hostProblem) return { ok: false, reason: hostProblem }; |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift
Keep remote Docker hosts out of the non-blocking capacity path.
The new result model conflates an unmeasurable local filesystem with an unsupported remote daemon, causing installation to continue across the host boundary.
src/lib/inference/vllm-storage.ts#L163-L201: return a distinct blocking outcome for remote endpoints and named non-local contexts.src/lib/inference/vllm.test.ts#L689-L709: assert that the remote-endpoint case stops before pull, download, and run; use a genuinely local inspection failure for non-blocking coverage.
As per path instructions, “Trace every in-scope entrypoint and lifecycle path, including fresh execution.”
📍 Affects 2 files
src/lib/inference/vllm-storage.ts#L163-L201(this comment)src/lib/inference/vllm.test.ts#L689-L709
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/lib/inference/vllm-storage.ts` around lines 163 - 201, Update
localDockerHostProblem and the surrounding resolveDockerStorageLocations result
model to distinguish remote Docker endpoints and named non-local contexts as
blocking outcomes, preserving non-blocking behavior only for genuinely local
filesystem inspection failures. Trace every in-scope entrypoint and lifecycle
path, including fresh execution, so these outcomes stop before pull, download,
and run. In src/lib/inference/vllm-storage.ts lines 163-201, implement the
distinct blocking result; in src/lib/inference/vllm.test.ts lines 689-709,
assert the remote-endpoint case halts before those operations and replace its
non-blocking fixture with a genuinely local inspection failure.
Source: Path instructions
<!-- markdownlint-disable MD041 --> ## Summary DGX Station now uses the existing express-install and onboarding FSM to offer a one-confirmation managed-vLLM install. The express default is the canonical pinned NVIDIA Nemotron 3 Ultra 550B recipe; `--station-deepseek` selects the existing DeepSeek V4 Flash recipe for demos. This PR does not add a parallel launcher, a local-machine image dependency, or a new network mode. The previously proposed `experimental-single-user` profile has been removed because its qualified Docker config ID was not published as a registry manifest. Supersedes #6881 with a clean history after #6875 merged; repository policy disables force-pushing the original PR branch. ## Changes - Detect DGX Station in the existing installer path and offer express setup with no follow-up model/configuration choices after confirmation. - Select Nemotron 3 Ultra by default for Station express setup while preserving `--station-deepseek` as the explicit DeepSeek V4 Flash override. - Keep managed vLLM on NemoClaw's existing Docker bridge topology: `--ipc=host`, explicit `-p 8000:8000`, no `--network host` override. - Pin Ultra to: - model `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` - revision `183968f87ae4cedce3039313cac1fd43d112c578` - served identity `nvidia/nemotron-3-ultra-550b-a55b` - context length `262144` - runtime `vllm/vllm-openai@sha256:0fec7ec5f3e6bc168e54899935fb0557da908a4832a1dbc88e2debcf2f889416` - 150 GiB CPU offload, 16 GiB shared memory, memlock/stack ulimits, MTP, `nemotron_v3`, and `qwen3_coder` - Add the approximately 352 GB Hugging Face cache preflight, post-image-pull capacity recheck, a 3600-second Ultra startup timeout, managed-container ownership protection, and expected-versus-detected model handling when port 8000 is occupied. Inconclusive model-cache probes now require explicit interactive confirmation and fail closed in non-interactive setup unless the exact disk-space override is set. - Normalize the canonical Ultra served alias back to the registered installer slug before managed-vLLM selection. Validate explicit Station-only and conflicting flags before license state, Docker setup, OpenShell build dependencies, or any other host mutation. - Require every effective managed-vLLM runtime to use a pullable immutable `repository@sha256:<manifest>` reference. Bare Docker image/config IDs and mutable tags fail before callbacks, prompts, pulls, or container launch. - Keep explicit pulls against the immutable digest even on cache hits; download and long-lived containers use `--pull=never` afterward so Docker cannot substitute another image. - Update Station express, managed-vLLM, storage, security, and Deferred-validation documentation. ## Distribution and Network Boundary All four shipped managed-vLLM refs were resolved directly from their registries without pulling layers. Each returned HTTP 200 and a `Docker-Content-Digest` equal to the requested digest: - `vllm/vllm-openai@sha256:0fec7ec5f3e6bc168e54899935fb0557da908a4832a1dbc88e2debcf2f889416` — multi-arch index containing Linux ARM64 and AMD64. - `nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2` — Linux ARM64. - `nvcr.io/nvidia/vllm@sha256:447995cbb57e6c7cf792cab95e9852e5f62b5fb6d2f39e030fa4eda9a54eadb4` — Linux ARM64. - `nvcr.io/nvidia/vllm@sha256:7be6c2f676c36059a494fe17254e69ae5c677535ba6191044e5fc8e42a91c773` — Linux AMD64. The Station Ultra runtime follows the same network boundary as standard managed vLLM. `0.0.0.0` is inside the container network namespace and Docker publishes only port 8000. Because Docker's default publication can bind on all host interfaces, the existing default-deny firewall guidance still applies; this PR introduces no additional host-network exception. ## Runtime Selection | Station path | Selection | Runtime image | Network | |---|---|---|---| | Express default | Nemotron 3 Ultra 550B | published immutable Docker Hub digest above | existing bridge + `-p 8000:8000` | | `--station-deepseek` | DeepSeek V4 Flash | published immutable NGC digest above | existing bridge + `-p 8000:8000` | | Interactive managed vLLM | existing Station registry/default behavior | published immutable registry digest | existing bridge + `-p 8000:8000` | There is no installer-selectable experimental/local-only profile in this PR. A future qualified single-user recipe can be proposed only after its exact runtime is published as a pullable immutable manifest and integrated through this same registry/FSM path. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [x] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Quality Gates - [x] Tests added or updated for changed behavior - [x] Docs updated for user-facing behavior changes - [x] Sensitive paths changed (preflight, onboarding, inference, and container launch) - [x] Product/design scope is being coordinated directly with the PM team; code/security review remains requested on the exact head. - [ ] Non-success, skipped, or missing CI check accepted by maintainer — no waiver requested. ## Verification Exact local head: `7624d02c7da6d96bb49058bd49474941740e9bd1`. - [x] Commit and pre-push hooks passed, including repository checks, formatting, lint, ShellCheck, secret scan, installer env-var documentation, and CLI typecheck. - [x] Cumulative focused changed-surface verification: 342 passed, 1 existing skip across installer, vLLM registry/runtime/storage, onboarding FSM, Hermes config/dashboard, CLI dispatch, and docs-contract suites; the latest storage/ordering remediation subset is 144 passed, 1 existing skip. - [x] `npm run typecheck` passed. - [x] `npm run docs:strict` passed with zero errors and two existing Fern warnings. - [x] `npm run check:installer-hash`, `bash -n install.sh scripts/install.sh`, and `shellcheck install.sh scripts/install.sh` passed. - [x] Synthetic merge-tree comparison against the pre-experiment boundary plus current merged-main state found only the intended alias normalization, canonical command assertions, occupied-port assertion, and registry-digest enforcement; canonical Ultra and `--station-deepseek` behavior are unchanged by the cleanup. - [x] All shipped managed-vLLM image manifests resolve remotely at their exact pinned digests. - [ ] Fresh physical DGX Station qualification remains tracked by the existing Deferred platform status; this PR does not claim to advance that status. - [x] No secrets, API keys, credentials, local image IDs, or host-network runtime overrides are committed. --- Signed-off-by: Aaron Erickson <aerickson@nvidia.com> --------- Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
Summary
DGX Station's initial installer now offers one express-install confirmation and then completes onboarding without additional provider, model, policy, or sandbox-name choices. The express path translates the validated Nemotron playbook into a pinned managed-vLLM recipe for NVIDIA Nemotron 3 Ultra 550B while keeping direct managed-vLLM onboarding defaults unchanged.
Changes
nemotron-3-ultra-550b-a55bautomatically when DGX Station express install is accepted.Type of Change
Quality Gates
Verification
Signed-off-by:line and every commit appears asVerifiedin GitHubpre-commit,commit-msg, andpre-pushhooks passed, ornpm run check:diffpassed when hooks were skipped or unavailablevitest run src/lib/onboard/setup-nim-flow.test.ts src/lib/onboard/machine/handlers/provider-inference.test.ts src/lib/onboard/machine/transition-traces.test.ts src/lib/inference/vllm-models.test.ts src/lib/inference/vllm.test.ts test/inference-options-docs.test.ts test/install-express-prompt.test.ts --testTimeout=15000(168 passed, 1 existing skip)npm testfor broad runtime/test-harness changes;npm run checkfor repo-wide validation/coverage changes — not applicable; this bounded installer/managed-vLLM change is covered by targeted tests and normal hooks.npm run docsbuilds without warnings (doc changes only) —npm run docs:strictpassed with zero errors and two pre-existing Fern warnings.Signed-off-by: Aaron Erickson aerickson@nvidia.com
Summary by CodeRabbit
New Features
Documentation
Bug Fixes
Tests