Skip to content

VDR 5 checklist: DGX Station & DGX Spark golden path #6743

Description

@sandl99

VDR 5 Checklist: DGX Station and DGX Spark

This checklist reconciles the DGX Station and DGX Spark findings from VDR 5 with Will Curran's July 8, 2026 progress update. The tested evidence includes older releases, while the progress response uses .67 as the current validation target. A related pull request or closed issue is evidence of progress, not proof that the original platform-specific finding is fixed.

Status as of July 13, 2026:

  • [x] The original VDR behavior is fixed, addressed, no longer applicable, or otherwise validated.
  • [~] Relevant improvements landed, but the original Station or Spark behavior still needs a hands-on .67 retest.
  • [ ] The original finding remains open, lacks an exact fix, or has not been proven fixed.

DGX Station

Remaining VDR findings

  • Piped installer and sandbox-name flow do not complete cleanly out of box

    • The Station Playbook's curl | bash path did not reliably complete setup.
    • Re-running with --name <sandbox> was unclear or failed because arguments were being passed through a piped shell installer.
    • The working NEMOCLAW_SANDBOX_NAME=test-sandbox environment-variable path was not documented.
    • Acceptance: one supported copy-paste command completes setup or prints one unambiguous continuation command, including the sandbox name.
  • Disk-space preflight before vLLM, image, and model downloads

    • #6757 — exact managed-vLLM storage-preflight follow-up and native sub-issue.
    • Status: Failed on DGX Station with Docker's containerd image store limited to 2 GiB.
    • The NGC image pull downloaded and verified multiple layers before failing with no space left on device; no capacity warning appeared before the pull.
    • Acceptance: preflight the actual Docker image store before an uncached pull and the model-cache filesystem before hf download; report available and required space plus remediation, and stop safely by default unless explicitly overridden.
    • Verify classic DockerRootDir and containerd image-store backends, plus the bounded-storage DGX Station reproduction.
    • Do not claim this fixed without backend-aware unit coverage and a clean bounded-storage Station retest.
  • Host GPU detected, but sandbox reports GPU disabled

    • #5772 — related, closed.
    • #6172 — related, closed.
    • Acceptance: a clean current Station install detects and enables the expected sandbox GPU path.
    • Status: Passed, NemoClaw detected GPU.
  • Managed local vLLM endpoint validation failed

    • #5744 — related, closed; generic Linux/A100 custom-local-vLLM proxy failure, not a Station-exact fix.
    • #6177 — related, closed.
    • #6314 — related, closed.
    • Acceptance: managed local vLLM configures, validates, and serves successfully out of box on a current DGX Station.
    • Status: Passed
  • Piped installer : Onboard --resume --refresh contains bugs, bad UX

    • #6500 — a stopped registered sandbox blocks strict pre-upgrade backup during an installer rerun.
    • #6562 — canceling onboarding after sandbox-name capture can leave resume without its route reservation.
    • Status: Failed on the DGX Station manual retest at v0.0.81-4-gdeed1aaa3.
    • The stopped-sandbox path fails before onboarding after the current CLI is linked, with OpenShell installation still deferred.
    • The cancel/resume path replays uncheckpointed choices, then fails with sandbox route reservation 'tm' disappeared while onboarding was in progress and recommends the same nemoclaw onboard --resume recovery.
    • --refresh is not a supported flag.
    • Acceptance: installer reruns must provide one clear recovery path; resume must recover the original sandbox name and route reservation before later prompts, or fail fast with one actionable --fresh path.
  • Failed preflight logs need stronger visual contrast

    • #6752 — exact VDR follow-up and native sub-issue.
    • Status: Failed on the DGX Station manual retest.
    • The port-8080 conflict heading and explanation render as plain white text instead of using warning emphasis.
    • With a silent listener holding port 8080, accepting Continue with onboarding? [y/N] can hang after ✓ openshell CLI: openshell 0.0.72 before the conflict remediation is printed.
    • Acceptance: use the shared warning/error presentation and fail promptly with actionable PID and alternate-port guidance; direct onboarding and the installer path must not hang.

Additional Station follow-up

  • DGX Station Playbook still pins v0.0.55
    • #6755 — exact documentation follow-up and native sub-issue.
    • Status: Confirmed on July 13, 2026. The live official playbook calls v0.0.55 the currently recommended stable release and explicitly sets NEMOCLAW_INSTALL_TAG=v0.0.55.
    • The hosted installer defaults to the maintained lkg channel.
    • #6064 Closed corrected the DGX Spark playbook, not this Station page.
    • SWQA bug 6407748 tracks the stale-command concern.
    • Acceptance: remove the concrete version override and stale stable claim, use the supported hosted-installer path, and validate it on a clean Station.

Progress cited in Will's update

Related Station and managed-vLLM work landed through #4888, #4867, #4676, #5215, and #5612. The guided express-install prompt is also intended to steer Station users toward managed local vLLM and suggested policy defaults. These are progress signals; the unchecked and partially checked findings above remain until Station validation succeeds.

DGX Spark

Remaining VDR findings

  • Optional Telegram integration can abort express onboarding

    • When Telegram was unreachable from the corporate network, setup aborted instead of skipping the optional integration.
    • Acceptance: skip the optional channel or clearly direct the user to continue interactively with nemoclaw onboard.
    • Retest on .67.
  • Telegram configuration state and health diagnostics are misleading feat(messaging): telegram channels-status health probe #6887

    • Telegram could be shown as configured even when express setup did not finish.
    • nemoclaw <name> status did not show channel-specific health, while failed / fetch failed output lacked actionable token or network guidance.
  • nemoclaw onboard --resume repeats completed prompts

    • #6932 — exact VDR follow-up and native sub-issue.
    • #6934 — implementation under review.
    • Status: reproduced after completing sandbox-name, Brave Search, and messaging setup, then interrupting at resource-profile selection; resume repeated those completed prompts and credentials.
    • #5961 — related, closed; it covers interrupted non-interactive onboarding, not the exact repeated-prompt behavior.
    • #6179 — related, closed.
    • Acceptance: resume preserves sandbox name, web-search selection, messaging selection and non-secret settings, resource profile, and all later completed choices. Validated web/messaging credentials are reused only through same-session exact OpenShell bindings; raw secrets are never persisted in session state.
  • Empty resource-profile input does not select the displayed default

    • Pressing Enter at Choose [6]: should select profile 6 instead of failing validation.
  • Brave Search prompt needs a usable skip/back path

    • A user without an API key could not leave the prompt with q or Escape and had to interrupt onboarding with Ctrl-C.
    • Coverage and related work exist, but the no-key path still needs a hands-on .67 check.
  • Shields can block chat or re-enable unexpectedly

    • #5922 — exact match, closed.
    • #6182 — related, closed.
    • Retest that shields-up allows session/history operation without failing to create agents/main/sessions/, and that shields do not turn back on during a chat without user action.
  • Working local Ollama endpoint can fail provider verification

    • #1060 — related, open.
    • Switching from an NVIDIA endpoint failed verification even though Ollama was reachable and inference worked with --no-verify.
  • llama.cpp / custom OpenAI-compatible endpoint remains unproven

    • Onboarding appeared to succeed, but connecting to the sandbox failed.
    • This was outside the VDR 5 rating scope, but remains an unresolved Spark compatibility finding.
  • Installed OpenClaw version is not proven current

    • VDR observed an older OpenClaw inside the sandbox than the then-current release. Validate the installed version and supported update path on .67.
  • Gateway lifecycle and repeat-onboard cleanup remain unclear

    • There was no graceful gateway stop command; users had to kill the process.
    • A second onboard could fail because lingering OpenShell gateway or SSH-proxy processes still held ports 8080 or 18789.
    • Tunnel, dashboard, and gateway recovery improved, but this exact cleanup UX was not explicitly proven fixed.
    • Acceptance: provide a clear stop/cleanup command and make repeat onboarding recover automatically or explain exactly which prior process/state to clear.
  • nemoclaw status unhealthy-probe wording is too opaque

    • Define which output is the "host-side delivery chain."
    • Explain that OpenShell-managed runtimes include NemoClaw sandboxes.
    • Replace bare #3975 with the repository and a link to NemoClaw Issue #3975.
  • OpenClaw TUI displays noisy HEARTBEAT_OK messages

    • These consume space in the chat view and should be suppressed or reduced
  • nemoclaw --help lacks default-sandbox guidance

    • Expose or document how users select the default sandbox
  • Unknown-provider errors lack registration guidance

    • nemoclaw inference set is expected to reject unregistered providers, but the error should tell users to register them through nemoclaw onboard.
  • Remote dashboard access lacks a copy-paste SSH forwarding command

    • Onboard output only showed the localhost URL.
    • Quickstart did not cross-reference a concrete SSH port-forwarding command with placeholders for the remote host and user.

Progress cited in Will's update

Fixed, addressed, no longer applicable, or descoped

  • Pip installation without a virtual environment has documented workarounds.
  • Warm local latency is about 7–8 seconds and is comparable to the local Ollama baseline; the earlier 16-second concern improved.
  • nemoclaw status now provides context for starting stopped cloudflared.
  • Dead dashboard-forward recovery was added.
  • Spark-specific setup guide/navigation findings are no longer applicable; the remaining onboarding step was added to the installation documentation.
  • WhatsApp integration was descoped from VDR 5.
  • Global npm installation no longer requires the reported sudo workaround.
  • nemoclaw debug --quick no longer depends on failing ps and free commands inside the sandbox.
  • Stale onboard-session output from nemoclaw debug was fixed.
  • Onboard now selects another dashboard port when Remote SSH already holds 18789.
  • Model/provider update guidance was added to inference documentation.
  • nemoclaw connect installation/linking on Spark was fixed.
  • Spark GPU detection was fixed.
  • Long-output command-line guidance was added to the documentation.
  • Sandbox-name validation now makes the no-spaces requirement clear.
  • Documentation was corrected from openclaw nemoclaw status to the supported nemoclaw status path.

Validation rule

Do not close a [~] or [ ] item solely because a related issue or pull request closed. Close it only when the original behavior is either directly covered by code evidence or reproduced successfully on the target platform and release.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

VDRLinked to VDR findingplatform: dgx-sparkAffects DGX Spark hardware or workflowsplatform: dgx-stationAffects DGX Station hardware or workflows

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions