Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 12 additions & 2 deletions docs/inference/set-up-vllm.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -372,7 +372,17 @@ When you start managed vLLM outside the installer express flow, NemoClaw uses th
| Linux with an NVIDIA GPU | `nvidia/NVIDIA-Nemotron-3-Nano-4B-FP8` |

For the managed one-host DGX Spark and N1x `nvidia/Qwen3.6-35B-A3B-NVFP4` profiles, NemoClaw enables async scheduling and does not enable multi-token prediction (MTP) speculative decoding.
The DGX Spark profile retains `--gpu-memory-utilization 0.4`, a `262144`-token context window, `4` concurrent sequences, and an `8192`-token batch limit.
DGX Spark automatically selects its Qwen recipe from the detected unified-memory capacity:

| Detected capacity | Context tokens | Concurrent sequences | Batched tokens | GPU-memory utilization |
|---|---|---|---|---|
| At least 60 GB and less than 64 GB | 32,768 | 1 | 4,096 | 0.5 |
| At least 64 GB | 262,144 | 4 | 8,192 | 0.4 |

The smaller recipe admits a nominal 64 GB DGX Spark that reports about 61.6 GB of usable unified memory.
It remains Experimental with software validation; physical 64 GB Spark qualification is pending.
Unknown capacity or less than 60 GB does not qualify for either Qwen recipe.
Before installation, NemoClaw displays the selected hardware profile and context limit.
The Deferred N1x profile uses `--gpu-memory-utilization 0.6`, a `32768`-token context window, `2` concurrent sequences, and a `4096`-token batch limit.
The two-sequence limit permits vLLM to schedule a compaction summary while an agent response uses the other sequence.
Both profiles retain chunked prefill, prefix caching, and their registered parsers and acceleration backends.
Expand Down Expand Up @@ -720,7 +730,7 @@ Follow the diagnostic to restore valid telemetry, correct the selected GPU, or f
Then run `$$nemoclaw onboard --resume`.

To bound resource use while investigating long-context workflows on one DGX Spark, select the Qwen profile.
The following override disables async scheduling and lowers the context window, concurrent-sequence limit, and batch limit to the current N1x defaults.
For a Spark with at least 64 GB of detected unified memory, the following override disables async scheduling and lowers the context window, concurrent-sequence limit, and batch limit to the current N1x defaults.

```bash
NEMOCLAW_PROVIDER=install-vllm \
Expand Down
2 changes: 1 addition & 1 deletion docs/reference/commands.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -1358,7 +1358,7 @@ Nonempty `--message-file` arguments and calls without a recognized inline messag
The sandbox resolves message-file paths, including symlinks to stdin.
For example, the wrapper preserves redirected input for `printf 'ping' | $$nemoclaw my-assistant agent --agent main`.

When the top-level OpenClaw `--json` output flag is present, the wrapper uses a captured no-TTY path with a `64 MiB` buffer so `stdout` stays parseable JSON. Raw `stderr`, including structured JSON diagnostics, is forwarded unchanged. NemoClaw appends failed-tool or untrusted-child provenance only from the `stdout` JSON. The wrapper reads completion markers only from the final matching OpenClaw response envelope: a local `{ payloads, meta }` response or a gateway `{ status, result: { payloads, meta } }` response. It ignores earlier JSON progress or log records. It exits with status `1` when that metadata contains `error.kind: "incomplete_turn"`, `livenessState: "abandoned"`, `replayInvalid: true`, or a `timeoutPhase` value, even when the envelope reports success. Marker-shaped values inside tool results, tool-call arguments, or other descendants do not change the exit status. A turn can run every tool successfully and still become abandoned before it produces a reply. The wrapper writes the unchanged JSON trace to `stdout` before it reports the incomplete turn, so the partial tool trace remains available. The wrapper writes the verdict, the detected markers, and verify-before-retry guidance to `stderr`. A `timeoutPhase` value names the phase the deadline fired in, so the wrapper writes deadline guidance in place of the generic incomplete-turn text. Tool calls in a partial trace may have already applied side effects, so verify what the turn changed before you retry it. The wrapper passes through an upstream non-zero exit status unchanged. Literal `--json` values consumed by flags such as `-m` or `--reply-channel`, or arguments after `--`, stay on the normal passthrough path. Documented value flags written as `--flag=value`, such as `--session-id=s1`, are recognized the same way as separated value flags. If an unrecognized OpenClaw option appears before `--json`, NemoClaw also keeps the command on the normal passthrough path so OpenClaw remains the argv source of truth.
When the top-level OpenClaw `--json` output flag is present, the wrapper uses a captured no-TTY path with a `64 MiB` buffer so `stdout` stays parseable JSON. Raw `stderr`, including structured JSON diagnostics, is forwarded unchanged. NemoClaw appends failed-tool or untrusted-child provenance only from the `stdout` JSON. The wrapper reads completion markers only from the final matching OpenClaw response envelope: a local `{ payloads, meta }` response or a gateway `{ status, result: { payloads, meta } }` response. It ignores earlier JSON progress or log records. It exits with status `1` when that metadata contains `error.kind: "incomplete_turn"`, `livenessState: "abandoned"`, or a `timeoutPhase` value, even when the envelope reports success. `replayInvalid: true` means replaying the turn may repeat side effects. The wrapper accepts it only with evidence of a completed run, a successful tool summary, and final assistant text matching the reply payload. Otherwise it exits with status `1` and the existing inspect-before-retry guidance. Marker-shaped values inside tool results, tool-call arguments, or other descendants do not change the exit status. A turn can run every tool successfully and still become abandoned before it produces a reply. The wrapper writes the unchanged JSON trace to `stdout` before it reports the incomplete turn, so the partial tool trace remains available. The wrapper writes the verdict, the detected markers, and verify-before-retry guidance to `stderr`. A `timeoutPhase` value names the phase the deadline fired in, so the wrapper writes deadline guidance in place of the generic incomplete-turn text. Tool calls in a partial trace may have already applied side effects, so verify what the turn changed before you retry it. The wrapper passes through an upstream non-zero exit status unchanged. Literal `--json` values consumed by flags such as `-m` or `--reply-channel`, or arguments after `--`, stay on the normal passthrough path. Documented value flags written as `--flag=value`, such as `--session-id=s1`, are recognized the same way as separated value flags. If an unrecognized OpenClaw option appears before `--json`, NemoClaw also keeps the command on the normal passthrough path so OpenClaw remains the argv source of truth.

Common OpenClaw flags include `-m <text>`, `--session-id <id>`, `--agent <id>`, `--model <id>`, `--thinking <level>`, `--json`, `--deliver`, `--reply-channel <channel>`, and `--timeout <seconds>`. For OpenClaw sandboxes and registry fallbacks, `$$nemoclaw <name> agent --help` prints the wrapper-level summary locally. Invoke `$$nemoclaw <name> exec -- openclaw agent --help` to view the upstream OpenClaw help text directly. For registered terminal-runtime sandboxes, bare invocations and `--help` are forwarded to the terminal command, so a LangChain Deep Agents Code sandbox receives `dcode` for `$$nemoclaw <name> agent` and `dcode --help` for `$$nemoclaw <name> agent --help`.

Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,112 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0

apiVersion: nemoclaw.nvidia.com/managed-inference/v1
kind: ServingPreset

metadata:
id: vllm.dgx-spark-gb10.single-64gb.qwen3-6-35b-a3b-nvfp4
displayName: Qwen3.6 35B-A3B NVFP4 on one 64 GB DGX Spark
supportState: experimental
validation:
level: software
evidence: src/lib/inference/vllm-runtime-selection.test.ts

spec:
selection: automatic
# Prefer the existing larger-memory profile whenever its capacity floor is met.
priority: 549

requirements:
all:
- readiness:
scope: everyNode
kind: qualification
id: host.platform.dgx_spark
status: qualified
- readiness:
scope: everyNode
kind: capability
id: host.platform.supported
state: present
- readiness:
scope: everyNode
kind: capability
id: host.platform.dgx_spark
state: present
- readiness:
scope: everyNode
kind: capability
id: host.docker.available
state: present
- readiness:
scope: everyNode
kind: capability
id: host.docker.daemon_reachable
state: present
- readiness:
scope: everyNode
kind: capability
id: host.docker.runtime_supported
state: present
- readiness:
scope: everyNode
kind: capability
id: host.docker.storage_compatible
state: present
- readiness:
scope: everyNode
kind: capability
id: host.gpu.nvidia_available
state: present
- readiness:
scope: everyNode
kind: capability
id: host.gpu.container_toolkit_available
state: present
- readiness:
scope: everyNode
kind: capability
id: host.gpu.cdi_healthy
state: present
- readiness:
scope: everyNode
kind: observation
id: host.os.platform
comparison:
operator: equals
value: linux
- readiness:
scope: everyNode
kind: observation
id: host.os.architecture
comparison:
operator: equals
value: arm64
- readiness:
scope: everyNode
kind: observation
id: host.docker.runtime
comparison:
operator: equals
value: docker
- readiness:
scope: everyNode
kind: observation
id: host.gpu.count
comparison:
operator: at-least
value: 1
- readiness:
scope: everyNode
kind: observation
id: host.gpu.driver_version
comparison:
operator: version-at-least
value: 580.65.06

plan:
backend: vllm
platform: spark
interactive: true
recipeRef: vllm.qwen3-6-35b-a3b-nvfp4.spark-single-64gb.v1
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0

apiVersion: nemoclaw.nvidia.com/managed-inference/v1
kind: ServingRecipe

metadata:
id: vllm.qwen3-6-35b-a3b-nvfp4.spark-single-64gb.v1
displayName: Qwen3.6 35B A3B NVFP4 with vLLM on 64 GB DGX Spark

spec:
backend: vllm

modelRef: vllm.qwen3-6-35b-a3b-nvfp4.v1
runtime:
image: nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2
imageDownloadSizeBytes: 9603085145
imageUnpackedSizeBytes: 27658526720
minimumComputeCapability: 121
# A nominal 64 GB Spark reports about 61.6 GB of usable unified memory.
minimumGpuMemoryBytes: 60000000000
pullTimeoutSeconds: 43200
architecture: arm64
networkMode: bridge
ipcMode: host
sharedMemoryBytes: 68719476736
gpuRequest: all
devices: []
ulimits:
memlock: -1
stackBytes: 67108864
modelCache:
source: huggingface-cache
target: /root/.cache/huggingface
temporaryFilesystems: []
environment: {}

execution:
materializerRef: vllm.host-local/v1
lifecycleRef: vllm.host-local.lifecycle/v1
orchestrationRef: vllm.host-local.standard/v1

serve:
authentication: bearer
directInstall:
authentication: none
fixedArguments: false
catalogReceipt: false
executable: /usr/local/bin/vllm
arguments:
- name: --max-model-len
value: 32768
- name: --gpu-memory-utilization
value: 0.5
- name: --dtype
value: auto
- name: --quantization
value: modelopt
- name: --kv-cache-dtype
value: fp8
- name: --attention-backend
value: flashinfer
- name: --moe-backend
value: marlin
- name: --max-num-seqs
value: 1
- name: --max-num-batched-tokens
value: 4096
- name: --enable-chunked-prefill
- name: --async-scheduling
- name: --enable-prefix-caching
- name: --enable-auto-tool-choice
- name: --tool-call-parser
value: qwen3_coder
- name: --reasoning-parser
value: qwen3
- name: --load-format
value: fastsafetensors

readiness:
timeoutSeconds: 1800
expectedModel: nvidia/Qwen3.6-35B-A3B-NVFP4
68 changes: 44 additions & 24 deletions src/lib/actions/sandbox/agent/passthrough-json.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -388,30 +388,50 @@ describe("runAgentJsonPassthrough", () => {
expect(exit).toHaveBeenCalledWith(1);
});

it("keeps a completed turn at exit 0 so the incomplete-turn check does not misfire", async () => {
const payload = JSON.stringify({
status: "ok",
summary: "completed",
result: { payloads: [{ text: "PONG" }], meta: { livenessState: "working" } },
});
const runDispatch = vi.fn(async (_request: OpenShellSandboxSessionRequest) => ({
outcome: { kind: "exited" as const, exitCode: 0 },
stdout: payload,
stderr: "",
}));
const { exit, proc } = makeProc();

await expect(
runAgentJsonPassthrough("alpha", ["openclaw", "agent", "--json"], proc, {
getGatewayName: () => null,
getOpenshellBinary: () => "openshell",
runDispatch,
stdinIsTty: () => false,
}),
).rejects.toThrow("__exit:0");

expect(exit).toHaveBeenCalledWith(0);
});
it.each([
{ replayInvalid: false, corroborated: true, exitCode: 0 },
{ replayInvalid: true, corroborated: true, exitCode: 0 },
{ replayInvalid: true, corroborated: false, exitCode: 1 },
])(
"returns $exitCode for replayInvalid=$replayInvalid with corroborated=$corroborated",
async ({ replayInvalid, corroborated, exitCode }) => {
const payload = JSON.stringify({
status: "ok",
summary: "completed",
result: {
payloads: [{ text: "PONG" }],
meta: {
aborted: false,
replayInvalid,
stopReason: "stop",
finalAssistantVisibleText: "PONG",
...(corroborated ? { toolSummary: { calls: 1, failures: 0, tools: ["exec"] } } : {}),
},
},
});
const runDispatch = vi.fn(async (_request: OpenShellSandboxSessionRequest) => ({
outcome: { kind: "exited" as const, exitCode: 0 },
stdout: payload,
stderr: "",
}));
const { exit, proc, stderr, stdout } = makeProc();

await expect(
runAgentJsonPassthrough("alpha", ["openclaw", "agent", "--json"], proc, {
getGatewayName: () => null,
getOpenshellBinary: () => "openshell",
runDispatch,
stdinIsTty: () => false,
}),
).rejects.toThrow(`__exit:${String(exitCode)}`);

expect(exit).toHaveBeenCalledWith(exitCode);
expect(stdout.join("")).toBe(payload);
expect(stderr.join("").includes("did not complete")).toBe(exitCode === 1);
expect(stderr.join("").includes("replayInvalid=true")).toBe(exitCode === 1);
expect(stderr.join("").includes("Inspect the partial JSON trace")).toBe(exitCode === 1);
},
);

it("keeps a healthy response at exit 0 after a marker-bearing JSON log record", async () => {
const payload = [
Expand Down
2 changes: 2 additions & 0 deletions src/lib/inference/serving/host-local-vllm-selection.ts
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ import type {
} from "./types.js";

export interface MaterializedHostLocalVllmSelection {
readonly displayName?: string;
readonly profile: VllmProfile;
readonly model: VllmModelDef;
readonly presetId: string;
Expand Down Expand Up @@ -140,6 +141,7 @@ export function materializeHostLocalVllmSelection(
const model = materializeHostLocalVllmModel(recipe, directInstall, baseProfile.platform);
const gpuMemoryUtilization = hostLocalVllmGpuMemoryUtilization(recipe);
return {
displayName: preset.metadata.displayName,
presetId: preset.metadata.id,
recipeId: recipe.metadata.id,
model,
Expand Down
31 changes: 31 additions & 0 deletions src/lib/inference/vllm-fixed-catalog-install.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -326,6 +326,37 @@ describe("fixed catalog vLLM installs", () => {
);
}

it.each(["automatic", "picker", "resume"] as const)(
"installs the bounded Spark recipe through %s selection at the reported 64 GB capacity",
async (mode) => {
const profile = detectVllmProfile({ platform: "spark", type: "nvidia" })!;
const readinessReports = vllmInstallTestReadinessAtMemory(profile, 61_614_325_760);
await withActualSelectionGuard(readinessReports);
mockSuccessfulVllmInstall(mocks, profile.containerName);

const result = await installVllm(profile, {
hasImage: true,
nonInteractive: mode !== "picker",
promptFn: vi.fn(async (question: string) => (question.includes("Continue") ? "y" : "1")),
...(mode === "resume" ? { modelIntent: "qwen3.6-35b-a3b-nvfp4" } : {}),
readinessReports,
});

expect(result, spies.errSpy.mock.calls.flat().join("\n")).toEqual({ ok: true });
expect(spies.logSpy).toHaveBeenCalledWith(
" Selected for your hardware: Qwen3.6 35B-A3B NVFP4 on one 64 GB DGX Spark",
);
expect(spies.logSpy).toHaveBeenCalledWith(" Context limit: 32768 tokens");
expect(mocks.dockerRunDetached).toHaveBeenCalledOnce();
const command = mocks.dockerRunDetached.mock.calls[0]![0].at(-1) as string;
expect(command).toContain("vllm serve nvidia/Qwen3.6-35B-A3B-NVFP4");
expect(command).toContain("--max-model-len 32768");
expect(command).toContain("--max-num-seqs 1");
expect(command).toContain("--max-num-batched-tokens 4096");
expect(command).toContain("--gpu-memory-utilization 0.5");
},
);

it("resumes a checkpointed model under an explicitly selected serving preset", async () => {
const profile = detectVllmProfile({ platform: "spark", type: "nvidia" })!;
const modelIntent = "muse-glimmer-30b";
Expand Down
4 changes: 3 additions & 1 deletion src/lib/inference/vllm-models.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -605,7 +605,9 @@ describe("vllm model registry", () => {

it("builds the MTP-free NVFP4 serve command for DGX Spark (#7127)", () => {
const qwen35b = VLLM_MODELS.find((m) => m.envValue === "qwen3.6-35b-a3b-nvfp4");
const cmd = buildVllmServeCommand(qwen35b!);
const sparkProfile = detectVllmProfile({ platform: "spark", type: "nvidia" })!;
const sparkModel = resolveVllmModelRuntime(sparkProfile, qwen35b!, "arm64").model;
const cmd = buildVllmServeCommand(sparkModel);
// The current NVIDIA model card no longer needs Spark-specific env exports.
expect(cmd).not.toContain("VLLM_USE_FLASHINFER_MOE_FP4");
expect(cmd).not.toContain("VLLM_FP8_MOE_BACKEND");
Expand Down
Loading
Loading