Repository navigation
fix(platform): skip an unidentifiable Docker socket instead of aborting detection - #10253
Conversation
…ng detection detectDockerHost returned null for the whole function the moment it found a reachable-but-unclassified socket, instead of trying the remaining candidates. On macOS, a stale Colima socket that answers but can't be classified broke Docker auto-detection entirely, even when a later candidate (e.g. Docker Desktop) would have resolved cleanly. Skip an unknown-identity candidate rather than aborting (#10248). Signed-off-by: Jason Ma <jama@nvidia.com>
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall line coverage in commit 1012a71 in the TypeScript / code-coverage/cliThe overall line coverage in commit 1012a71 in the Updated |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
Included review availability: Your plan provides up to 12 included reviews per hour; 5 remain after this review. 📝 WalkthroughWalkthrough
ChangesDocker host detection
Estimated code review effort: 2 (Simple) | ~10 minutes Merge Risk: ⚪ Minimal · up to The fix lets Docker auto-detection continue past a reachable but unidentifiable socket so a later valid engine can still be selected, with a regression test covering the scenario; no actionable merge-blocking risk remains at the current head. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Linked Issues checkExplanation The PR satisfies issue ✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Comment |
…on test The codebase-growth-guardrails check flags new if statements added to test files. Rewrite the mock as a ternary chain instead, keeping the same probe behavior (#10248). Signed-off-by: Jason Ma <jama@nvidia.com>
PR Review Advisor — No blocking findings reportedAdvisor assessment: No blocking advisor findings reported E2E guidanceAdvisory only. A maintainer can dispatch the default E2E suite for the commit under review. Recommended E2E: None Manual-only E2E: This automated review informs maintainers. Warnings and suggestions do not require a response. A maintainer decides whether to merge. |
Signed-off-by: Apurv Kumaria <akumaria@nvidia.com>
apurvvkumaria
left a comment
There was a problem hiding this comment.
Security review: PASS for the current revision.
- Secrets and sensitive data: PASS. No credential or diagnostic-content changes.
- Input validation and injection: PASS. Candidate socket paths still come from the existing trusted discovery list; no command construction changes.
- Authentication and authorization: PASS. No permission, identity, or daemon-access expansion.
- Dependencies and supply chain: PASS. No dependency, workflow, or artifact-source changes.
- Error handling and information disclosure: PASS. An unidentifiable reachable socket is skipped, while two conflicting identified engines still fail closed.
- Cryptography and data protection: PASS. No cryptographic or persisted-data changes.
- Configuration and deployment: PASS. Explicit Docker host and context precedence are unchanged.
- Security testing: PASS. All 35 platform integration tests, repository checks, CLI build/type checking, and normal hooks pass.
- System security: PASS. The change selects only a later candidate whose identity probe succeeds and does not weaken the conflicting-engine guard.
No documentation update is needed for this internal detection correction. The change adds 29 lines and removes one, below the large-change threshold. Cross-issue review found only issue #10248 and this PR, with no competing open implementation. Jason Ma remains the primary contributor; Apurv Kumaria’s merge commit only refreshes the branch.
|
Refreshed the branch from current main without conflicts. The signed merge commit is Verified. Post-refresh validation passes: 35 platform integration tests, repository checks, CLI build and type checking, and normal push hooks. Jason Ma remains the sole substantive contributor; the maintainer commit is branch synchronization only. |
…10379) ## Summary NemoClaw probes Docker reachability by running `docker version` in a four-name environment (`HOME`, `USER`, `LOGNAME`, `PATH`), but every Docker command it runs afterwards gets the full subprocess allowlist. On a host whose daemon answers only through one of the dropped names, the probe reports the host's own default authority unreachable, detection falls through to the socket candidates, and the CLI pins `DOCKER_HOST` to Podman's rootless socket. Preflight then reports `Docker is not reachable` and points at the docker group, so onboarding stops at its first step on a host whose Docker is healthy. After this change the probe asks the same question the later commands answer, and detection never redirects the CLI on no evidence. ## Related Issue Closes #10367 This removes the mechanisms that produce the reported outcome: a probe environment narrower than the one the predicted commands run in, and a probe that reaches no verdict counting as a refusal. Either can send `DOCKER_HOST` to Podman's socket on a host whose Docker daemon is live. One honest caveat for whoever merges this. The reporter runs a DGX Spark with Docker and Podman installed; I have no such host and they have not yet answered the two diagnostic commands I asked for on the issue, so the cure is reasoned from the code path, not observed on their machine. If their `docker version` under the old four-name environment turns out to exit `0` quickly, neither fix explains their failure and the issue should be reopened rather than left closed. Two details from the report stay out of scope either way: the docker-group remediation text that names the wrong cause, and the `docker info` versus `docker version` disagreement on an unhealthy daemon. ## Changes - `buildDockerProbeEnv` now selects names with `isSubprocessEnvNameAllowed`, the same allowlist `buildSubprocessEnv` gives real Docker commands, and drops an ambient `DOCKER_HOST` so the probe still pins the authority under test. The probe predicts whether those commands reach a daemon, so it must not ask under a narrower environment: `SSH_AUTH_SOCK` authenticates an `ssh://` Docker context and the proxy names decide how a `tcp://` one is routed. (An earlier revision of this description claimed `XDG_RUNTIME_DIR` selects a rootless daemon socket for the Docker CLI. I tested that and it is false — the CLI ignores a listening `docker.sock` in the runtime directory — so the justification is corrected here and in the code comment.) - `probeDockerHost` reports `inconclusive` when the Docker CLI cannot be spawned or the 3-second probe timeout kills it, and `detectDockerHost` holds the host default in that case. A probe that never answered is not an observed refusal, so it must not move the whole CLI to a fallback socket. - Linux socket candidates are now ordered `/run/docker.sock`, `/var/run/docker.sock`, `/run/user/<uid>/docker.sock`, then Podman's. Rootless Docker's socket sits beside Podman's in the same runtime directory and was never a candidate. - `buildDockerProbeEnv` also applies `withLocalNoProxy`, which `buildSubprocessEnv` gives every real Docker command. Without it, forwarding the proxy names could route a probe of a local `tcp://` authority through a host proxy that the real commands bypass — the same defect class, reintroduced by the fix. - `ci/source-architecture-budget.json`: reading the shared allowlist raises the recorded fan-in of `src/lib/subprocess-env.ts` from 24 to 25. - Onboarding now bounds the existing `docker info` and `docker version` preflight calls at 15 seconds, so preserving an inconclusive default authority cannot leave onboarding waiting without a limit. ## Risk family `src/lib/platform.ts` puts this PR in the tier-3 `platform-install` family, whose required job is `cloud-onboard`. That workflow has no `pull_request` trigger, so it selects on the post-merge push to `main` rather than here. Say the word if you want a `cloud-onboard` run before merge and I will arrange it. ## Not in this PR CodeRabbit's merge-risk note and the PR Review Advisor both point at the mixed-identity bail: when the default authority is dead and both a Docker socket and a Podman socket answer, `detectDockerHost` returns `null` and the CLI keeps its dead default. That path is pre-existing and unchanged here, and removing it reverses a decision recorded in #8823 and #10253, whose security review cited it as a pass criterion. It is a maintainer call, so it is a separate stacked PR — #10387 — with the reversal argued. This PR leaves the guard exactly as it was. ## Type of Change - [x] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Quality Gates - [x] Tests added or updated for changed behavior - [ ] Existing tests cover changed behavior — justification: - [ ] Tests not applicable — justification: - [x] Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging) - [ ] Sensitive-path review completed or maintainer-approved waiver recorded — reviewer/approval link/justification: pending review on this PR - [ ] Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue: The probe environment stays an allowlist. `test/e2e-runtime/platform.test.ts` fails the probe binary when `NVIDIA_INFERENCE_API_KEY` crosses the boundary, in the new test and in the existing `#8816` one. ## Verification - [x] PR description includes a `Signed-off-by:` line and every commit appears as `Verified` in GitHub - [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or `npm run validate:pr` passed after refreshing `origin/main` when hooks were skipped or unavailable - [x] Targeted behavior tests pass for the current change set, or tests are marked not applicable above — command/result: `npx vitest run test/e2e-runtime/platform.test.ts` gives 38 passed, and a focused sweep over the Docker-authority files (`platform`, `runner`, `preflight-docker-host`, `domain/docker-host`, `subprocess-env`, `readiness/host`, `container-engine`, `docker-authority-profile`) gives 187 passed. `npm run typecheck:cli` and `npm run lint` pass. The focused platform and Docker-preflight timeout suites cover 2 files and 40 tests, and the codebase growth guardrails cover 33 tests. All three original probe changes were confirmed red first: without the probe-environment change the default-authority test returns `unix:///run/user/1000/podman/podman.sock` where `null` is expected; without `withLocalNoProxy` that same test fails on the proxy-exclusion guard; and without the no-verdict branch, the test whose Docker CLI dies without an exit status selects the Podman socket. - [ ] Applicable broad gate passed — command/result: not run. The change set is two source functions and their tests. - [x] Quality Gates section completed with required justifications or waivers - [x] No secrets, API keys, or credentials committed - [ ] `npm run docs` builds without warnings (doc changes only) - [ ] Doc pages follow the style guide (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) --- Signed-off-by: Dongni Yang <dongniy@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved Docker environment detection across Linux setups, including rootless Docker and Podman installations. * Prioritized native Docker sockets for more accurate runtime detection. * Prevented incorrect Docker or Podman classification when the Docker CLI is unavailable or unresponsive. * Preserved relevant runtime and proxy settings while excluding ambient configuration that could cause misleading results. * **Tests** * Expanded coverage for socket prioritization, environment handling, proxy behavior, and inconclusive Docker probes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> ## Merge with `main` (`7d1476a7a1`, 2026-09-09) `reviewed-npm-audit` failed on `b20e7fcd22` with `Tencent WeChat plugin 2.4.3 locked runtime graph lock SHA-256 mismatch`: the branch carried the pre-#11023/#11253 expected hash in `ci/reviewed-npm-audit.json` while the trusted action computes the current one. `main` already records the current hash, so this is a clean merge of `main` (104 commits, no conflicts) with no change to the fix itself. It also picks up the patched `js-yaml` pin from #11264. --------- Signed-off-by: Dongni Yang <dongniy@nvidia.com> Signed-off-by: Apurv Kumaria <akumaria@nvidia.com> Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> Co-authored-by: Apurv Kumaria <akumaria@nvidia.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Prekshi Vyas <prekshiv@nvidia.com> Co-authored-by: Prekshi Vyas <34834085+prekshivyas@users.noreply.github.com>
…r-group fix (#11393) Closes #10622 ## Problem When `DOCKER_HOST` is unset, the default Docker authority does not answer, and two discovered sockets answer as different engines (Docker and Podman), `detectDockerHost` deliberately selects neither (#8816, #10253). Nothing named that conflict. Preflight saw an installed Docker, an active service and an unreachable daemon, and printed `docker_group_permission` with `sudo usermod -aG docker $USER`, which does not address the cause. This is the diagnostic the #10387 review asked for. The selection policy is unchanged. ## Change - `src/lib/platform.ts`: the socket-candidate loop moves into a private `selectDockerAuthority` that returns `{ selection, conflict }`. `detectDockerHost` returns `selection` and behaves exactly as before. The new export `observeDockerAuthorityConflict` returns `conflict`: the two candidates, in probe order, when a reachable known engine differs from the one already selected. The fail-closed outcome is unchanged. - `src/lib/onboard/preflight.ts`: `HostAssessment.dockerAuthorityConflict` (optional). `assessHost` observes it only when Docker is installed, `DOCKER_HOST` is unset and valid, the daemon is unreachable, and the probe reached a verdict. The default observer probes this host's own sockets, so it is used only when the assessment probes the local host itself. Callers that inject Docker evidence or a command transport (unit tests, the remote DGX Spark transport) get no observer unless they inject `observeDockerAuthorityConflictImpl`. Cost: nothing on a healthy host; one `docker version` per candidate socket on the failure path. - `src/lib/advisories/checks/host/docker.ts`: new blocking advisory `docker_authority_conflict` ("Choose the Docker authority") after `docker_probe_inconclusive`. It names both sockets and engines, says NemoClaw did not choose and did not diagnose why the default is unreachable, and prints `export DOCKER_HOST='unix://…'` for each socket, single-quoted so a home directory with a space survives. `docker_group_permission`, `start_docker` and `enable_docker_desktop_wsl_integration` stay silent while a conflict is observed, so the conflict is the single Docker advisory. - `docs/reference/system-readiness.mdx`: two sentences next to the existing mixed-fallback statement. ## Scope ruling #10622 proposed (a) a new advisory id or (b) no new id, and asked for a maintainer ruling. None arrived in ten days. This PR implements (a): the existing advisories' titles do not fit a conflict, and an honest message needs its own id. If maintainers prefer (b), the check can fold into the unreachable-Docker path without changing the observation. ## Tests - `test/e2e-runtime/platform.test.ts`: `observeDockerAuthorityConflict` reports both candidates in probe order while `detectDockerHost` still returns `null` on the same inputs. Same-engine candidates, a reachable or inconclusive ambient probe, an unknown identity, and a set `DOCKER_HOST` all yield `null` with no extra probes. The existing `detectDockerHost` tests are untouched (154 lines added, 0 removed). - `src/lib/onboard/preflight-docker-authority-conflict.test.ts` (new): the assessment carries the conflict and plans `docker_authority_conflict` without `docker_group_permission` or `start_docker`; the observer receives `{ env, platform }`; it is not called when `DOCKER_HOST` is set, the daemon is reachable, Docker is not installed, or the probe timed out; a caller that injects `dockerInfoOutput` never reaches the real observer (spied), so the existing hermetic `assessHost` tests spawn no `docker version`. - `src/lib/advisories/checks/host/docker.test.ts`: id order; full reason and command assertions; the conflict is the only advisory with an inactive Linux service, on macOS, and on WSL (each case fails if its gate is removed); no conflict advisory when the daemon is reachable, Docker is missing, or `DOCKER_HOST` is invalid; quoting of a path with a space. `index.test.ts`: registry order. - `src/lib/onboard/preflight.test.ts` is untouched. 255 tests pass across the nine affected files, and `npm run typecheck:cli` reports no error in a changed file. ## Review before opening An adversarial review ran on the draft before this PR was opened: three reviewers, eleven verified findings, all fixed here. They were the hermetic default observer, the WSL gate, the wording about an undiagnosed cause, the shell quoting, and the gate tests listed above. CodeRabbit's second pass asked that the resolver tests assert observable results rather than callback identity (path rule). `2fb87831b7` invokes the selected observer with fixed options: a spied real observer returns the conflict for a local assessment, and an injected observer wins over a spied real one that returns null. Same hook disclosure applies. ## Hooks `c20d266489` was committed with `--no-verify` because the only failing pre-commit hook was the pre-existing `pi-qualification-receipt-refresh` check (`bad object 609d60a…` in this worktree, as on `main`); every other pre-commit hook passed, and the pre-push TypeScript hooks passed once `nemoclaw/dist` was built. ## CodeRabbit round on `c20d266489` → `22ec9c3737` CodeRabbit asked for coverage of the default observer wiring, which no `assessHost` case could reach hermetically: the runner is destructured from a CommonJS `require` at module load, so a spy cannot intercept the real `docker info`. The selection now lives in an exported `resolveDockerAuthorityConflictObserver`, tested directly: the real `observeDockerAuthorityConflict` for a local assessment, no observer when `dockerInfoOutput`, `runCaptureImpl` or `runCaptureExImpl` is injected, and the injected observer when one is given. `assessHost` calls that resolver, so its behaviour is unchanged. The docstring-coverage warning is addressed with one-line docstrings on the touched functions. Committed with `--no-verify` for the same sole failing Pi receipt hook. Signed-off-by: Dongni Yang <dongniy@nvidia.com> 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added onboarding detection for conflicting Docker and Podman authorities. * Displays reachable sockets and detected engines, with instructions to set `DOCKER_HOST` explicitly. * Suppresses conflicting or misleading Docker setup advisories during authority conflicts and inconclusive detection. * **Documentation** * Documented the blocking Docker authority-conflict advisory. * Added guidance for native rootless Podman onboarding. * **Tests** * Added coverage across Linux, macOS, WSL, onboarding preflight, and runtime detection scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Dongni Yang <dongniy@nvidia.com> Signed-off-by: Julie Yaunches <jyaunches@nvidia.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Julie Yaunches <jyaunches@nvidia.com>
<!-- markdownlint-disable MD041 --> ## Summary `nemoclaw host probe` reported `supported` on a host whose selected Docker context was unreachable, while `nemoclaw onboard` refused on the same host seconds later. Docker authority detection treated `DOCKER_HOST` as an operator-chosen authority but not `DOCKER_CONTEXT`, so an unreachable context was silently replaced by a discovered fallback socket and every later consumer measured — and certified — a daemon the operator never selected. Closes #11719. ## Reproduction ```bash docker context create qa-unreachable --docker host=unix:///run/user/$(id -u)/no-such.sock DOCKER_CONTEXT=qa-unreachable docker version # fails: cannot connect unset DOCKER_HOST DOCKER_CONTEXT=qa-unreachable nemoclaw host probe; echo "exit=$?" DOCKER_CONTEXT=qa-unreachable nemoclaw onboard --name dock-conflict; echo "exit=$?" DOCKER_HOST=unix:///run/user/$(id -u)/no-such.sock nemoclaw host probe; echo "exit=$?" ``` **Environment** - Test machine: ubuntu (Ubuntu 24.04.4 LTS, x86_64, no GPU) - Docker 29.8.0 + podman answering on its own socket, matching the reporter's host - NemoClaw `main` HEAD `a05d1f238f84e59e6d4fbdc30e547ec7281772c4`; also reproduced on `v0.0.123` as reported - Sandbox onboarded through a host-local Ollama provider **Observed on `main` (before fix)** ``` $ DOCKER_CONTEXT=qa-unreachable nemoclaw host probe System readiness: supported Findings: none exit=0 $ DOCKER_CONTEXT=qa-unreachable nemoclaw onboard --name dock-conflict ✗ Docker is not reachable. Please fix Docker and try again. - Choose the Docker authority (docker_authority_conflict): The default Docker authority did not answer, and two engines answered on discovered sockets ... exit=1 $ DOCKER_HOST=unix:///run/user/$(id -u)/no-such.sock nemoclaw host probe System readiness: incompatible - [blocking] host.docker.daemon_unreachable: The Docker daemon is unreachable. exit=2 ``` **Observed on `fix/docker-context-authority-11719` (after fix)** ``` $ DOCKER_CONTEXT=qa-unreachable nemoclaw host probe System readiness: incompatible - [blocking] host.docker.daemon_unreachable: The Docker daemon is unreachable. exit=2 ``` The probe now matches the `DOCKER_HOST` control the issue names as correct. Three further live cases: ``` # context selects a different live daemon: the probe describes that daemon, not the default one $ DOCKER_CONTEXT=qa-podman nemoclaw host probe - [blocking] host.docker.runtime_unsupported: The detected container runtime is unsupported. exit=2 # context selects a TCP endpoint: refused in 1.5s without contacting it $ DOCKER_CONTEXT=qa-remote nemoclaw onboard --name dock-remote - Fix the DOCKER_CONTEXT endpoint (invalid_docker_host): DOCKER_CONTEXT selects the Docker context 'qa-remote', which does not resolve to an endpoint onboarding can use ... exit=1 # context reduces to a socket that does not exist $ DOCKER_CONTEXT=qa-unreachable nemoclaw onboard --name dock-missing - Fix the selected Docker endpoint (docker_endpoint_socket_missing): The selected Docker endpoint unix:///run/user/1000/no-such.sock has no socket at that path ... exit=1 ``` ## Analysis `selectDockerAuthority` in `src/lib/platform.ts` accepted `DOCKER_HOST` as the selected authority and returned immediately. With `DOCKER_HOST` unset it probed the *ambient* authority instead — which honors `DOCKER_CONTEXT` through `buildDockerSubprocessEnv` — saw the selected context refuse, and then walked the local socket candidates and adopted one. That fallback exists for a host that never selected an authority; an unreachable *selected* context is a different fact, and treating it the same way redirected NemoClaw to another daemon. `src/lib/runner.ts` then applied that selection process-wide at module load: it set `process.env.DOCKER_HOST` to the fallback socket **and deleted `process.env.DOCKER_CONTEXT`**, so the operator's selector was gone before any command ran. That produced the two surfaces disagreeing on the reporter's host: - **`host probe`** builds a replacement child environment from an exact allow-list in `src/lib/readiness/probe-env.ts`. `DOCKER_HOST` is carried through when it is a supported local socket; `DOCKER_CONTEXT` is not on the list. With the selector already deleted, `docker info` resolved through Docker's default context to the local socket, so `assessHost` saw a reachable daemon and the report certified `supported`. - **`onboard`** runs on the same host, but there both docker and podman answered on discovered sockets, so detection declined to choose (#8816, #10253), left the unreachable selection in place, and reported `docker_authority_conflict`. On a single-engine host the two surfaces agreed — on the wrong daemon: `onboard` completed successfully against the default socket while the operator had selected an unreachable context. `scripts/install.sh` already refuses to proceed when `DOCKER_CONTEXT` does not select the local default target, so the installer and the CLI disagreed about whether a context selection may be ignored. ## Fix **An explicit `DOCKER_CONTEXT` selects the authority exactly as `DOCKER_HOST` does.** `selectDockerAuthority` now resolves the selected context once, locally, with `docker context inspect --format '{{.Endpoints.docker.Host}}'` — a read of the local context store that never contacts a daemon — and: - an absolute local `unix://` socket becomes the selected `DOCKER_HOST`, so readiness probes, preflight, and the container-engine authority all measure the same daemon; - any other endpoint, and a context that does not resolve, keeps the host default. The selection is never widened to an endpoint a later consumer would not otherwise be allowed to contact. In both cases the fallback-socket scan is skipped, so a host that chose an authority is never redirected. Because a selected context is not a host without an authority, `observeDockerAuthorityConflict` no longer names both engines for it — that advisory's remedy (`export DOCKER_HOST=...`) was a misdiagnosis of the operator's own dead context. **A selector that survives to `assessHost` marks the endpoint unusable.** Detection has already reduced every supported context to `DOCKER_HOST`, so a remaining `DOCKER_CONTEXT` names an endpoint onboarding cannot use. `dockerHostInvalid` now covers it, which flows through all five existing consumers of that flag, including `host.docker.endpoint_supported` and onboarding admission. The blocked endpoint is never contacted: `assessHost` skips its Docker probes entirely when the flag is set. `invalid_docker_host` names the variable that actually made the selection, and shell-quotes the context name it prints back into a command. **An explicitly selected endpoint whose socket is absent is no longer diagnosed as a permissions or stopped-daemon problem.** Nothing can listen on a path with no socket, so the `docker_group_permission` remedy (`sudo usermod -aG docker $USER`, a root-equivalent group grant) and the `start_docker` remedy both acted on a ruled-out cause. A new blocking `docker_endpoint_socket_missing` advisory names the absent socket, and the two misdiagnosing advisories stay silent while it is observed — the same gating pattern #10622 established for the authority conflict. Docker's **default** endpoint is deliberately excluded: a missing default socket is exactly the stopped-daemon case `start_docker` exists for. This reproduces on `main` through `DOCKER_HOST` alone, so it is fixed for the whole class, not only the context path. ### Tests `src/lib/onboard/preflight-docker-context-authority.test.ts` pins the contract in both directions: - a selected context naming a supported socket is recorded as the endpoint, and the fallback sockets are never probed; - a context that does not resolve, and one naming `tcp://` / `ssh://` / a relative path, keep the host default; - `DOCKER_CONTEXT` overrides `DOCKER_HOST`, matching Docker CLI precedence; - a blank selector still falls through to the fallback scan (regression lock); - no engine conflict is reported while a context owns the selection, and the conflict is still reported when none does (regression lock); - an unreduced selector blocks with a `DOCKER_CONTEXT`-specific remedy, with shell quoting asserted on a hostile name; - an absent selected socket yields `docker_endpoint_socket_missing` and withholds the group and start remedies, while an existing selected socket still yields `docker_group_permission` and a missing **default** socket still yields `start_docker` (regression locks). - a production-path runner test exercises the real context resolver and verifies the selected endpoint and Docker config reach the final Docker child; - terminal and bidirectional control characters make a selected endpoint invalid before it can reach diagnostics. Each test would have failed before this change on the behavior it pins. ## Changes - `src/lib/domain/docker-host.ts`: reject terminal and bidirectional control characters in selected endpoints. - `src/lib/platform.ts`: resolve an explicit `DOCKER_CONTEXT` to its endpoint and treat it as the selected authority; never replace it with a fallback socket. - `src/lib/runner.ts`: drop the context selector for any non-`env` selection, including one derived from that context. - `src/lib/onboard/preflight.ts`: mark an unreduced context selector as an unusable endpoint; detect an explicitly selected endpoint whose socket is absent. - `src/lib/advisories/checks/host/docker.ts`: name the selecting variable in `invalid_docker_host`; add `docker_endpoint_socket_missing` and silence the two advisories that misdiagnose it. - `src/lib/readiness/host.ts`: `host.docker.host_invalid` names the endpoint rather than one of the two variables that can select it. - `src/lib/onboard/preflight-docker-context-authority.test.ts`: new coverage for the selection, classification, and advisory contract. - `src/lib/advisories/checks/host/docker.test.ts`, `src/lib/advisories/checks/host/index.test.ts`: register the new advisory id in the pinned remediation order. - `docs/reference/system-readiness.mdx`, `docs/reference/troubleshooting.mdx`: document context selection and the new advisory. ### Maintainer follow-up (`c6ed738`) Reviewed the complete hosted Advisor run `34945998133`. Commit `c6ed738` closes its three valid blockers: WSL credential-isolation docs now describe context precedence, a production-path test proves the actual resolver reaches the final Docker child, and selected Docker endpoints reject terminal and bidirectional controls before diagnostics. Exact local evidence: 99 focused CLI tests, 63 runner integration tests, docs validation with 0 errors and 2 warnings, and `npm run validate:pr`. The local Advisor was attempted, but its disposable bootstrap could not resolve the isolated checkout shared dependencies. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [x] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Verification - [x] `npx prek run --all-files` passes - [x] `npm test` passes (touched files at minimum) - [x] Tests added or updated for new or changed behavior - [x] No secrets, API keys, or credentials committed - [x] Docs updated for user-facing behavior changes - [ ] `make docs` builds without warnings (doc changes only) - [x] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) ## AI Disclosure - [x] AI-assisted — tool: Claude Code Signed-off-by: Yanyun Liao <yanyunl@nvidia.com> Signed-off-by: Rebecca Sliter <571084+rsliter@users.noreply.github.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Docker contexts now take precedence over `DOCKER_HOST` when selecting the active endpoint. - Invalid, unsupported, unresolved, or inaccessible endpoints are reported without silent fallback. - Missing or unusable Unix sockets now receive targeted guidance identifying the affected endpoint. - Docker and daemon remediation suggestions are suppressed when they do not apply to the selected endpoint. - Docker environment cleanup prevents stale context or host selections from affecting subsequent commands. - Docker host values containing hidden control or bidirectional text characters are rejected. - **Documentation** - Added troubleshooting guidance for invalid endpoints, missing sockets, verification, and resetting Docker selections. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Yanyun Liao <yanyunl@nvidia.com> Signed-off-by: Rebecca Sliter <571084+rsliter@users.noreply.github.com> Co-authored-by: Rebecca Sliter <571084+rsliter@users.noreply.github.com>
Summary
detectDockerHostreturnednullfor the whole function the moment it found a reachable-but-unidentifiable socket in its candidate loop, instead of skipping to the next candidate. On macOS, a stale Colima socket that answers but can't be classified broke Docker auto-detection entirely, even when a later candidate (e.g. Docker Desktop) would have resolved cleanly.Related Issue
Fixes #10248
Changes
src/lib/platform.ts: changedreturn nulltocontinueon theobservation.identity === "unknown"branch insidedetectDockerHost's socket-candidate loop, so one ambiguous candidate is skipped rather than aborting detection of all remaining candidates — matching the reporter's suggested fix exactly.test/e2e-runtime/platform.test.ts: added a regression test for the exact repro — an earlier reachable-but-unidentifiable socket (stale Colima) must not prevent a later valid candidate (Docker Desktop) from being selected.Note: the reporter also flagged the next line's conflicting-identity check (
if (selectedIdentity && observation.identity !== selectedIdentity) return null) as "worth reviewing" for a similar concern. Left that unchanged — the issue only asks to review it, not fix it, and that check's abort-on-conflict behavior may be intentional (refusing to guess between two differently-identified reachable engines rather than picking one). Flagging it as a possible separate follow-up rather than bundling an unrequested behavior change into this fix.Type of Change
Quality Gates
DGX Station Hardware Evidence
Verification
Signed-off-by:line and every commit appears asVerifiedin GitHubpre-commit,commit-msg, andpre-pushhooks passed, ornpm run validate:prpassed after refreshingorigin/mainwhen hooks were skipped or unavailablenpx vitest run --project integration test/e2e-runtime/platform.test.ts— 35/35 passed locally and again on a clean host checkout (fresh clone +npm ci) on the auto-fix loop's ubuntu verify host.npm testfor broad runtime/test-harness changes;npm run checkfor repo-wide validation/coverage changes — command/result:npm run docsbuilds without warnings (doc changes only)Signed-off-by: Jason Ma jama@nvidia.com
Summary by CodeRabbit
Bug Fixes
Tests