Skip to content

Sandbox supervisor reports terminal Error phase for a healthy container during GPU-patch recreate #2117

Description

@latenighthackathon

Surfaced via NVIDIA/NemoClaw#5662 (native-Linux GPU onboard). During NemoClaw's GPU-patch recreate (a docker stop + docker run to add device passthrough), openshell sandbox list / get reports the sandbox in a terminal Error phase while the underlying container is running, healthy, and exit_code=0.

NemoClaw only reads the reported phase and its own code documents the ownership boundary: the preferred fix lives at the OpenShell gateway/supervisor, and a NemoClaw-side health-aware retry was explicitly rejected (src/lib/onboard/docker-gpu-supervisor-reconnect.ts). NemoClaw #4316 / #4407 already fixed the classification and timing on the NemoClaw side (the fast-fail message), so this report is specifically the upstream condition.

Expected: the supervisor phase reflects the healthy / recreating container during a stop+run recreate rather than surfacing a terminal Error.

Repro context: native Linux, GPU-enabled sandbox, during the GPU-patch recreate window; reporter diagnostics show phase Error while the container is running/healthy with exit_code=0.

Activity

  1. self-assigned this
    on Jul 3, 2026
  2. elezar commented on Jul 3, 2026

    @elezar
    Member

    Investigation confirms that the reported Error phase can result from the Docker driver observing both the stopped backup and the healthy replacement with the same sandbox identity labels, then selecting the stopped container for the cached snapshot.

    PR #2116 is related, but should not yet be treated as the fix. This issue and the changes proposed in #2116 raise a broader ownership question: should matching, copyable Docker labels be sufficient for an out-of-band container to become the authoritative driver-managed instance?

    Note that this is also not GPU-specific and affects any out-of-band container modifications.

  3. cv commented on Jul 6, 2026

    @cv
    Contributor

    Maintainer status ping: NVIDIA/NemoClaw#5662 is currently a v0.0.75 test blocker and is waiting on this OpenShell issue. Is there a confirmed fix path or ETA we can align the release train to? If another issue or PR now owns the work, please point us there. Thanks.

  4. added
    area:sandboxSandbox runtime and isolation work
    test:e2eRequires end-to-end coverage
    and removed
    state:triage-neededOpened without agent diagnostics and needs triage
    on Jul 13, 2026
  5. added theissue type on Jul 13, 2026
  6. github-actions commented on Aug 29, 2026

    @github-actions

    This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:computearea:sandboxSandbox runtime and isolation workstate:staleInactive item at risk of automatic closure.test:e2eRequires end-to-end coveragetopic:compatibilityCompatibility-related work

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions