Skip to content

Fix Docker orphan cleanup and restart-safe recreation handoff #8720

Description

@apurvvkumaria

Description

Four failures in E2E run 31367710033 at commit 3ac3a77557aea903c62a91634e5c4aebc5d7a414 converge on the Docker/OpenShell container-ownership handoff used by rebuild and restart-safe recreation.

Two failure modes are visible:

  1. After OpenShell reports the new sandbox Ready, NemoClaw replaces its Docker container with the persistent startup-command clone. The OpenShell supervisor never reconnects and the required forward never appears, leaving the sandbox unready.
  2. During upgrades from NemoClaw 0.0.74 and 0.0.89, the sandbox remains in the NemoClaw registry but is absent from the upgraded live gateway. An old, exactly labeled Docker container survives. Rebuild creates a replacement, and later privileged lifecycle work refuses to continue because two running containers have the same sandbox ownership label.

Affected jobs:

The passing 0.0.55 upgrade control follows the normal live-gateway deletion path and does not leave the old container behind. The passing Hermes channel-cycle control also avoids the failing OpenClaw recreation outcome.

This work is separate from #8662 and #8683, which address the public Shields stop/start recovery path.

Reproduction

Run the four existing E2E targets above against 3ac3a77557aea903c62a91634e5c4aebc5d7a414.

Key observed output:

Sandbox reported Ready before create stream exited; continuing.
Recreating OpenShell Docker sandbox container with restart-safe startup...
Waiting for OpenShell supervisor to reconnect to the recreated container...
Still waiting for forward on port 18789 to register...
sandbox is not ready

For the upgrade cases:

Sandbox is registered locally but absent from the live OpenShell gateway.
Multiple running OpenShell containers are labeled for sandbox 'e2e-gw-survivor'; refusing ambiguous lifecycle execution.

Expected behavior

  • Rebuild/recreation finishes with exactly one running OpenShell-owned container for the sandbox.
  • A registry/live-gateway mismatch cleans up only the exactly owned orphan before replacement creation.
  • Restart-safe recreation does not race the still-active OpenShell create operation.
  • Failure before ownership cutover preserves or restores the original container; success removes the rollback backup.
  • The supervisor reconnects and the sandbox returns to Ready with required forwards registered.

Implementation constraints

Keep this a small ownership-handoff fix:

  • Do not introduce a new lifecycle state machine, journal, coordinator, or recovery framework.
  • Do not add diagnostic collection, new diagnostic formats, or broad error-reporting work.
  • Keep the production-code change below 200 changed lines.
  • Avoid unrelated refactors and new abstractions beyond one narrowly scoped helper if needed.
  • Add only targeted tests for the exact orphan-cleanup and recreation-handoff branches.
  • Do not add broad test coverage or expand the general E2E matrix.
  • Prepare a release-ready PR immediately after the targeted tests pass; normal required CI still applies.

Acceptance criteria

  • An exactly owned orphan is removed before rebuilding a registry-only sandbox; foreign or ambiguous containers remain fail-closed.
  • Restart-safe recreation cannot overlap an unfinished OpenShell create ownership transition.
  • Successful cutover leaves one running labeled container and no stopped rollback backup.
  • Failed cutover restores the original container without starting two labeled containers.
  • The existing sessions-agents-cli and OpenClaw channels-stop-start scenarios pass.
  • The existing 0.0.74 and 0.0.89 gateway-upgrade scenarios pass.
  • Targeted unit/integration tests cover the two regression branches.
  • Production-code changes remain below 200 changed lines with no new lifecycle state machine or diagnostics expansion.

Environment

  • GitHub Actions Linux x86_64 runners
  • Docker/OpenShell driver
  • OpenShell 0.0.101 after upgrade
  • NemoClaw commit 3ac3a77557aea903c62a91634e5c4aebc5d7a414

Checklist

  • Reproduced in the linked E2E run
  • Searched existing open issues; no duplicate found

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area: e2eEnd-to-end tests, nightly failures, or validation infrastructurearea: onboardingOnboarding FSM, provider setup, sandbox launch, or first-run flowarea: sandboxOpenShell sandbox lifecycle, runtime, config, or recoveryneeds: triageAwaiting maintainer classificationplatform: containerAffects Docker, containerd, Podman, or images

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions