Skip to content

Self-hosted environment: worker spawned per claimed work item is told to shut down before executing (sandbox-per-session handoff broken; in-process pattern works) #1779

Description

@raghu-giatec

Summary

In a self-hosted environment (Managed Agents), the documented "sandbox per session" pattern — a dispatcher claims a work item via ant beta:worker poll --on-work spawn.sh, then spawns a fresh container whose worker process is supposed to execute that already-claimed item — reliably fails. The spawned worker starts fine, reaches api.anthropic.com fine, then shuts itself down within 2–6 seconds without executing anything (exit code 0). The session is left permanently stuck at stop_reason: requires_action and the work item is never re-offered to any worker.

We reproduced the identical failure with two independent client implementations:

  • the ant CLI (ant beta:worker run), versions 1.15.0 and 1.17.0
  • this SDK's EnvironmentWorker.handle_item()

Both fail at the same point, in the same way, which suggests the issue is server-side (or an undocumented requirement of the claim-handoff contract) rather than a client defect. Filing here because the SDK repro is the cleanest minimal case, and we don't have a direct support line.

Environment

  • Environment ID: env_01UPQK3XYX67tnteLTvZw891 (type self_hosted)
  • Agent ID: agent_014h5P54ckDkXpUG1NVdqY4G
  • Runtime: Linux amd64, Kubernetes Job pod, Debian bookworm-slim base
  • Reproduced identically across 5+ sessions, e.g. sesn_011zBktK84XqgQapwZCSunbU, sesn_017jQMiZ3mSv2bbtPFAWtayn, sesn_01GSgDdFAJ9JPKgx8oSt1UA6, sesn_01BjrqoHYCysBHz4tFboyfDU, sesn_014f4Y2BQrjJFaHZavJCxxVH
  • SDK repro session: sesn_01GvoUeeRqmHe6dKbfPkCtcH

Reproduction steps

  1. Create a self-hosted environment (POST /v1/environments, config.type: self_hosted); generate its environment key via the Console.
  2. Run a dispatcher: ant beta:worker poll --on-work ./spawn.sh, where spawn.sh follows the documented shape — one fresh container per claimed work item, with ANTHROPIC_SESSION_ID, ANTHROPIC_WORK_ID, ANTHROPIC_ENVIRONMENT_ID, ANTHROPIC_ENVIRONMENT_KEY injected as env vars (our implementation spawns a Kubernetes Job rather than docker run; functionally equivalent).
  3. Create an Agent (POST /v1/agents) and a Session (POST /v1/sessions referencing agent + environment_id).
  4. Send a trivial user.message event, e.g. "List every file in /workspace and print their sizes."

SDK variant of the spawned worker (fails identically to the CLI)

async with AsyncAnthropic(auth_token=environment_key) as client:
    worker = EnvironmentWorker(
        client, environment_id=environment_id,
        environment_key=environment_key, workdir="/workspace",
    )
    await worker.handle_item(
        work_id=work_id, environment_id=environment_id,
        session_id=session_id, environment_key=environment_key,
    )

Observed behavior

  • The dispatcher polls, claims exactly one work item (work_type=session), and spawns exactly one container per claim — this part works on every attempt.

  • The spawned worker's logs show, consistently within 2–6 seconds of startup:

    msg="heartbeat reports shutdown" state=stopping
    msg="reconcile list failed" component=session-tool-runner error="context canceled"
    

    and the process exits 0. The SDK variant behaves the same (ran ~3s, exit 0, no error output).

  • Meanwhile, via the Sessions API, the conversation proceeds normally on the model side: user.message → agent.thinking → agent.message → agent.tool_use (bash, evaluated_permission: allow), then stop_reason: requires_action referencing that tool_use, and session.status_idle.

  • The tool_use is never fulfilled. No second work item is ever claimed for the session. It is stuck permanently.

Ruled out

  • Our dispatch logic — found and fixed two real bugs in it (Jobs keyed by session_id instead of work_id; a heredoc quoting bug); failure persisted unchanged after both fixes.
  • Env vars — verified all four are non-empty and correct on the running container via kubectl get pod -o jsonpath.
  • Network — DNS + TLS to api.anthropic.com confirmed working from inside the spawned pod.
  • --max-idle timing — tried 300s vs. the 60s default, plus explicit --environment-id; no change. Reverted to the exact zero-flag documented pattern for this report.
  • Stale CLI — bumped 1.15.0 → 1.17.0 (the version pinned in the self-hosted-sandboxes doc's own Dockerfile examples as of 2026-07-17); identical failure.

Confirmed workaround (what narrowed it down)

The "in-process" pattern — a single long-running process that both claims and executes, ant beta:worker poll --workdir /workspace with no --on-work — works correctly on the first try: full tool-call cycle, results posted, clean stop_reason: end_turn.

So the failure is specifically in resuming an already-claimed work item in a process other than the one that claimed it. The observable shape (server heartbeat replying "shutdown" to the freshly spawned worker) is consistent with the claim/lease being bound to the claiming process's identity, so the handed-off worker is told to stop — but that's inference; we can't see the server side.

Question

Is there a known issue — or an undocumented handshake/requirement — for how a spawned ant beta:worker run / EnvironmentWorker.handle_item() process adopts a work item claimed by a different (dispatcher) process via the injected env vars? The session IDs above should allow server-side correlation. We've temporarily rearchitected to the in-process pattern, but we'd like to return to sandbox-per-session isolation once this is resolved.

Activity

  1. craigie-ant commented on Jul 24, 2026

    @craigie-ant
    Contributor

    Thanks for the exceptionally thorough report — the session IDs, the ruled-out list, and especially the confirmed workaround made this quick to pin down. Good news: nothing is broken server-side or in the claim hand-off, and there's no process-identity binding on the lease. The failure comes from one line in the dispatcher.

    What's happening

    ant beta:worker poll --on-work spawn.sh runs the script synchronously for each claimed item and treats the script's exit as "this work item is finished": as soon as spawn.sh returns, the poller posts a graceful stop for that work item (the underlying Go WorkPoller documents this: "After the consumer finishes with a yielded item (the next Next call or Close), the poller posts Stop for that work item"). The documented spawn.sh relies on that — exec docker run --rm … runs the container in the foreground, so the script doesn't return until the sandbox exits.

    Your spawn.sh creates a Kubernetes Job, and kubectl create/apply returns as soon as the API object is accepted, so:

    1. the dispatcher claims the item; spawn.sh creates the Job and exits 0 within milliseconds;
    2. the poller immediately posts stop → the work item transitions to stopping;
    3. 2–6 s later the pod starts, its worker sends its first heartbeat, the response carries state: "stopping", and it shuts down gracefully — that's your heartbeat reports shutdown state=stopping line, followed by the session runner's context being cancelled (reconcile list failed … context canceled) and a clean exit 0. EnvironmentWorker.handle_item() does the same (it logs the shutdown at INFO, which is why you saw no output).

    Fixing

    I believe if you update spawn.sh to block until the sandbox finishes, you should get the behaviour you expect.

    Our docs are not clear about this so I'm working on improving them.

    Let me know if that doesn't make sense / doesn't work!

  2. priyanka25aug commented on Jul 28, 2026

    @priyanka25aug

    Thanks for the explanation above — that confirms exactly what I traced through the code.

    I've put together a complementary SDK-side fix in src/anthropic/lib/environments/_worker.py (the hand-written carve-out, not generated code): it adds a heartbeat_count tracker to _heartbeat_loop and emits a log.warning with a specific diagnostic message when shutdown is signalled on the very first beat — naming the root cause and pointing to the blocking requirement (docker run foreground vs -d, kubectl wait vs kubectl apply). Also added a Note: to the handle_item docstring.

    Opening a PR now — it's a logging-only change so existing behaviour for normal teardown is unchanged.

  3. raghu-giatec commented on Jul 29, 2026

    @raghu-giatec
    Author

    Confirmed — that was exactly it. Updated our spawn script to block until the Kubernetes Job reaches a terminal condition (polling both Complete and Failed, bounded by the Job's activeDeadlineSeconds) instead of returning right after kubectl apply, and the sandbox-per-session pattern now works end-to-end on the first try: dispatcher claims → Job spawned → ant beta:worker run executes write/read/bash tools → results posted → clean end of turn → dispatcher unblocks. Verified live on our real deployment today.

    For anyone else landing here from Kubernetes: kubectl apply/create returning is NOT your sandbox finishing — you must wait on the Job (and kubectl wait --for=condition=complete alone will hang on a failed Job until timeout, so poll both terminal conditions). A docs note on the --on-work exit semantics would indeed have saved us two weeks — glad to hear that's coming. Thanks for the fast and thorough answer, closing this out.

  4. dtmeadows-ant commented on Aug 4, 2026

    @dtmeadows-ant
    Contributor

    Hi @raghu-giatec, glad that unblocked you — closing this out per your confirmation. As covered in #1779 (comment), the poller treats the --on-work script's exit as the work item finishing, so the script must block until the sandbox completes; a docs clarification on those exit semantics is in the works.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions