Repository navigation
Self-hosted environment: worker spawned per claimed work item is told to shut down before executing (sandbox-per-session handoff broken; in-process pattern works) #1779
Description
Activity
Thanks for the exceptionally thorough report — the session IDs, the ruled-out list, and especially the confirmed workaround made this quick to pin down. Good news: nothing is broken server-side or in the claim hand-off, and there's no process-identity binding on the lease. The failure comes from one line in the dispatcher.
What's happening
ant beta:worker poll --on-work spawn.sh runs the script synchronously for each claimed item and treats the script's exit as "this work item is finished": as soon as spawn.sh returns, the poller posts a graceful stop for that work item (the underlying Go WorkPoller documents this: "After the consumer finishes with a yielded item (the next Next call or Close), the poller posts Stop for that work item"). The documented spawn.sh relies on that — exec docker run --rm … runs the container in the foreground, so the script doesn't return until the sandbox exits.
Your spawn.sh creates a Kubernetes Job, and kubectl create/apply returns as soon as the API object is accepted, so:
- the dispatcher claims the item; spawn.sh creates the Job and exits 0 within milliseconds;
- the poller immediately posts stop → the work item transitions to stopping;
- 2–6 s later the pod starts, its worker sends its first heartbeat, the response carries state: "stopping", and it shuts down gracefully — that's your heartbeat reports shutdown state=stopping line, followed by the session runner's context being cancelled (reconcile list failed … context canceled) and a clean exit 0. EnvironmentWorker.handle_item() does the same (it logs the shutdown at INFO, which is why you saw no output).
Fixing
I believe if you update
spawn.shto block until the sandbox finishes, you should get the behaviour you expect.Our docs are not clear about this so I'm working on improving them.
Let me know if that doesn't make sense / doesn't work!
Thanks for the explanation above — that confirms exactly what I traced through the code.
I've put together a complementary SDK-side fix in
src/anthropic/lib/environments/_worker.py(the hand-written carve-out, not generated code): it adds aheartbeat_counttracker to_heartbeat_loopand emits alog.warningwith a specific diagnostic message when shutdown is signalled on the very first beat — naming the root cause and pointing to the blocking requirement (docker runforeground vs-d,kubectl waitvskubectl apply). Also added aNote:to thehandle_itemdocstring.Opening a PR now — it's a logging-only change so existing behaviour for normal teardown is unchanged.
Confirmed — that was exactly it. Updated our spawn script to block until the Kubernetes Job reaches a terminal condition (polling both Complete and Failed, bounded by the Job's activeDeadlineSeconds) instead of returning right after
kubectl apply, and the sandbox-per-session pattern now works end-to-end on the first try: dispatcher claims → Job spawned →ant beta:worker runexecutes write/read/bash tools → results posted → clean end of turn → dispatcher unblocks. Verified live on our real deployment today.For anyone else landing here from Kubernetes:
kubectl apply/createreturning is NOT your sandbox finishing — you must wait on the Job (andkubectl wait --for=condition=completealone will hang on a failed Job until timeout, so poll both terminal conditions). A docs note on the --on-work exit semantics would indeed have saved us two weeks — glad to hear that's coming. Thanks for the fast and thorough answer, closing this out.Hi @raghu-giatec, glad that unblocked you — closing this out per your confirmation. As covered in #1779 (comment), the poller treats the
--on-workscript's exit as the work item finishing, so the script must block until the sandbox completes; a docs clarification on those exit semantics is in the works.
Summary
In a self-hosted environment (Managed Agents), the documented "sandbox per session" pattern — a dispatcher claims a work item via
ant beta:worker poll --on-work spawn.sh, then spawns a fresh container whose worker process is supposed to execute that already-claimed item — reliably fails. The spawned worker starts fine, reachesapi.anthropic.comfine, then shuts itself down within 2–6 seconds without executing anything (exit code 0). The session is left permanently stuck atstop_reason: requires_actionand the work item is never re-offered to any worker.We reproduced the identical failure with two independent client implementations:
antCLI (ant beta:worker run), versions 1.15.0 and 1.17.0EnvironmentWorker.handle_item()Both fail at the same point, in the same way, which suggests the issue is server-side (or an undocumented requirement of the claim-handoff contract) rather than a client defect. Filing here because the SDK repro is the cleanest minimal case, and we don't have a direct support line.
Environment
env_01UPQK3XYX67tnteLTvZw891(typeself_hosted)agent_014h5P54ckDkXpUG1NVdqY4Gsesn_011zBktK84XqgQapwZCSunbU,sesn_017jQMiZ3mSv2bbtPFAWtayn,sesn_01GSgDdFAJ9JPKgx8oSt1UA6,sesn_01BjrqoHYCysBHz4tFboyfDU,sesn_014f4Y2BQrjJFaHZavJCxxVHsesn_01GvoUeeRqmHe6dKbfPkCtcHReproduction steps
POST /v1/environments,config.type: self_hosted); generate its environment key via the Console.ant beta:worker poll --on-work ./spawn.sh, wherespawn.shfollows the documented shape — one fresh container per claimed work item, withANTHROPIC_SESSION_ID,ANTHROPIC_WORK_ID,ANTHROPIC_ENVIRONMENT_ID,ANTHROPIC_ENVIRONMENT_KEYinjected as env vars (our implementation spawns a Kubernetes Job rather thandocker run; functionally equivalent).POST /v1/agents) and a Session (POST /v1/sessionsreferencing agent + environment_id).user.messageevent, e.g. "List every file in /workspace and print their sizes."SDK variant of the spawned worker (fails identically to the CLI)
Observed behavior
The dispatcher polls, claims exactly one work item (
work_type=session), and spawns exactly one container per claim — this part works on every attempt.The spawned worker's logs show, consistently within 2–6 seconds of startup:
and the process exits 0. The SDK variant behaves the same (ran ~3s, exit 0, no error output).
Meanwhile, via the Sessions API, the conversation proceeds normally on the model side:
user.message→agent.thinking→agent.message→agent.tool_use(bash,evaluated_permission: allow), thenstop_reason: requires_actionreferencing that tool_use, andsession.status_idle.The
tool_useis never fulfilled. No second work item is ever claimed for the session. It is stuck permanently.Ruled out
session_idinstead ofwork_id; a heredoc quoting bug); failure persisted unchanged after both fixes.kubectl get pod -o jsonpath.api.anthropic.comconfirmed working from inside the spawned pod.--max-idletiming — tried 300s vs. the 60s default, plus explicit--environment-id; no change. Reverted to the exact zero-flag documented pattern for this report.Confirmed workaround (what narrowed it down)
The "in-process" pattern — a single long-running process that both claims and executes,
ant beta:worker poll --workdir /workspacewith no--on-work— works correctly on the first try: full tool-call cycle, results posted, cleanstop_reason: end_turn.So the failure is specifically in resuming an already-claimed work item in a process other than the one that claimed it. The observable shape (server heartbeat replying "shutdown" to the freshly spawned worker) is consistent with the claim/lease being bound to the claiming process's identity, so the handed-off worker is told to stop — but that's inference; we can't see the server side.
Question
Is there a known issue — or an undocumented handshake/requirement — for how a spawned
ant beta:worker run/EnvironmentWorker.handle_item()process adopts a work item claimed by a different (dispatcher) process via the injected env vars? The session IDs above should allow server-side correlation. We've temporarily rearchitected to the in-process pattern, but we'd like to return to sandbox-per-session isolation once this is resolved.