On a self-hosted environment, the outcome grader's tool calls are executed by the worker, but posting their results back is rejected with a 400. The grader then waits indefinitely and outcome_evaluations[].result stays pending forever. The same agent and the same rubric grade normally on a type: cloud environment.
The rejection is self-inconsistent. GET /v1/sessions/{session_id}/events returns the grader's agent.tool_use events. POST /v1/sessions/{session_id}/events with a user.tool_result for one of those same event ids returns:
`tool_use_id` "sevt_01C7v5e8VgTT71gnLSYkUnn3" does not match any `agent.tool_use` event in this session
A writer-thread tool result posted to that identical endpoint seven seconds earlier succeeds with a 200. The only difference I can find between the two is that the grader's events carry a session_thread_id and the writer's do not.
Environment
anthropic==0.120.0 (Python), running EnvironmentWorker in a long-lived Kubernetes deployment
- Self-hosted environment,
agent_toolset_20260401, permission policy always_allow
- Agent model
claude-opus-5, effort low, max_iterations: 1
src/anthropic/lib/tools/_beta_session_runner.py and src/anthropic/lib/environments/ are byte-identical across v0.117.0, v0.120.0, and v0.120.2, so this is not specific to the version we run.
Reproduction
Create a session on a self-hosted environment with a single user.define_outcome initial event whose rubric requires the grader to read the deliverable back from the sandbox:
{
"type": "user.define_outcome",
"description": "Create a file /workspace/probe.txt whose only line is exactly OUTCOME-PROBE-OK. Then stop.",
"rubric": {
"type": "text",
"content": "## File exists and is correct\nThe file /workspace/probe.txt exists and its only line is exactly `OUTCOME-PROBE-OK`. You MUST verify this by reading the file from the sandbox filesystem with a tool call."
},
"max_iterations": 1
}
The writer completes and writes the file. The grader starts, issues two tool calls, and the evaluation never finishes.
Worker log
One session, one worker, no stream reconnect between these lines:
19:11:35 _poller INFO claimed work work_id=sesn_01JeEaW44N8VQs3r23sHqaAG work_type=session
19:11:36 _beta_session_runner INFO session tool runner starting session_id=sesn_01JeEaW44N8VQs3r23sHqaAG
19:11:37 httpx GET /v1/sessions/sesn_01JeEaW44N8VQs3r23sHqaAG/events/stream "200 OK"
19:11:37 httpx GET /v1/sessions/sesn_01JeEaW44N8VQs3r23sHqaAG/events?limit=1000 "200 OK"
19:11:37 _beta_session_runner INFO executing tool tool=write tool_use_id=sevt_013KfLN9Vgqh4wcLJPA4oU8W
19:11:37 httpx POST /v1/sessions/sesn_01JeEaW44N8VQs3r23sHqaAG/events "200 OK"
19:11:44 _beta_session_runner INFO executing tool tool=bash tool_use_id=sevt_01C7v5e8VgTT71gnLSYkUnn3
19:11:44 httpx POST /v1/sessions/sesn_01JeEaW44N8VQs3r23sHqaAG/events "400 Bad Request"
19:11:44 _beta_session_runner ERROR failed to send tool result tool_use_id=sevt_01C7v5e8VgTT71gnLSYkUnn3
error=Error code: 400 - {'type': 'error', 'error': {'type': 'invalid_request_error',
'message': '`tool_use_id` "sevt_01C7v5e8VgTT71gnLSYkUnn3" does not match any
`agent.tool_use` event in this session'}, 'request_id': 'req_011CdYmkQdC493iPnk2tkFPZ'}
19:11:44 _beta_session_runner INFO executing tool tool=read tool_use_id=sevt_015ZVxsWuDgG2ywnmpTwBa8p
19:11:44 httpx POST /v1/sessions/sesn_01JeEaW44N8VQs3r23sHqaAG/events "400 Bad Request"
19:11:44 _beta_session_runner ERROR failed to send tool result tool_use_id=sevt_015ZVxsWuDgG2ywnmpTwBa8p
error=Error code: 400 - ... same message, request_id: req_011CdYmkRzYFj4URAB2hY44N
19:12:08 onward: work-item heartbeat every ~30s, indefinitely
The first POST is the writer's write result and succeeds. The second and third are the grader's bash and read results and are rejected.
The rejected events are present in the session
GET /v1/sessions/sesn_01JeEaW44N8VQs3r23sHqaAG/events returns both of them:
{"evaluated_permission": "allow", "id": "sevt_01C7v5e8VgTT71gnLSYkUnn3", "input": {"command": "ls -la /workspace/ 2>&1 | head -50"}, "name": "bash", "processed_at": "2026-07-30T19:11:44.005467Z", "session_thread_id": "sthr_01Xc5CCvKXaeFNReSQFyiexs", "type": "agent.tool_use"}
{"evaluated_permission": "allow", "id": "sevt_015ZVxsWuDgG2ywnmpTwBa8p", "input": {"file_path": "/workspace/probe.txt"}, "name": "read", "processed_at": "2026-07-30T19:11:44.105355Z", "session_thread_id": "sthr_01Xc5CCvKXaeFNReSQFyiexs", "type": "agent.tool_use"}
The session's outcome, two hours later:
{"outcome_id": "outc_01BScKAoJTMYrrc1wF6jpcer", "result": "pending", "iteration": 0, "completed_at": null, "explanation": null}
Across four self-hosted runs the arithmetic is identical every time: total agent.tool_use minus total user.tool_result equals exactly 2, and the two unanswered calls are always the grader's.
The grader's thread is not addressable
The session_thread_id on those events is sthr_01Xc5CCvKXaeFNReSQFyiexs. GET /v1/sessions/{id}/threads does not return it — it lists only the writer thread, sthr_01JoM7eNWheWGhMvV5QFQfZd. Retrieving it directly 404s:
GET /v1/sessions/sesn_01JeEaW44N8VQs3r23sHqaAG/threads/sthr_01Xc5CCvKXaeFNReSQFyiexs
404 {"type": "not_found_error", "message": "Thread not found: sthr_01Xc5CCvKXaeFNReSQFyiexs"}
So the grader's thread is observable through session events but not through the threads API.
Why there is no client-side workaround
If the 400 means that a result for a grader-thread tool call must be scoped to that thread, the SDK offers no way to express it:
BetaManagedAgentsUserToolResultEventParams carries tool_use_id, content, and is_error. There is no session_thread_id. Same for the custom-tool variant.
resources/beta/sessions/threads/events.py exposes list and stream only. There is no send.
SessionToolRunner never references session_thread_id. It reads events.list(session_id) and events.stream(session_id) and posts to events.send(session_id).
The worker has no other move available to it.
Secondary: the grader's tool calls are re-executed indefinitely
Because the POST fails, no user.tool_result event is ever recorded, so _answered never contains the id. On every stream reconnect, _reconcile re-lists history, finds the call still unanswered, and executes it again. On a worker that stayed up for several hours I watched the same grader tool_use_ids re-run on each reconnect, each followed by the same 400.
For a grader that only inspects state this is wasted work, but the tools are general-purpose (bash, write, edit) and the re-execution is unbounded.
SEND_RETRIES = 3 applies within a single attempt, after which the runner logs at ERROR and abandons the call. Nothing about the failure appears in the session event stream, so from the API side a permanently broken run is indistinguishable from a slow grader.
Ruled out
To save you the first few replies, all of these were tested and are not the cause:
- Rubric length. Tightened from 1,388 to 745 words; stalled identically. The 745-word version is the one that graded cleanly on cloud.
- Criterion count. Seven criteria, each decomposed and scored correctly on cloud.
- Deliverable location. The rubric targets
/workspace, which matches the self-hosted guidance that the /mnt/session/outputs instruction is omitted and deliverables land under the working directory.
- Confirmation gating. Every
agent.tool_use event in these sessions carries evaluated_permission: "allow", so _route_tool_event is not holding them awaiting a user.tool_confirmation.
- Worker health. The run above is on a worker restarted immediately beforehand, polling cleanly, serving only this session.
Identifiers
| Role |
Session |
Environment |
| Clean repro, minimal rubric |
sesn_01JeEaW44N8VQs3r23sHqaAG |
env_01JPL6GM5GxUXcfLkyji11sb (self-hosted) |
| Stall, seven-criterion rubric |
sesn_01Bc96NBxTeHwhWS6hok5YDL |
env_01PCEAkT6FTZgo97u3zvEQyS (self-hosted) |
| Stall, seven-criterion rubric |
sesn_01SDHETG8r6ckbqpmU2k1r3j |
env_01DVvdTG8qCfNBMRQ7CRq7UR (self-hosted) |
| Stall, tightened rubric |
sesn_01MKrw6ELdwGvNP64W5s6kJL |
env_01JPL6GM5GxUXcfLkyji11sb (self-hosted) |
| Control, graded successfully |
sesn_01SRfYcwC4GsX4bathVPAThh |
env_01F4KWpuuBJpqPJwb4ysH1pD (type: cloud) |
The control environment is retained in case you want its trace.
What would help
Confirm whether a tool result for a grader-thread tool call is meant to be posted session-scoped. If it is, the 400 looks like a server-side bug in how tool_use_id is resolved for non-writer threads. If it is not, then the SDK is missing the ability to post a thread-scoped result, and a self-hosted worker cannot service an outcome grader at all.
For what it's worth, nothing under examples/ or tests/ in this repository mentions outcomes, and no example combines a self-hosted environment with an outcome. That is why I suspect this combination is untested rather than deliberately unsupported. Filing here rather than elsewhere following the precedent of #1779.
🤖 Co-authored with Claude Code
On a self-hosted environment, the outcome grader's tool calls are executed by the worker, but posting their results back is rejected with a 400. The grader then waits indefinitely and
outcome_evaluations[].resultstayspendingforever. The same agent and the same rubric grade normally on atype: cloudenvironment.The rejection is self-inconsistent.
GET /v1/sessions/{session_id}/eventsreturns the grader'sagent.tool_useevents.POST /v1/sessions/{session_id}/eventswith auser.tool_resultfor one of those same event ids returns:A writer-thread tool result posted to that identical endpoint seven seconds earlier succeeds with a 200. The only difference I can find between the two is that the grader's events carry a
session_thread_idand the writer's do not.Environment
anthropic==0.120.0(Python), runningEnvironmentWorkerin a long-lived Kubernetes deploymentagent_toolset_20260401, permission policyalways_allowclaude-opus-5, effortlow,max_iterations: 1src/anthropic/lib/tools/_beta_session_runner.pyandsrc/anthropic/lib/environments/are byte-identical acrossv0.117.0,v0.120.0, andv0.120.2, so this is not specific to the version we run.Reproduction
Create a session on a self-hosted environment with a single
user.define_outcomeinitial event whose rubric requires the grader to read the deliverable back from the sandbox:{ "type": "user.define_outcome", "description": "Create a file /workspace/probe.txt whose only line is exactly OUTCOME-PROBE-OK. Then stop.", "rubric": { "type": "text", "content": "## File exists and is correct\nThe file /workspace/probe.txt exists and its only line is exactly `OUTCOME-PROBE-OK`. You MUST verify this by reading the file from the sandbox filesystem with a tool call." }, "max_iterations": 1 }The writer completes and writes the file. The grader starts, issues two tool calls, and the evaluation never finishes.
Worker log
One session, one worker, no stream reconnect between these lines:
The first POST is the writer's
writeresult and succeeds. The second and third are the grader'sbashandreadresults and are rejected.The rejected events are present in the session
GET /v1/sessions/sesn_01JeEaW44N8VQs3r23sHqaAG/eventsreturns both of them:{"evaluated_permission": "allow", "id": "sevt_01C7v5e8VgTT71gnLSYkUnn3", "input": {"command": "ls -la /workspace/ 2>&1 | head -50"}, "name": "bash", "processed_at": "2026-07-30T19:11:44.005467Z", "session_thread_id": "sthr_01Xc5CCvKXaeFNReSQFyiexs", "type": "agent.tool_use"} {"evaluated_permission": "allow", "id": "sevt_015ZVxsWuDgG2ywnmpTwBa8p", "input": {"file_path": "/workspace/probe.txt"}, "name": "read", "processed_at": "2026-07-30T19:11:44.105355Z", "session_thread_id": "sthr_01Xc5CCvKXaeFNReSQFyiexs", "type": "agent.tool_use"}The session's outcome, two hours later:
{"outcome_id": "outc_01BScKAoJTMYrrc1wF6jpcer", "result": "pending", "iteration": 0, "completed_at": null, "explanation": null}Across four self-hosted runs the arithmetic is identical every time: total
agent.tool_useminus totaluser.tool_resultequals exactly 2, and the two unanswered calls are always the grader's.The grader's thread is not addressable
The
session_thread_idon those events issthr_01Xc5CCvKXaeFNReSQFyiexs.GET /v1/sessions/{id}/threadsdoes not return it — it lists only the writer thread,sthr_01JoM7eNWheWGhMvV5QFQfZd. Retrieving it directly 404s:So the grader's thread is observable through session events but not through the threads API.
Why there is no client-side workaround
If the 400 means that a result for a grader-thread tool call must be scoped to that thread, the SDK offers no way to express it:
BetaManagedAgentsUserToolResultEventParamscarriestool_use_id,content, andis_error. There is nosession_thread_id. Same for the custom-tool variant.resources/beta/sessions/threads/events.pyexposeslistandstreamonly. There is nosend.SessionToolRunnernever referencessession_thread_id. It readsevents.list(session_id)andevents.stream(session_id)and posts toevents.send(session_id).The worker has no other move available to it.
Secondary: the grader's tool calls are re-executed indefinitely
Because the POST fails, no
user.tool_resultevent is ever recorded, so_answerednever contains the id. On every stream reconnect,_reconcilere-lists history, finds the call still unanswered, and executes it again. On a worker that stayed up for several hours I watched the same gradertool_use_ids re-run on each reconnect, each followed by the same 400.For a grader that only inspects state this is wasted work, but the tools are general-purpose (
bash,write,edit) and the re-execution is unbounded.SEND_RETRIES = 3applies within a single attempt, after which the runner logs at ERROR and abandons the call. Nothing about the failure appears in the session event stream, so from the API side a permanently broken run is indistinguishable from a slow grader.Ruled out
To save you the first few replies, all of these were tested and are not the cause:
/workspace, which matches the self-hosted guidance that the/mnt/session/outputsinstruction is omitted and deliverables land under the working directory.agent.tool_useevent in these sessions carriesevaluated_permission: "allow", so_route_tool_eventis not holding them awaiting auser.tool_confirmation.Identifiers
sesn_01JeEaW44N8VQs3r23sHqaAGenv_01JPL6GM5GxUXcfLkyji11sb(self-hosted)sesn_01Bc96NBxTeHwhWS6hok5YDLenv_01PCEAkT6FTZgo97u3zvEQyS(self-hosted)sesn_01SDHETG8r6ckbqpmU2k1r3jenv_01DVvdTG8qCfNBMRQ7CRq7UR(self-hosted)sesn_01MKrw6ELdwGvNP64W5s6kJLenv_01JPL6GM5GxUXcfLkyji11sb(self-hosted)sesn_01SRfYcwC4GsX4bathVPAThhenv_01F4KWpuuBJpqPJwb4ysH1pD(type: cloud)The control environment is retained in case you want its trace.
What would help
Confirm whether a tool result for a grader-thread tool call is meant to be posted session-scoped. If it is, the 400 looks like a server-side bug in how
tool_use_idis resolved for non-writer threads. If it is not, then the SDK is missing the ability to post a thread-scoped result, and a self-hosted worker cannot service an outcome grader at all.For what it's worth, nothing under
examples/ortests/in this repository mentions outcomes, and no example combines a self-hosted environment with an outcome. That is why I suspect this combination is untested rather than deliberately unsupported. Filing here rather than elsewhere following the precedent of #1779.🤖 Co-authored with Claude Code