Repository navigation
fix(think, sandbox-coding-agent): model id validation, ai@7 peer, sandbox cleanup and hardening - #2388
Conversation
- think: require ai@^7 / @ai-sdk/react@^4 (bundled providers are v4 models) - think: resolveModel fails fast on malformed string model ids; tests for invalid ids, missing AI binding, LanguageModel pass-through, and string models from beforeTurn/beforeStep - think README: dynamic-config snippet no longer uses an unimported helper - sandbox-coding-agent: destroy containers when runs finish and on clear; sandbox id hashes orchestrator name + run id; missing exit / non-zero exit / failed git commands fail the run; validate Claude session ids; allowlist proxied Anthropic endpoints and cap max_tokens; README security section; align Dockerfile base image with SDK 0.12.2 and pin Claude Code 2.1.283 - rfc-coding-agent: fix dangling section refs, reconcile topology with the accepted user-chat DO RFC, cross-reference the Codex harness RFC, add open questions Co-authored-by: Cursor <cursoragent@cursor.com>
🦋 Changeset detectedLatest commit: ae2d3de The changes in this PR will be included in the next version bump. This PR includes changesets to release 2 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
|
✅ agents import sizes: no significant changes ( |
| override async onAgentToolFinish(run: AgentToolRunInfo): Promise<void> { | ||
| await this.destroySandbox(run.runId); |
There was a problem hiding this comment.
🔴 Interrupted containers outlive completed child runs
When an interrupted child finishes after releaseIdleSandboxes sees it running, its container receives no cleanup. onStart runs the sweep only on orchestrator wakes; child completion does not trigger another wake. Those containers can exhaust the five-instance limit before they sleep.
Learn more
The orchestrator tracks each container in storage. A finish callback can mark a run interrupted while its child continues, so the callback leaves the container alive. The sweep runs once when the parent wakes and skips any child it finds running. The child can finish after that sweep without sending another finish callback to the parent. Without another parent wake, the container stays allocated until the sandbox's 15-minute idle sleep; the example permits only five concurrent containers.
Example: A child is still running when its parent recovers and reports interrupted. The wake sweep sees running and retains its container. The child finishes one minute later, but the parent has no other requests for ten minutes. Further delegated tasks can reach max_instances: 5 despite the earlier children having finished.
Recommended fix: Schedule a durable follow-up cleanup for tracked interrupted runs, such as a periodic alarm or scheduled callback that re-inspects the child until it is terminal and calls destroySandbox. Avoid relying on an unrelated request to wake the orchestrator.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Valid — fixed in e163805. onAgentToolFinish now keeps the sandbox when the status is interrupted and childStillRunning !== false (unknown treated as maybe-running); it's freed later by the wake sweep, a later finish, or clearDelegatedRuns.
agents
@cloudflare/ai-chat
@cloudflare/codemode
hono-agents
@cloudflare/shell
@cloudflare/think
@cloudflare/voice
@cloudflare/worker-bundler
commit: |
- think: resolveModel accepts only @cf/ and @hf/ Workers AI ids with a model name; other @-prefixed ids fail with the invalid-id error. @hf/ ids get the same Workers AI settings as @cf/ ids - sandbox-coding-agent: keep the container of an interrupted run whose child may still be working; sweep tracked containers on wake and destroy those whose child run is no longer running (covers failed starts) Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # packages/think/src/think.ts
Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # packages/think/src/tests/hooks.test.ts # packages/think/src/think.ts
| if (failure) { | ||
| const tail = stderr.trim().split("\n").slice(-12).join("\n"); | ||
| writer.write({ type: "text-start", id: "error" }); | ||
| writer.write({ | ||
| type: "text-delta", | ||
| id: "error", | ||
| delta: | ||
| `\n\n**Claude Code error**\n\n${detail}` + | ||
| `\n\n**Claude Code error**\n\n${failure}` + | ||
| (tail ? `\n\n\`\`\`\n${tail}\n\`\`\`` : "") | ||
| }); | ||
| writer.write({ type: "text-end", id: "error" }); |
There was a problem hiding this comment.
🔴 Recovered Claude failures become successful runs
If a child restarts after writing an error message but before recording its outcome, inspectAgentToolRun treats that assistant message as success. The parent reports the failed coding task as completed.
Learn more
A Claude failure writes human-readable error text into the assistant stream before throwing. If the child Durable Object restarts between persisting that text and finalizing the agent-tool row, reconciliation treats any assistant message after the run started as proof of success. This error message satisfies that condition, so inspectAgentToolRun reports completed even though Claude failed.
Example: Claude exits with code 1 and streams the error text; the child is evicted before startAgentToolRun stores error. On recovery, the error assistant message causes a completed status, and the orchestrator reports the task as done.
Recommended fix: Persist a durable failure marker the child reconciliation can read, or update AIChatAgent reconciliation to respect a persisted stream-error outcome before inferring completion from assistant messages. Test a restart between the error chunk and terminal row update.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Accurate, but it's an SDK-level reconciliation gap rather than an issue in this example: AIChatAgent's stale-run reconciliation (and Think's mirror) infers completed from any assistant message after the run started, so an eviction between persisting the error text and finalizing the row reads as success. The right fix is in the reconciliation itself (respect a durable failure/stream-error marker before inferring completion), which affects every agent tool, so I'm tracking it as a separate follow-up rather than widening this PR.
There was a problem hiding this comment.
Fixed in #2390: the child now records its stream error on the open run row as it happens, and stale-row reconciliation seals such a row as error instead of inferring completed from the assistant reply (both AIChatAgent and Think).
…ire a bounded max_tokens Co-authored-by: Cursor <cursoragent@cursor.com>
| if (result.status === "error" && (await this.isChildLive(run.runId))) { | ||
| return; |
There was a problem hiding this comment.
🔴 Finished child sandboxes exhaust container slots
When a child stream fails, onAgentToolFinish retains the live child's sandbox without arranging cleanup after it finishes. releaseIdleSandboxes runs only on orchestrator startup, so an active orchestrator leaves these sandboxes allocated. Repeated failures can exhaust the five-container limit and block new tasks.
Learn more
The parent tracks a sandbox when a run starts, and its finish hook normally destroys that sandbox. A child-stream failure can mark the parent run error while the child still reports running. This branch retains the sandbox, but releaseIdleSandboxes runs only in onStart. An orchestrator that stays active does not run that sweep again when the child finishes. Its retained sandbox continues occupying a container slot until the container sleeps or the orchestrator restarts.
Example: Five child streams fail while their children keep editing. Each child finishes later, but the same orchestrator remains active. All five sandboxes remain allocated, so a sixth task cannot start within the configured max_instances: 5 limit.
Recommended fix: Schedule a durable follow-up inspection for retained runs, or arrange a child-completion callback that destroys the sandbox once inspectAgentToolRun reports a terminal state. Ensure the follow-up survives orchestrator eviction and retries transient inspection failures.
Was this helpful? React with 👍 or 👎 to provide feedback.
|
|
||
| - **`/agents/*` has no authentication.** Anyone who can reach the Worker can chat with any orchestrator by name, start containers, and spend your Workers AI and AI Gateway budget. Put it behind [Cloudflare Access](https://developers.cloudflare.com/cloudflare-one/policies/access/) or add an auth check before `routeAgentRequest` before deploying. | ||
| - **The container runs model-directed commands with internet access.** Claude Code runs with `--permission-mode bypassPermissions`, so it executes whatever shell commands it decides on, without approval, and the container can reach the public internet (`enableInternet = true`, needed for the `git clone`). Don't put anything in the container you wouldn't hand to an untrusted process. | ||
| - **The Anthropic proxy is bounded, not locked down.** It only forwards `POST /v1/messages` and `POST /v1/messages/count_tokens` and rejects message requests whose `max_tokens` is missing or above `MAX_OUTPUT_TOKENS`, but any process in the container can still call those endpoints on your account. Set spend limits / rate limits on the gateway. |
There was a problem hiding this comment.
🔍 Cleanup example differs from the implementation
The README's finish-hook example destroys a sandbox after an error, even when its child is still editing. The implementation keeps that sandbox. Update the example so readers do not copy the conflicting cleanup behavior.
Was this helpful? React with 👍 or 👎 to provide feedback.
Follow-ups from a review of #1832 (bundled workers-ai-provider), #1830 (sandbox-coding-agent example) and #1831 (CodingAgent RFC).
@cloudflare/think
ai@^7(and@ai-sdk/react@^4). The bundledworkers-ai-provider@4/@ai-sdk/*@4produce v4 models, whichai@6rejects withUnsupportedModelVersionError, so string model ids already failed onai@6. Shipped as a patch; consider whether you want a minor.resolveModelfails fast on ids that are neither@...nor<provider>/<model>(e.g.gpt-5), instead of failing at inference time inenv.AI.run.LanguageModelpassthrough, string models frombeforeTurn/beforeStep.createWorkersAI.examples/sandbox-coding-agent
clearDelegatedRuns(previously the 6th task within 15 minutes exceededmax_instances: 5).getWorkspaceDiffis replaced bygetLastResult, since the container is gone after a run.POST v1/messagesandv1/messages/count_tokens, and rejectsmax_tokens > 32000. README gains a Security section (no auth on/agents/*, internet +bypassPermissions, gateway spend).0.12.2); Claude Code pinned to2.1.283.design/rfc-coding-agent.md
§10references; notesrfc-think-multi-session.mdis rejected and aligns topology with the acceptedrfc-user-chat-durable-objects.md; cross-referencesrfc-codex-harness-capability.md; adds open questions (container quotas/cleanup, egress cost/abuse, auth/multi-tenancy, concurrent turns, container-death contract).Verified on a combined branch with the other review follow-ups:
pnpm run check, agents chat/workers/react, ai-chat, and Think suites all pass.