Repository navigation
fix(agents): agent-tool replay sequencing, re-attach, and fiber recovery - #2384
Conversation
…ery review fixes Co-authored-by: Cursor <cursoragent@cursor.com>
🦋 Changeset detectedLatest commit: 2ecbad8 The changes in this PR will be included in the next version bump. This PR includes changesets to release 4 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🟡 agents import sizes: 1 entry point grew
Changed exports (2)
How this worksEach runtime export is bundled on its own, minified, and gzipped. Changes smaller than 100 B, or smaller than 1% and 1 KiB, are ignored. Growth over 10% or 5 KiB is marked 🔴. This report is informational and does not fail CI. The workflow artifact contains every measurement. Compared |
agents
@cloudflare/ai-chat
@cloudflare/codemode
hono-agents
@cloudflare/shell
@cloudflare/think
@cloudflare/voice
@cloudflare/worker-bundler
commit: |
…on, stop progress frames shifting child tail numbering - Record the fiber body's outcome (completed/error/aborted + message) on the run row with its settlement stamp, and settle a managed ledger whose own settle write failed with that outcome instead of always 'completed'. Schema v14 adds cf_agents_runs.outcome / error_message. - Connect-time replay now bounds child resolution too, sharing one per-run budget across resolve, chunk read, and milestone inspection. - The shared broadcast interceptor no longer advances the child's live counter for progress/milestone frames, and ai-chat/Think tails emit those frames outside the stored-position high-water dedupe, so a chunk stored and broadcast during a tail's drain is not duplicated and progress is not dropped. Co-authored-by: Cursor <cursoragent@cursor.com>
…e-attach, and repeated progress - A chunk too large to store is broadcast live but never replayed. The child's broadcast snoop now gives it the next stored position without consuming it plus a unique `unstoredId`; child tails forward it outside the stored-position dedupe, the parent does not advance its sequence for it, and `agentToolEventDedupeKey` keys it on the id. Live and replayed numbering stay aligned after a skipped chunk. - ai-chat and Think tails seed a cold live counter from the stored backlog before registering the forwarder, instead of realigning after the drain / inspection, so a chunk broadcast during those awaits after a restart is forwarded rather than dropped. - Each `reportProgress` frame carries a unique `id` that the dedupe key uses, so identical progress payloads are no longer dropped by `useAgentToolEvents`. Co-authored-by: Cursor <cursoragent@cursor.com>
| keeps working. A client that drills in to the child sees the stored | ||
| conversation when it connects and the final messages when the run ends, but | ||
| does not see text stream live while the run is in flight. Use the default | ||
| `eventDelivery` for runs you expect users to watch. Detached runs do not |
There was a problem hiding this comment.
Checked: cloudflare-docs doesn't document eventDelivery / terminal delivery yet (no matches in src/content/docs/agents/), so there's no published section to keep in sync. This paragraph will go upstream with the rest of that section when it's ported; no change in this PR.
…ery-review Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # packages/agents/src/chat/resumable-stream.ts
| // | ||
| // Seed a cold live counter first, so a chunk broadcast while this | ||
| // tail drains or inspects continues the stored numbering. | ||
| const seeded = this._seedAgentToolLiveSequence(runId); |
There was a problem hiding this comment.
🟡 Finished child retains live broadcast counter
If a recovered child finishes during tail drainage, _seedAgentToolLiveSequence leaves its counter behind. The early return bypasses counter cleanup, and _closeAgentToolTailers does not remove it. Later broadcasts keep parsing frames and looking up run ownership.
Learn more
A child tail registers a live counter before awaiting its stored backlog. Terminalization can close that tail during the backlog read, causing the if (closed) return branch to skip the later terminal inspection and its counter deletion. The normal child finalizer clears the counter, but a recovered turn never reaches that finalizer; _closeAgentToolTailers currently leaves it set. Every future broadcast then takes the snoop path and can query SQLite for attribution.
Example: A recovered child has three stored chunks. Its parent starts a tail, then the child completes while getAgentToolChunks awaits. The drain returns early and the run remains in _agentToolLiveSequences after completion.
Recommended fix: Clear _agentToolLiveSequences when _closeAgentToolTailers finalizes a recovered run, or ensure every early tail exit clears a counter it seeded without erasing another active tail's counter.
Was this helpful? React with 👍 or 👎 to provide feedback.
| const { live, replay } = await agent.captureLiveAndReplayForTest({ | ||
| runId, | ||
| chunkBodies: [textStart, textDelta("A"), textDelta("B"), textDelta("C")], | ||
| unstoredChunks: [{ beforeChunk: 2, body: textDelta("X") }] |
There was a problem hiding this comment.
🔍 Reconnect test simulates rather than stores oversized chunks
The reconnect fixture marks a small chunk unstored without exercising the storage limit. Host tests cover oversized forwarding, but reconnect coverage does not test the actual storage-skip boundary.
Was this helpful? React with 👍 or 👎 to provide feedback.
## pnpm-workspace.yaml (default) ## Dependency Updates | Package | From | To | Type | | --- | --- | --- | --- | | `agents` | 0.24.0 | 0.25.0 | minor | ## Release Notes <details> <summary><b>agents</b> (0.24.0 → 0.25.0)</summary> ### Minor Changes - [#2005](cloudflare/agents#2005) [`c2f7672`](cloudflare/agents@c2f7672) Thanks [@<!---->cjol](https://github.com/cjol)! - Native RPC calls to async Agent and Think methods now start lifecycle initialization first; address Agents by name because raw IDs from `newUniqueId()` and `idFromString()` now fail their first async RPC. See [Lifecycle](https://github.com/cloudflare/agents/blob/main/docs/agents/lifecycle.md). ### Patch Changes - [#2390](cloudflare/agents#2390) [`c55ec80`](cloudflare/agents@c55ec80) Thanks [@<!---->threepointone](https://github.com/threepointone)! - Report a failed agent-tool child as failed even when it was evicted before recording the failure. See [Agent tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md). - [#2384](cloudflare/agents#2384) [`f904999`](cloudflare/agents@f904999) Thanks [@<!---->threepointone](https://github.com/threepointone)! - Fix agent-tool chunks being duplicated or dropped on reconnect, child re-attach, and fiber recovery. See [Agent tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md). - [#2364](cloudflare/agents#2364) [`5e0507e`](cloudflare/agents@5e0507e) Thanks [@<!---->threepointone](https://github.com/threepointone)! - Add `eventDelivery: "terminal"` to `runAgentTool` to forward only lifecycle, progress, and milestone events for a run. See [Agent tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md). - [#2448](cloudflare/agents#2448) [`54f9ca7`](cloudflare/agents@54f9ca7) Thanks [@<!---->aron-cf](https://github.com/aron-cf)! …[full notes](https://github.com/cloudflare/agents/releases/tag/agents%400.25.0) </details> --- *This PR was auto-generated by [catalog-update-action](https://github.com/brandhaug/catalog-update-action).* Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
## pnpm-workspace.yaml (default) ## Dependency Updates | Package | From | To | Type | | --- | --- | --- | --- | | `@cloudflare/ai-chat` | 0.12.0 | 0.12.1 | patch | ## Release Notes <details> <summary><b>@<!---->cloudflare/ai-chat</b> (0.12.0 → 0.12.1)</summary> ### Patch Changes - [#2390](cloudflare/agents#2390) [`c55ec80`](cloudflare/agents@c55ec80) Thanks [@<!---->threepointone](https://github.com/threepointone)! - Report a failed agent-tool child as failed even when it was evicted before recording the failure. See [Agent tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md). - [#2384](cloudflare/agents#2384) [`f904999`](cloudflare/agents@f904999) Thanks [@<!---->threepointone](https://github.com/threepointone)! - Fix agent-tool chunks being duplicated or dropped on reconnect, child re-attach, and fiber recovery. See [Agent tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md). - [#2364](cloudflare/agents#2364) [`5e0507e`](cloudflare/agents@5e0507e) Thanks [@<!---->threepointone](https://github.com/threepointone)! - Add `eventDelivery: "terminal"` to `runAgentTool` to forward only lifecycle, progress, and milestone events for a run. See [Agent tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md). - [#2334](cloudflare/agents#2334) [`7f564e7`](cloudflare/agents@7f564e7) Thanks [@<!---->threepointone](https://github.com/threepointone)! - AI Chat and Think send the terminal `done` frame after persisting and broadcasting the assistant reply, so later sends are not overwritten. See [Chat agents](https://github.com/cloudflare/agents/blob/main/docs/agents/chat-agents.md). - [#2352](cloudflare/agents#2352) [`449ac27`](cloudflare/agents@449ac27) Thanks [@<!---->threepointone](https://github.com/threepointone)! - Do not start an automatic continuation after an …[full notes](https://github.com/cloudflare/agents/releases/tag/%40cloudflare/ai-chat%400.12.1) </details> --- *This PR was auto-generated by [catalog-update-action](https://github.com/brandhaug/catalog-update-action).* Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Ports three upstream changes to the legacy AIChatAgent's agent-tool child adapter onto AGUIChatAgent: - cloudflare#2364: opt-in terminal-only event delivery. Runs started with eventDelivery: "terminal" are tracked (and persisted on the run row) so their chunks are stored but not broadcast, including after a restart. - cloudflare#2384: seed a cold live sequence from the stored backlog before a tail attaches (replacing the post-drain realign), and inspectAgentToolRun(runId, { reconcile: false }). - cloudflare#2390: record a run's stream error on its open row so a stale-row reconcile after eviction reports error. The engine sends no error frame for an in-band RUN_ERROR, so that path records it directly. The test worker's unported() casts for _agentToolTerminalOnlyRuns and the inspectAgentToolRun options are replaced with the real members, and its finalize gate moves to _saveAGUIMessages, which is what the engine's agent-tool lifecycle calls.
Follow-ups from a review of recently landed agent-tool and fiber PRs (#2364, #2363, #1836, #1827, #2362).
ABbecameABBand chunks streamed while offline were lost. Stored chunks are now numbered by stored position on both paths, and progress/milestone frames don't consume a sequence.afterSequence, which production passes as the last stored index; the only test passed-1.error) a stale child run on every client connect.inspectAgentToolRunaccepts{ reconcile: false }.fiber:recovery:detected/fiber:run:interrupted.completed_at IS NULL); the agent-tool child-run table is migrated before rebinding a recovered turn (previously could fail with "no such column" on first wake after upgrade); terminal-only run entries are cleared when a recovered turn settles, in both Think and ai-chat.Verified on a combined branch with the other review follow-ups:
pnpm run check, agents chat/workers/react, ai-chat, and Think suites all pass.Notes for review
docs/agents/agent-tools.mdnow says drill-in connections see the stored conversation and final messages but no live streaming. Streaming to them would remove the child-side half ofeventDelivery: "terminal", since they're the child's only broadcast audience, so that's left as a product decision.durable-execution.mdnow saysonFiberRecoveredis best-effort (it can still run if recording the finish fails), so hooks should be idempotent.completed/error/aborted) and error message alongsidecompleted_at, and recovery settles the managed ledger with it (schema version 14). A row withcompleted_atbut no recorded outcome goes through normal interrupted recovery.