Repository navigation
Coalesce high-frequency assistant streaming deltas - #4323
colonelpanic8 wants to merge 21 commits into
Conversation
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
ApprovabilityVerdict: Needs human review This PR introduces significant new runtime behavior for coalescing high-frequency streaming assistant deltas, including new state caches, retry/deferral mechanisms, and multiple new processing paths. The scope and complexity of these orchestration layer changes warrant human review. You can customize Macroscope's approvability policy. Learn more. |
415a93a to
591b639
Compare
591b639 to
87bae1d
Compare
87bae1d to
fc188be
Compare
fc188be to
f8855ac
Compare
6519c5a to
983609d
Compare
fdc405d to
032a352
Compare
d3d98b4 to
9cec4c5
Compare
c45471d to
a38265c
Compare
The terminal fallback for an exhausted assistant finalization sent the undispatched buffer as the completion's message text. A non-streaming thread.message-sent replaces the projected text in every projector (read model, persisted projection, and client reducer), so any delta that had already been streamed and persisted was dropped and the message collapsed to its trailing suffix. Only undispatched text is ever buffered, so the completion command now carries it as appendText: the decider emits it as one more streaming delta before settling the message with an empty text, which the projectors already treat as "keep the existing text". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
a38265c to
770d450
Compare
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 770d450. Configure here.
…eaming deltas)
…eaming deltas)
|
Closing in favor of #2829 (orchestration V2). #2829 deletes the V1 orchestration layer this PR builds on — This is not a judgement on the change itself. Several of these are real gaps we still want fixed; the base just moved out from under them. Once #2829 merges, please rebase onto |

What Changed
Why
Providers can emit assistant text one token or character at a time. Persisting and broadcasting every tiny delta creates thousands of durable events, repeatedly rebuilds the client message, and forces the renderer to reconcile and re-render Markdown at token frequency. In long-running sessions this can drive the Electron renderer past its practical memory limit even while the backend remains healthy.
The bounded coalescing keeps the streaming UI responsive while reducing persistence, replay, and renderer work by roughly two orders of magnitude for high-frequency streams. Completion and pause paths still flush all pending text, so no content is lost.
UI Changes
No visual changes. Assistant text continues to stream, with updates grouped into intervals of at most 100 ms under normal timestamped provider traffic.
Checklist
pnpm --config.enableGlobalVirtualStore=false exec vp test run apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.test.tspnpm --config.enableGlobalVirtualStore=false exec vp checkpnpm --config.enableGlobalVirtualStore=false exec vp run typecheckNote
High Risk
Changes the core provider-to-orchestration message pipeline, event ordering, and completion semantics; regressions could lose assistant text, duplicate completions, or stall threads on stuck retries.
Overview
Provider runtime ingestion now batches high-frequency assistant text when streaming is enabled: pending text flushes on a 100ms interval or when 512 characters accumulate, routed through dedicated worker inputs (
assistant-delta,assistant-flush,assistant-finalize) instead of dispatching every token immediately.Failed
thread.message.assistant.delta/completedispatches retain buffered text and retry with bounded backoff; turn completion, session exit, runtime errors, and approval pauses wait until deferred assistant work drains so lifecycle events do not overtake partial messages. Sessionthread.session.setfor terminal boundaries can be deferred until assistant finalization finishes.thread.message.assistant.completegains optionalappendText; the decider emits a final streaming delta for undispatched text, then an empty-text completion so projectors do not replace already-streamed content. Terminal fallback completion after exhausted retries uses the same path.finalizeTrackedAssistantMessagesis exported for partial turn finalization (release segment state only for messages that actually completed). Integration tests cover coalescing event counts, dispatch failures, timer flush, ordering vs plans/requests, and deduplication.Reviewed by Cursor Bugbot for commit 8638e03. Bugbot is set up for automated code reviews on this repo. Configure here.
Note
Coalesce high-frequency assistant streaming deltas with retry and deferred finalization
STREAMING_ASSISTANT_FLUSH_INTERVAL_MILLIS) and size threshold (MAX_COALESCED_STREAMING_ASSISTANT_CHARS) in ProviderRuntimeIngestion.ts to reduce high-frequency tiny delta updates.ThreadMessageAssistantCompleteCommandgains an optionalappendTextfield in orchestration.ts; decider.ts emits a preceding streaming delta event when this field is set, enabling best-effort fallback completion.Macroscope summarized 8638e03.