Repository navigation
fix(think): keep the prompt-cache prefix stable under truncation (#2200) - #2356
Merged
threepointone merged 2 commits intoSep 24, 2026
Merged
threepointone merged 2 commits into
threepointone merged 2 commits into
Conversation
threepointone
added this pull request to stack #2357
September 24, 2026 01:08
🦋 Changeset detectedLatest commit: 33b1e73 The changes in this PR will be included in the next version bump. This PR includes changesets to release 2 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
Contributor
⚪ agents import sizesMeasured 344 runtime imports as minified bundles. The primary size is gzip; raw minified size is included for diagnosis. An existing import growing by more than 10% is marked red. This report is informational.
Compared No import sizes changed. All 344 current runtime imports
Reported by agent-think[bot]. |
threepointone
marked this pull request as ready for review
September 24, 2026 01:12
agents
@cloudflare/ai-chat
@cloudflare/codemode
hono-agents
@cloudflare/shell
@cloudflare/think
@cloudflare/voice
@cloudflare/worker-bundler
commit: |
threepointone
force-pushed
the
fix/think-stable-truncation-prefix
branch
from
September 24, 2026 01:35
b772d69 to
5a8e256
Compare
threepointone
force-pushed
the
fix/think-stable-truncation-prefix
branch
from
September 24, 2026 01:43
5a8e256 to
5071db3
Compare
threepointone
force-pushed
the
fix/think-stable-truncation-prefix
branch
from
September 24, 2026 02:03
5071db3 to
2bdee44
Compare
Measures how each context-reduction mechanism changes the prompt prefix turn over turn. Media eviction and compaction rewrite it once; read-time truncation rewrote it every turn because its cutoff slid with the message count. Cut at a multiple of 8 messages instead. Closes #2200. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
threepointone
force-pushed
the
fix/think-stable-truncation-prefix
branch
from
September 24, 2026 07:03
2bdee44 to
33b1e73
Compare
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This measures how each context-reduction mechanism changes the prompt prefix from one turn to the next, and fixes the one that rewrote it every turn.
Measurement
A new test agent runs real Think turns against a mock model and records the full prompt of every model request. For each turn it compares the first request with the request sent before it. Providers cache on a byte-identical prefix, so the shared prefix is the part the cache can hit.
packages/think/src/tests/prompt-cache.test.tsasserts the results:The sliding cutoff (
messages.length - 4) rewrote a message near the end of the prefix on every turn. The measured cost uses character units, with cached input billed at a tenth of fresh input:main)So on
main, truncation cost about 2.2 times what sending the untruncated history would. With 12,000-character user messages the ratio is 1.4 times.Fix
Think now cuts at a multiple of 8 messages, using
keepRecent = 4 + ((length - 4) % truncationStep), and passes the samekeepRecenttotruncateOlderMessagesandtruncateOlderToolResults. The model still always sees at least the 4 most recent messages in full, and at most 11 between cuts. The hydration floor and the media eviction clamp are unchanged, because they only need the most recent 4. A newtruncationStepclass field (default 8) sets the step;1restores the previous cut-every-turn behavior for models with a small context window.A compaction threshold that the compacted history still exceeds compacts on every append, which rewrites the summary every turn. The docs now warn about this.
hydrationByteBudget(32 MiB by default) is larger than any context window, so it does not affect caching in practice.Closes #2200
Docs: a new "Prompt caching" section in
docs/think/index.md, and a stepped-cutoff recipe for directtruncateOlderMessagesusers indocs/agents/sessions.md.Not done here: moving these mechanisms into
agents/context, as #2200 suggests. The stepped cutoff is the reference point that move would need.