Skip to content

[think] Tool outputs are silently truncated to 500 chars #2014

Description

@urjitc

ai slop but still true:

Describe the bug

Think._assembleModelMessages calls truncateOlderMessages unconditionally:

// @cloudflare/think@0.15.1 dist/think.js:2463
const truncated = truncateOlderMessages(await this._repairTranscriptForProvider(this.messages));

No options are passed, so the defaults bind: keepRecent: 4, maxToolOutputChars: 500.

For oversized structured outputs, truncateToolOutput preserves the container shape but replaces the contents with marker keys:

// agents@0.20.1 dist/tool-output-truncation-CNnnGZQ3.js:3-4
const TRUNCATED_FLAG = "__truncated";
const TRUNCATED_CHARS = "__truncatedChars";

Those markers do not satisfy the tool's declared output schema. Any toModelOutput that validates its input — for example with a zod .parse() — throws.

The failure is permanent rather than transient, because the AI SDK invokes toModelOutput from inside convertToModelMessages, which re-runs on every replay of the thread. So once a tool output ages past the 4-message window, every subsequent turn throws the same error at assembly time, the model request is never sent, and the thread is unusable until history is cleared.

To Reproduce

Steps to reproduce the behavior:

  1. Build a Think agent with a tool whose toModelOutput validates the persisted output against its declared schema (a throwing .parse).
  2. Call that tool so it returns a structured output larger than 500 characters.
  3. Send five or more further turns, pushing the tool call out of the keepRecent: 4 window.
  4. Send one more message.
  5. See validation fail on the {__truncated, __truncatedChars} markers. Every subsequent turn fails identically.

Expected behavior

Truncation should not produce tool outputs that violate the schema the tool itself declares. Options:

  • exclude structured outputs with a declared schema from truncation, or
  • expose the truncation options on Think (keepRecent, maxToolOutputChars, maxTextChars, including opting out), or
  • at minimum, make a truncation-induced validation failure non-fatal so one aged tool output cannot wedge a thread permanently.

Screenshots

N/A.

Version:

@cloudflare/think@0.15.1 with agents@0.20.1.

Additional context

This is the same failure mode as #1498, which was closed by making Think's own read.toModelOutput tolerant of truncated input. The truncation side was not changed — it now preserves the container shape instead of stringifying the whole output, but the contents still violate strict schemas, so any tool that validates hits it again.

We currently mitigate twice: safeParse with a raw-output fallback in our own toModelOutput, and a patch removing the truncateOlderMessages call from _assembleModelMessages. The first costs us citations on truncated turns; the second is a fork of the SDK.

Activity

  1. SimplyCorey commented on Aug 5, 2026

    @SimplyCorey

    I'm seeing the same errors when websearch is used.

    Another repro of this, with no custom toModelOutput anywhere in the path — the truncation markers break the AI SDK's own provider-executed tool schema, so there is no user code involved in the failure at all.

    Setup — a Think chat agent registering Anthropic's server-side web search:

    tools: { web_search: anthropic.tools.webSearch_20250305({ maxUses: 3 }) }

    What happens — each search result carries an opaque encryptedContent blob (0.6–3.4 KB in our traffic). Once the tool call ages past keepRecent: 4, _assembleModelMessages → truncateOlderMessages(…, maxToolOutputChars: 500) budgets each of the 10 result objects at childMaxChars(500, 10) = 80 chars. Every one falls through truncateObject → shrinkStringFields → compactObjectMarker, which discards all original keys:

    [{"__truncated":true,"__truncatedChars":1566}, {"__truncated":true,"__truncatedChars":2350}, … ]

    convertToAnthropicMessagesPrompt then validates that array against webSearch_20250305OutputSchema and throws AI_TypeValidationError (url, encryptedContent, type all undefined), inside getArgs — so the request never leaves the Worker. Because truncation is recomputed from stored history every turn, the conversation is permanently unusable rather than transiently degraded. Every subsequent prompt fails; a new conversation works.

    Why provider-executed results are a distinct case: they can never be usefully truncated. encryptedContent is opaque to the caller and the provider requires it back byte-for-byte, so shrinking it has exactly two outcomes — unchanged, or broken. Truncation is correct for a large ordinary tool output; it is categorically wrong here. Of the three options in the issue, exempting parts flagged providerExecuted looks like the cleanest floor: no new configuration surface, and it can't regress the context-saving truncation exists to provide.

    Versions — first hit on think@0.10.0 + agents@0.17.3. Still reproduces on think@0.15.1 + agents@0.20.1: think.js:2463 still calls truncateOlderMessages with no options, and tool-output-truncation-CNnnGZQ3.js is byte-identical (md5 9b86d4922ca229a1d185ee6b2bdbf5c4) across 0.17.3 → 0.20.1 with no providerExecuted reference. Independent of the AI SDK major — the same call is in @ai-sdk/anthropic@4.0.30 (dist/index.js:3233).

    Workaround — truncateOlderMessages doesn't persist, so the intact output is still in this.messages. A beforeStep hook can restore any providerExecuted result by toolCallId before conversion.

  2. ppcvote commented on Aug 7, 2026

    @ppcvote

    Reproduced against the shipped module rather than the dist bundle, and there is a detail that I think decides the shape of the fix: this violates the module's own documented contract, so it is a bug against spec rather than a trade-off someone chose.

    sanitize.ts says what compaction is supposed to do:

    Compact tool outputs over 1KB with truncateToolOutput, preserving the structured output shape, and annotate metadata.compactedToolOutputs with the compacted tool-call IDs.

    Shape preservation is the stated behaviour. The marker paths are the ones that break it.

    Reproduction

    A tool output with a declared shape ({status, totalCount, rows: {id, title, score}[]}), through the real truncateToolOutput:

    === maxChars = 500 ===
    rows= 8  original= 799ch  truncated=true   schemaValid=false  <- rows[0] missing id
    rows=40  original=3908ch  truncated=true   schemaValid=false  <- rows[0] missing id
    
    === maxChars = 1000 ===
    rows= 8  original= 799ch  truncated=false  schemaValid=true
    rows=40  original=3908ch  truncated=true   schemaValid=false  <- rows[0] missing id
    

    The produced output:

    {"status":"ok","totalCount":8,"rows":[{"__truncated":true,"__truncatedChars":761,"note":"Tool output omitted because…"}]}

    status and totalCount survive, so the value still looks well-formed at the top level and fails a level down. That is why it surfaces as a thrown parse rather than something obviously malformed.

    Worth noting the aggressiveness separately from the schema question: 799 characters against a 500-character budget discards the entire rows array rather than shortening the title strings, which would have fit comfortably.

    Five paths that break the shape, not one

    Both container types have them, so a fix at the object level alone would not close it:

    # site what it produces
    1 truncateObject injects __truncated / __truncatedChars into an object with a declared key set
    2 compactObjectMarker replaces the object entirely with {__truncated, __truncatedChars, note}
    3 truncateArray pushes "Array output truncated …", a string into an array of objects
    4 compactArrayMarker replaces the array with [{__truncated, …}]
    5 compactArrayMarker or with ["Array output omitted …"], again a string

    Before proposing removal, the argument against it

    The obvious fix is to stop emitting the markers, and I do not think that is right on its own. The markers plausibly exist so the model reading the transcript knows the output was cut. Deleting them and keeping only metadata.compactedToolOutputs would hand the model a plausible-looking complete object, which trades a loud failure for a silent one.

    Two things make that avoidable:

    • truncateString already discloses inline, appending ... [truncated N chars], and a shortened string is still a string, so it survives the schema. Disclosure to the model does not require breaking the shape; it already happens in the one place where it costs nothing.
    • The caller already carries the out-of-band signal. enforceRowSizeLimit records metadata.compactedToolOutputs, and compaction.ts branches on the returned truncated boolean. Neither reads the in-band markers.

    I checked that last claim across the repo: no production code reads __truncated or __truncatedChars. The only readers are two assertions in tests/experimental/memory/utils/compaction.test.ts (lines 98-99 and 204-205), which pin the current behaviour.

    That is the part that makes this a maintainer call rather than an obvious patch. The markers are tested behaviour. Changing them changes a test that someone wrote deliberately, even though the behaviour contradicts the docstring one module over.

    What I would propose

    Shape-preserving truncation only: shorten string leaves, never add keys, never replace a container, never change an element's type. Where that cannot reach the budget, return the shape-preserved best effort and let the caller act on truncated — both callers already have a downstream lever (enforceRowSizeLimit falls through to truncateTextParts, and it is a storage bound rather than a correctness one).

    The model still learns the output was cut, from the inline string suffixes. The schema still validates. The caller still knows, from a return value it already reads.

    Happy to send that as a PR including the two test updates, with the reasoning above in the description so the behaviour change is visible rather than smuggled. Would rather have a maintainer say the direction is right first, given it rewrites tested behaviour in a module two packages depend on.


    Context for why I am confident about the general shape rather than just this instance: a group of four independently written scanners spent the last weeks settling where a bound gets declared in a machine-readable artifact (envelope RFC). The rule that survived is that the bound is declared beside the payload, never inside it, precisely because an in-band disclosure has to satisfy the payload's contract and a disclosure that violates it is worse than none. This is the sharpest instance of that I have seen, because here the disclosure does not merely mislead a consumer, it makes the value unparseable.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't workingthink

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions