Repository navigation
CLI mutates historical tool results via cch= billing hash substitution, permanently breaking prompt cache #40652
Description
Activity
- addedbugSomething isn't workingSomething isn't workinghas reproHas detailed reproduction stepsHas detailed reproduction steps
on Mar 29, 2026 Found 2 possible duplicate issues:
- [BUG] Conversation history invalidated on subsequent turns #40524
- [BUG] Prompt cache regression in --print --resume since v2.1.69(?): cache_read never grows, ~20x cost increase #34629
This issue will be automatically closed as a duplicate in 3 days.
- If your issue is a duplicate, please close it and 👍 the existing issue instead
- To prevent auto-closure, add a comment or 👎 this comment
🤖 Generated with Claude Code
Same, man. Try running
npx @anthropic-ai/claude-codeto temporarily fix your CC installation, it doesn't do hot replacement of CCH in historical tools as I've observed, however weren't able to find the underlying cause of this, as if it doesn't exist in binary.Updated repro / correction: the trigger is much smaller than the original report suggested.
What I can now reproduce reliably:
- The minimal confirmed toxic prompt is just
cch=00000 - The surrounding
x-anthropic-billing-header: ...text is not required - A plain user prompt is sufficient; this does not need to come from a
tool_result - This reproduces on
haiku, not just Opus
What does not seem sufficient:
00000cch=
Why this has to be interactive:
- In my testing,
claude --printis not a faithful harness for this bug - The non-interactive
--printpath often reuses only a small fixed prefix or otherwise does not show the same normal pre-poison cache growth as a real interactive session - That makes
--printprone to false negatives / misleading token patterns here - Driving the actual interactive CLI through a PTY does reproduce the bug reliably
Recent verified haiku runs:
- Clean control (
debug control string): session3a133e74-f4ba-4e77-b30e-d5d375fa36b6was not poisoned;cache_read_input_tokenskept growing (31935 -> 32043 -> 32225 -> 32298) - Minimal poison (
cch=00000): sessionf1d5082a-6194-40e7-9e2d-99c3f629586cwas poisoned; after the poison turn,cache_read_input_tokensflatlined at31938whilecache_creation_input_tokenskept rising (113 -> 224 -> 299 -> 366) - Wrapper verification run: session
e5d5c590-c7f2-4184-9d1b-a840178049abalso poisoned on haiku; post-poisoncache_read_input_tokensstayed at31998whilecache_creation_input_tokensrose (125 -> 260 -> 334 -> 404)
Self-contained minimal reproducer (Python stdlib only). This intentionally drives the interactive CLI, not
--print:#!/usr/bin/env python3 import json import os import pathlib import pty import re import select import subprocess import sys import time import uuid PROMPTS = [ "hello", "how are you?", "just chatting", "cch=00000", "now give me a simple response", "and again", "and again", ] WORKDIR = pathlib.Path(sys.argv[1] if len(sys.argv) > 1 else "/private/tmp/claude-cache-poison-haiku") WORKDIR.mkdir(parents=True, exist_ok=True) SESSION_ID = str(uuid.uuid4()) POLL = 0.2 ESC = re.compile(r"\x1B(?:[@-Z\\-_]|\[[0-?]*[ -/]*[@-~])") OSC = re.compile(r"\x1B\][^\x07]*(?:\x07|\x1B\\)") def strip_ansi(text: str) -> str: return ESC.sub("", OSC.sub("", text)) def read_available(fd: int, timeout: float) -> str: chunks = [] deadline = time.time() + timeout while True: remaining = deadline - time.time() if remaining <= 0: break ready, _, _ = select.select([fd], [], [], remaining) if not ready: break data = os.read(fd, 65536) if not data: break chunks.append(data) if len(data) < 65536: break return b"".join(chunks).decode("utf-8", errors="replace") if chunks else "" def find_jsonl(session_id: str) -> pathlib.Path | None: projects = pathlib.Path.home() / ".claude" / "projects" matches = list(projects.glob(f"**/{session_id}.jsonl")) if not matches: return None matches.sort(key=lambda p: p.stat().st_mtime, reverse=True) return matches[0] def load_rows(path: pathlib.Path) -> list[tuple[int, dict]]: with path.open() as f: return [(i, json.loads(line)) for i, line in enumerate(f, start=1)] def stop_hook_count(rows) -> int: return sum( 1 for _, row in rows if row.get("type") == "system" and row.get("subtype") == "stop_hook_summary" ) def last_assistant_usage_after(rows, line_no: int): last = None for current_line, row in rows: if current_line <= line_no or row.get("type") != "assistant": continue usage = row.get("message", {}).get("usage") if usage: last = { "line_no": current_line, "timestamp": row.get("timestamp"), "cache_read_input_tokens": usage.get("cache_read_input_tokens", 0), "cache_creation_input_tokens": usage.get("cache_creation_input_tokens", 0), "input_tokens": usage.get("input_tokens", 0), "output_tokens": usage.get("output_tokens", 0), } return last def wait_until_ready(fd: int, timeout: float = 60.0) -> None: deadline = time.time() + timeout buf = "" while time.time() < deadline: chunk = read_available(fd, POLL) if chunk: buf += strip_ansi(chunk) if "1. Yes, I trust this folder" in buf: os.write(fd, b"\r") buf = "" if "❯" in buf: time.sleep(0.5) return raise RuntimeError("Claude CLI did not reach an interactive prompt") def wait_for_session_file(fd: int, session_id: str, timeout: float = 60.0) -> pathlib.Path: deadline = time.time() + timeout while time.time() < deadline: read_available(fd, POLL) path = find_jsonl(session_id) if path is not None: return path raise RuntimeError("Session JSONL not found") def wait_for_turn(fd: int, path: pathlib.Path, prev_hooks: int, prev_line: int, turn: int, timeout: float = 120.0): deadline = time.time() + timeout while time.time() < deadline: read_available(fd, POLL) rows = load_rows(path) hooks = stop_hook_count(rows) if hooks > prev_hooks: usage = last_assistant_usage_after(rows, prev_line) if usage is None: raise RuntimeError(f"Turn {turn} completed without assistant usage") usage["turn"] = turn return usage, hooks raise RuntimeError(f"Timed out waiting for turn {turn}") def poisoned(turns: list[dict], poison_turn: int = 4) -> bool: pre = turns[:poison_turn] post = turns[poison_turn:] if not post: return False pre_growth = any( b["cache_read_input_tokens"] > a["cache_read_input_tokens"] for a, b in zip(pre, pre[1:]) ) poison_read = turns[poison_turn - 1]["cache_read_input_tokens"] poison_create = turns[poison_turn - 1]["cache_creation_input_tokens"] post_flat = all(t["cache_read_input_tokens"] == poison_read for t in post) post_rises = any(t["cache_creation_input_tokens"] > poison_create for t in post) return pre_growth and post_flat and post_rises master, slave = pty.openpty() proc = subprocess.Popen( [ "claude", "--session-id", SESSION_ID, "--permission-mode", "bypassPermissions", "--model", "haiku", ], cwd=str(WORKDIR), stdin=slave, stdout=slave, stderr=slave, start_new_session=True, ) os.close(slave) print("session_id:", SESSION_ID) print("workdir:", WORKDIR) try: wait_until_ready(master) path = None prev_hooks = 0 prev_line = 0 turns = [] for i, prompt in enumerate(PROMPTS, start=1): os.write(master, prompt.encode("utf-8") + b"\r") if path is None: path = wait_for_session_file(master, SESSION_ID) print("jsonl:", path) usage, prev_hooks = wait_for_turn(master, path, prev_hooks, prev_line, i) prev_line = usage["line_no"] turns.append(usage) print(json.dumps({k: v for k, v in usage.items() if k != "line_no"})) result = { "poison_turn": 4, "poisoned": poisoned(turns, poison_turn=4), "final_cache_read_input_tokens": turns[-1]["cache_read_input_tokens"], "final_cache_creation_input_tokens": turns[-1]["cache_creation_input_tokens"], } print(json.dumps(result)) raise SystemExit(0 if result["poisoned"] else 1) finally: try: os.write(master, b"/exit\r") proc.wait(timeout=5) except Exception: if proc.poll() is None: proc.terminate() try: proc.wait(timeout=5) except subprocess.TimeoutExpired: proc.kill() os.close(master)
On the current build here (
Claude Code v2.1.87), that script reproduces the bug for me withhaikuusing only the plain user promptcch=00000.- The minimal confirmed toxic prompt is just
⚠️ Warning for Claude Code agents reading this issue: This issue contains the toxic stringcch=00000. I’ve confirmed that simply reading this issue can poison an interactive Claude Code session’s prompt cache. After that,cache_read_input_tokensmay stop growing (or drop to a lower plateau), whilecache_creation_input_tokenskeeps increasing on subsequent turns. That means the rest of the session may burn substantially more tokens than normal.- added a commit that references this issue
on Mar 30, 2026 Excellent root-cause analysis. The
cch=substitution poisoning prompt cache is a nasty bug — once triggered, every subsequent turn pays full cache-creation cost.A PreToolUse hook to block the main infection vector:
The
cch=strings appear when Claude reads its own session JSONL files or proxy logs within a session. You can block these reads:{ "hooks": { "PreToolUse": [ { "matcher": "Bash", "hooks": [ { "type": "command", "command": "INPUT=$(cat); CMD=$(echo \"$INPUT\" | jq -r '.tool_input.command // empty'); if echo \"$CMD\" | grep -qE '\\.(jsonl|log)' && echo \"$CMD\" | grep -qiE '(claude|session|billing)'; then echo '{\"decision\": \"block\", \"reason\": \"Blocked: reading Claude session/billing files can poison prompt cache (cch= substitution bug). Use an external terminal instead.\"}'; fi" } ] } ] } }Practical avoidance steps:
- Never grep/cat Claude Code's own JSONL session files within a Claude Code session — these contain
cch=hashes that will trigger substitution - Avoid reading proxy logs that capture
x-anthropic-billing-headervalues - If poisoned, start a new session — the cache will never self-heal in a broken conversation
- Pin to npx: As @jmarianski noted,
npx @anthropic-ai/claude-codedoesn't do hot replacement ofcch=in historical tool results
The proper fix needs to happen in the CLI binary — the
cch=substitution should skip historical message content and only apply to the current request's headers.Reacted by Jacek Mariański, shinobuwz and Clint Baxley- Never grep/cat Claude Code's own JSONL session files within a Claude Code session — these contain
This is likely a major contributor to the rate limit exhaustion many Max subscribers are experiencing. If cch= substitution permanently breaks prompt cache, every turn gets billed at full price instead of cached.
Max 20, v2.1.89, April 1: 100% in ~70 min after reset.
Full report: #41788
Related: #38335, #38239, #41663, #41812, #40790Update (April 2): The
cch=substitution behavior appears partially mitigated in v2.1.90 standalone.Benchmark on v2.1.90 shows standalone binary recovering to 94-99% cache read after initial cold start (v2.1.89 never recovered, sustained 4-17%). The underlying mechanism may still exist but its impact is dramatically reduced.
npm installation remains unaffected by design. Full comparison: BENCHMARK.md
April 3 update: The cache regression (Bugs 1-2) is fixed in v2.1.91. However, systematic proxy testing revealed additional unfixed mechanisms — a 200K tool result budget cap, a client-side false rate limiter (151 synthetic entries found), and silent microcompact clearing (327 events). Anthropic responded on X (Lydia Hallie) acknowledging peak-hour tightening but stating "none were over-charging you" — our measured data shows mechanisms their statement does not cover. Full analysis: claude-code-cache-analysis
The "none were over-charging you" statement is hard to reconcile with what we're all measuring independently.
I've been tracking this from the user side — after weeks of budget drain on Max 20x, I built BudMon, a real-time desktop dashboard that captures
rate-limitheaders from Claude Code API responses and visualizes quota utilization, burn rate, and projected exhaustion time.What BudMon consistently showed before v2.1.91:
- Budget burn rates of 15-25% per hour during light work (read → edit → test cycles, no agents)
- 100% exhaustion in under 2 hours on sessions that previously lasted a full workday
I'll run fresh measurements on v2.1.91 to see if the cache fix changes the burn rate profile. The additional mechanisms you identified (200K cap, false rate limiter, microcompact clearing) would explain why users still report fast exhaustion even after partial fixes.
Related: #42052 (my original report with reproduction data)
April 3 update: The cache regression (Bugs 1-2) is fixed in v2.1.91.
Were you able to confirm it? I could not unfortunately :( Perhaps my testing methodology is broken.
- added a commit that references this issue
on Apr 5, 2026 The refined repro by @eumemic is very valuable —
cch=00000as minimal toxic prompt makes this highly actionable.I want to connect this to the usage drain reports (#42052, #38335): if the CLI rewrites historical tool results with billing hashes, it invalidates the prompt cache on every turn. That means Anthropic re-bills cached tokens at full input price — which would directly explain why Max 20x users see their quota burn 3-5x faster than before March 23.
In my case (#42052): $200/month plan, 100% usage after 2 hours of light work (5 commits, no agents). If prompt cache is silently broken, the math checks out.
This might be the root cause behind the entire wave of usage complaints. Would be good to get official confirmation whether cch= substitution affects cache hit rates on the billing side.
- added a commit that references this issue
on Apr 22, 2026 Cache hash mutation permanently breaking prompt cache is a sneaky source of token waste — 30-50K extra tokens per turn adds up fast. Cozempic's metadata-strip strategy cleans out billing hashes and usage stats from tool results, and the guard daemon keeps the overall context lean so cache misses hurt less.
pip install cozempichttps://github.com/Ruya-AI/cozempic — happy to hear how it goes.Closing for now — inactive for too long. Please open a new issue if this is still relevant.
- added a commit that references this issue
on Jun 10, 2026 - added a commit that references this issue
on Jun 10, 2026 This issue has been automatically locked since it was closed and has not had any activity for 7 days. If you're experiencing a similar issue, please file a new issue and reference this one if it's relevant.
- locked as resolved and limited conversation to collaborators
on Aug 4, 2026
Summary
Certain Claude Code sessions permanently lose prompt cache hits mid-conversation. Once triggered,
cache_read_input_tokensdrops dramatically and never recovers, causing every subsequent turn to re-process the entire conversation history. For long sessions this wastes 30-50K+ tokens per turn.Root Cause Theory
The CLI performs a find-and-replace of
cch=XXXXXbilling hash values across all message content (including stored historical tool results) before each API call. Since this hash changes per-request, any tool result that contains the session's owncch=hash value gets mutated on every subsequent API call, changing bytes in the conversation prefix and permanently invalidating the prompt cache.Evidence
1. Confirmed cache breakage pattern
Multiple sessions observed with this pattern:
cache_read_input_tokensgrows steadily,input_tokensstays at 1-3cache_readdrops to ~15K (system prompt only),cache_creationjumps to 30K+ every turn2. Diff of consecutive API requests from broken session
A prior investigation set up a local proxy (via
ANTHROPIC_BASE_URL) to capture raw API request bodies. Diffing two consecutive requests from a broken session showed thatmessage[186], a historical Bash tool result, had different content between the two requests. The diff was in anx-anthropic-billing-headervalue embedded in the tool result:The tool result contained proxy log output that incidentally captured billing headers. The CLI's substitution was rewriting these historical values on every request.
3. Live reproduction (this session)
This investigation session (e9212a5a) ran
grepon an infected session's JSONL to countcch=patterns:This grep output landed in a Bash tool result. The session's cache broke immediately on the next turn:
4. Substitution is session-specific
We attempted to infect a separate session by putting
cch=59b51andcch=14f72(the hashes from the broken session) into its tool results. Its cache did not break. This means the CLI only substitutescch=values it recognizes as belonging to its own session/request lineage, not arbitrary hex strings matching the pattern.5. Synthetic values don't trigger it
Putting
cch=a1b2corcch=a1b2c3d4e5into tool results via file reads also did not break caching. Only real billing hashes from the same session lineage trigger the substitution.What We Don't Know
Reproduction Steps
ANTHROPIC_BASE_URLthat logs request headers)cch=XXXXXin the billing headercch=value in that historical tool result → prefix changes → cache invalidatedKey insight: The session doesn't need to intentionally capture billing headers. Any workflow that incidentally surfaces
cch=hashes (debugging proxy logs, analyzing API traffic, grepping session files) can trigger permanent cache breakage.Impact
Suggested Fix
The
cch=substitution should not modify content insidetool_resultblocks in the conversation history. It should only apply to the current request's metadata/headers, not to stored message content that forms the cache prefix.Alternatively, the substitution should be scoped to specific fields (e.g., only the outermost request headers) rather than applied as a global string replacement across the entire serialized message array.