Found in a token-efficiency audit of origin/main (9bf3e4f) against the founder's local session data (115 sessions, 60 days; 71 runtime-store turns).
Baseline
- deepseek-flash is 94.8% of 3.55B tokens (a miss costs 50× a hit).
- Mean request ~363k tokens; median task 5 requests.
- Fixed prefix ~22k tokens: system prompt 9.1k, tool schemas 13k (26 tools).
- Resent history is mostly reasoning (≤40%), bash output (23%), tool-call args (15%).
- Runtime-store cache hit rate 78.6%.
- Sub-agents ≈21% of cost.
Changes, ranked
- Telemetry first (direct). Record per-request usage by billing type (uncached input, cached input, output), prefix-change events, per-tool error class, sub-agent usage and task completion. Session files hold only a running total today.
- Cost-based compaction trigger (flag). Trigger around 250k tokens instead of 80% of a 1M window, with a cache-hitting summary call. Replay estimate: −15% input cost as compaction works today, −31% once the summary call is cached. Tell the model where the saved pre-compaction history is. Validate offline and online.
- Stable tool list (direct/flag). Loading a deferred tool re-pins the tool list; the next request's cache hit drops to 54% vs 96% (n=33/51,
turn_loop.rs:3621-3638, 4133-4150). Use a stable dispatcher or the sub-agent "hydrate" pattern. ≈2–4% of cost.
- Smaller static prefix (flag).
- The Skills index is ~19k chars at median but
load_skill ran 10 times in 11.6k calls: index names plus one line each.
- Drop duplicate/unused always-loaded tools (8 computer-use MCP tools from two duplicate servers;
workflow called 7 times).
- Move the MCP registry instruction into tool descriptions.
- ≈7k tokens per request.
- Small direct fixes.
- Send
prompt_cache_key on OpenAI routes.
- Use the unused Anthropic cache breakpoints (2 of 4 used).
- Keep bash output over 30 KB retrievable instead of dropping the middle.
- Delete the dead
CALM_PERSONALITY text.
- A goal edit should not re-pin the whole prefix.
- Prompt (flag). The "concise" verbosity setting tells the model to "minimize token usage" (
prompts.rs:123-133). Replace it with a description of the output style, since telling the model to save tokens degrades ambition. Merge the overlapping translation instruction layers.
- Proposals only. A cache-aware auto-router (14/115 sessions switched model mid-session); a cheaper sub-agent model only when the parent is frontier; an A/B on whether DeepSeek bills replayed reasoning.
Validation: the eval harness (#6506) is the gate. Cost per completed task falls, and success, tool errors, turns and cache hit rate don't regress.
Found in a token-efficiency audit of origin/main (9bf3e4f) against the founder's local session data (115 sessions, 60 days; 71 runtime-store turns).
Baseline
Changes, ranked
turn_loop.rs:3621-3638,4133-4150). Use a stable dispatcher or the sub-agent "hydrate" pattern. ≈2–4% of cost.load_skillran 10 times in 11.6k calls: index names plus one line each.workflowcalled 7 times).prompt_cache_keyon OpenAI routes.CALM_PERSONALITYtext.prompts.rs:123-133). Replace it with a description of the output style, since telling the model to save tokens degrades ambition. Merge the overlapping translation instruction layers.Validation: the eval harness (#6506) is the gate. Cost per completed task falls, and success, tool errors, turns and cache hit rate don't regress.