Skip to content

Token efficiency: per-request usage telemetry, cost-based compaction trigger, stable tool list, smaller static prefix #6541

Description

@Hmbown

Found in a token-efficiency audit of origin/main (9bf3e4f) against the founder's local session data (115 sessions, 60 days; 71 runtime-store turns).

Baseline

  • deepseek-flash is 94.8% of 3.55B tokens (a miss costs 50× a hit).
  • Mean request ~363k tokens; median task 5 requests.
  • Fixed prefix ~22k tokens: system prompt 9.1k, tool schemas 13k (26 tools).
  • Resent history is mostly reasoning (≤40%), bash output (23%), tool-call args (15%).
  • Runtime-store cache hit rate 78.6%.
  • Sub-agents ≈21% of cost.

Changes, ranked

  1. Telemetry first (direct). Record per-request usage by billing type (uncached input, cached input, output), prefix-change events, per-tool error class, sub-agent usage and task completion. Session files hold only a running total today.
  2. Cost-based compaction trigger (flag). Trigger around 250k tokens instead of 80% of a 1M window, with a cache-hitting summary call. Replay estimate: −15% input cost as compaction works today, −31% once the summary call is cached. Tell the model where the saved pre-compaction history is. Validate offline and online.
  3. Stable tool list (direct/flag). Loading a deferred tool re-pins the tool list; the next request's cache hit drops to 54% vs 96% (n=33/51, turn_loop.rs:3621-3638, 4133-4150). Use a stable dispatcher or the sub-agent "hydrate" pattern. ≈2–4% of cost.
  4. Smaller static prefix (flag).
    • The Skills index is ~19k chars at median but load_skill ran 10 times in 11.6k calls: index names plus one line each.
    • Drop duplicate/unused always-loaded tools (8 computer-use MCP tools from two duplicate servers; workflow called 7 times).
    • Move the MCP registry instruction into tool descriptions.
    • ≈7k tokens per request.
  5. Small direct fixes.
    • Send prompt_cache_key on OpenAI routes.
    • Use the unused Anthropic cache breakpoints (2 of 4 used).
    • Keep bash output over 30 KB retrievable instead of dropping the middle.
    • Delete the dead CALM_PERSONALITY text.
    • A goal edit should not re-pin the whole prefix.
  6. Prompt (flag). The "concise" verbosity setting tells the model to "minimize token usage" (prompts.rs:123-133). Replace it with a description of the output style, since telling the model to save tokens degrades ambition. Merge the overlapping translation instruction layers.
  7. Proposals only. A cache-aware auto-router (14/115 sessions switched model mid-session); a cheaper sub-agent model only when the parent is frontier; an A/B on whether DeepSeek bills replayed reasoning.

Validation: the eval harness (#6506) is the gate. Cost per completed task falls, and success, tool errors, turns and cache hit rate don't regress.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    • Status
      Backlog

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions