feat(context-budget): ship the startup-payload measurement plugin (0.6.0) - #2932
Conversation
Nine dispatched research runs plus first-hand measurement on CLI v2.1.232 establish the design for a new `context-budget` plugin whose skill measures and trims a session's fixed startup payload. Key results, all measured rather than asserted: - Rule shape, not the setting, decides whether a tool schema ships. A bare tool name in a deny rule removes the definition from the request; a scoped rule is a runtime guard whose schema is still billed every turn. - Deferral does not shrink the request, inverting the premise the source material's headline lever rests on. - The skill listing is hard-capped, so disabling skills saves nothing while over the cap. Disabling 45 of 65 plugins moved the row by zero. - Custom agents are not capped and do scale, but the deny form the sub-agents docs present for them leaves the description in the payload. - `System tools` has skill-frontmatter tokens subtracted, so it is not comparable across runs whose skill listing differs. Also records five corrections owed to this repository, including the safe-mode clean-room claim and unhobble's CLAUDE_CODE_SIMPLE gotcha. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
…ing, split the two SIMPLE env vars Second-eyes review over the locked brief. The /context headline excludes the deferred row (35.19k computed vs 35.3k reported), so the Goal's '35.9k of a 35.3k-headline session' phrasing was incoherent and is restated precisely. Binary inspection confirms CLAUDE_CODE_SIMPLE (--bare mode) and CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT are two distinct env vars, which upgrades the unhobble correction: its gotcha is wrong about documentation status AND attributes prompt-stripping to the wrong var. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
Phase 0 (docs/topics/context-budget/PLAN.md): the Agent SDK's getContextUsage() is real and probed live — exact integers matching the CLI, but systemTools/deferredBuiltinTools arrive unpopulated, so the engine is SDK-primary with A/B differencing for per-built-in-tool attribution. skillOverrides is documented (settings reference + skills page, fetched raw), with the plugin-skill carve-out; the fetch also surfaced skillListingBudgetFraction and skillListingMaxDescChars as documented levers. Corrections applied: - checks-and-sweep: safe-mode/CLAUDE_CONFIG_DIR are not clean rooms (dated correction with measured evidence) - coverage-matrix S7: deferred-tool half now owned by context-budget, premise corrected (deferral does not shrink the request) - permission-rule-hygiene: stale dated quote replaced with the page's current version-floor wording (fetched 2026-08-17); tightening gap recorded in the changelog - claude-config 0.38.7: unhobble's CLAUDE_CODE_SIMPLE gotcha rewritten against the binary (documented, and prompt-stripping belongs to the sibling CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT); eval 8 regraded onto the scope-boundary reasoning Filed #2895 (preload sentinel unsound as preload proof, follow-up to #2338) and #2896 (skill-quality verb-contract mismatch check). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
… cloud handoff The .work memory slice dies with the cloud container, and Phase 3's lever catalogue needs the per-claim citations, so the nine research slices plus the measurement/design/interview artifacts move to docs/topics/context-budget/research/ (typos exclusion added: the corpus quotes minified binary internals verbatim as evidence). The handoff file is committed for the same reason, with a .work/handoffs/ mirror for the standard local contract. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
….1.0 New plugin with single skill /context-budget:audit (read-only). The engine (skills/audit/scripts/measure.mjs) is SDK-primary: exact structured context usage over the Agent SDK with live tool enumeration, degrading to a version-aware parser of headless /context output (display-rounded, refuses loudly on format drift) and then to a structured error with a remediation - never a wrong number. Per-tool attribution of the built-in tool pools by bare-name-deny A/B differencing with an optional additivity verification; enforced comparability rules (skill-listing signature, one mode, one binary version); offline compare producing ledger rows; per-project state-keyed ledger (one file per run plus an appended history line). Every record stamped with the measured binary path/version, mode, precision, and session kind. Registered in the marketplace, catalog, audit leaf-name owner set, and the state-key sync cluster. Hermetic test suite covers the parser traps, compare comparability, and ledger retention. Verified live against a pinned binary: per-tool deltas reproduce and pass the additivity check. Phase 2 of docs/topics/context-budget/PLAN.md; Phases 3-6 remain. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
….2.0) skills/audit/reference/levers.json: 19 levers, each a data row carrying its honesty category (six-term vocabulary with an explicit request-vs-context- window dual-ledger distinction), category basis, measurement-resolved conditions, posture, detection, measurement route, exact emitted config, official citations, verified date, and recheck trigger. Net-negative and unverified rows are structurally barred from the recommendable posture; the memory-file and hook lanes are route-outs in the catalogue meta. levers.test.sh makes the honesty rules mechanical, including a scan that no row ships a token figure (it caught one authoring violation, which was fixed in the row, not the test). SKILL.md gains the lever-presentation step; README documents the catalogue; changelog and version bump per convention. Phase 3 of docs/topics/context-budget/PLAN.md; Phases 4-6 remain. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
skills/audit/reference/report.md fixes the audit's default deliverable: stamped header (binary, mode, precision, session kind, model, cwd, time), smart-zone headline leading with reclaimed reasoning space (context-guard zone framing presence-gated), measured category totals, ranked per-tool attribution where incomparable rows carry reasons instead of numbers and unmeasured tools are listed rather than omitted, lever findings grouped by honesty category with citations and emitted config, route-outs, and a degradations section. Reports persist one file per run under the keyed data dir. The what-the-report-never-does list bars external-arithmetic reconciliation, bucket merging, and any apply. Phase 4 of docs/topics/context-budget/PLAN.md; Phases 5-6 remain. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
…int (0.4.0) The fix path runs only on the explicit `fix` argument (verb-contract override): per-lever walkthrough in ranked-report order over recommendable-on-fit catalogue rows whose conditions this audit resolved by measurement, one lever at a time with a mandatory apply -> re-measure -> compare -> ledger loop. Write posture splits by scope: project settings editable after per-diff approval; user-global ~/.claude/settings.json print-only (never written -- auto mode's classifier can approve a protected- path write with no human, so the prompt cannot be the protection); managed policy never targeted; env levers printed. hooks/settings-write-ask.mjs (PreToolUse, exec-form node for Windows safety): permissionDecision "ask" on any Write/Edit targeting a settings surface, so auto mode prompts instead of silently approving. Documented as a checkpoint, not a guarantee. Fail-open; settings_write_ask_enabled userConfig kill switch via the hook mirror; hermetic contract test. Phase 5 of docs/topics/context-budget/PLAN.md; Phase 6 remains. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
…(0.5.0) Three fix-path eval cases join the six measurement-honesty cases: mutation only on the explicit fix argument, user-global print-only with the auto-mode classifier caveat as the stated reason, and one-lever-at-a-time apply -> re-measure -> ledger with batches refused on attribution grounds. Evals-quality gate passes with zero warnings. Mechanical acceptance sweeps recorded in the changelog and PLAN: the shipped plugin greps clean of every research-run figure, catalogue and hook contract tests pass, and the skill-layout gate reports zero errors. A fresh-context acceptance verifier over the six acceptance criteria is running; its verdict and any fixes close the phase in a follow-up commit. Phase 6 of docs/topics/context-budget/PLAN.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
… shipped (0.5.1) The fresh-context acceptance verifier returned OVERALL: ACCEPT (all six criteria PASS, all commanded runs green). Its one actionable finding is fixed: the catalogue test's token-figure scan now fails any k-suffixed figure in a lever row outright and catches plain integers adjacent to the word token in either order, verified against a seeded violation. PLAN marks Phase 6 shipped with the verdict recorded; the SKILL.md 227-line soft warning is accepted as load-bearing fix-path posture. Closes the build phases of docs/topics/context-budget/PLAN.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
…0.6.0) First end-to-end shakedown of the shipped skill against this container (state key -> baseline -> 5-tool attribution -> report -> ledger): deltas reproduced, four-lever additivity exact, report persisted per contract. It surfaced a measured anti-lever - denying the tool-search tool forces the entire deferred pool upfront - now a deny-bare-tool catalogue caveat. Two fresh-context probes resolved open questions for v2.1.232 headless: a PreToolUse ask fires and blocks even under bypassPermissions (surfacing as a tool error carrying the reason; interactive unmeasured, documented as such), and HTTP MCP tools measure DEFERRED in a dedicated "MCP tools (deferred)" category - upstream #40314's upfront loading does not reproduce. Engine cli-parse headline now excludes every "(deferred)" category to match; hook header, SKILL.md, engine.md, and the catalogue updated. FINDINGS gains the post-acceptance probe record (count_tokens billing probe blocked here - no API credential - with the two-call design recorded); PLAN gains the closeout section. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b1e4d0a292
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…ning (0.6.1) Two PR #2932 review findings fixed: a same-second rerun of the same lever now collides into a numbered-suffix run file instead of overwriting the earlier point (one-file-per-run held only by luck; test added), and binary resolution prefers claude.exe over claude.cmd with .cmd/.bat shims executed through the shell, since Node cannot spawn Windows command shims directly. The Windows path is an honest manual-verification gap (no Windows hardware in this environment). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
…-setup-xnwt0w # Conflicts: # scripts/skill-leaf-name-registry.txt
…lice Per the topic-docs contract-slice lifecycle (prune with pointer): durable method, citations, and honesty rules live in the shipped plugin (plugins/context-budget/skills/audit/reference/); the two remaining open measurements graduate to the work-item tracker as #2954; the evidence record stays retrievable via this PR's pre-prune SHA, named in the PR body. The _typos.toml research exclusion is reverted with the corpus it excluded, returning the managed synced copy to canonical (also resolves the PR review's P1). References into the slice from changelogs and the grandfathered context-engineering-claude-5 slice now point at the durable homes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
sync-plugin-options-docs gate: the settings_write_ask_enabled option shipped in 0.4.0 without its generated README options block; generated now, with a changelog line under 0.6.1. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
Claude finished the security review
Reviewed all 28 changed files (df2bb9e), focused on the security-relevant surface: the measurement engine's subprocess spawning ( FindingsIMPORTANT — case-insensitive filesystem bypasses the settings-write ask checkpoint
This matters because SKILL.md documents the checkpoint as covering "any settings-surface write" unqualified ( Posted as an inline comment with a one-line Reviewed and judged not security findings (in scope but no issue)
Two P1 correctness findings and one P2 robustness finding already posted by the Codex reviewer on this PR (missing |
…-out (0.6.2) Two further PR #2932 review findings: --out now creates missing parent directories so a fresh audit's first snapshot cannot ENOENT away an expensive measurement, and systemToolsComparable now includes every mismatch the row records as a reason - binary path (same version, different install) and a moved Skills bucket under a matching listing both poison the predicate instead of warning while the delta publishes. Tests cover all three cases (28/28). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
|
Last security-reviewed head: |
…checkpoint (0.6.3) Security-review finding: macOS and Windows filesystems resolve .claude/Settings.json to the same file as the lowercase name, so the case-sensitive match let a differently-cased write bypass the checkpoint silently on exactly the platforms it supports. Both path regexes now carry the i flag and the user-global comparison is case-folded; case-variant regression cases added (12/12). Also annotates a shell-portability false positive on an embedded JavaScript regex in levers.test.sh. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
check-shell-portability reads the exemption on the hit line itself; the line-above comment cleared nothing and splitting the expression exposed a second hit. Both embedded-JavaScript regex lines in levers.test.sh now carry inline portability-ok markers; gate green locally (9 files clean). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
plugin-gate, hygiene, and the review lane on the previous head all died downloading actions from codeload.github.com (429/502/503) before any step ran; no content changed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
Second infrastructure re-trigger; the prior head's failures (plugin-gate, hygiene, skill-quality-gate, review lanes) all died downloading actions from codeload.github.com before any step ran. Content unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
Claude finished the security review
Re-review after push (
|
plugin-gate's cheat-sheet drift check finally ran to completion once the codeload outage eased and caught the missing row for the new skill - the one real failure hiding under the infrastructure noise. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
The hygiene exec-bit check (whole-repo, ungated) flags tracked shebang files recorded 100644; the five engine/hook scripts and their tests were written without the bit. Second real failure surfaced as the codeload outage eased. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
Claude finished the security review
Re-review after push (
|
The hygiene shellcheck lane (whole-repo, ungated) finally ran to completion and flagged 34 SC2015 sites (the compact `test && ok || fail` idiom) plus one SC2181 across the two test suites. Rewritten onto explicit assert helpers (assert_eq / assert_asks / assert_silent) and if/else - same 28+12 cases, all green; shellcheck exits clean with the repo rcfile. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
|
Claude finished @kyle-sexton's task in 48s —— View job Claude finished the security review
Re-review after push (
|
|
Claude finished @kyle-sexton's task in 3m 0s —— View job Claude code review
ScopeThis repository carries a I read the full diff against FindingsNone at this lane's high-signal bar. Specifically checked and found correct:
The |
|
Claude has reviewed this PR 2 times. The lane skips further automatic reviews after 5; deleting this comment resets the count. |
…-setup-xnwt0w # Conflicts: # docs/SKILL-CHEAT-SHEET.md
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
|
Claude finished @kyle-sexton's task in 1m 34s —— View job Claude code review
Skip gate: previous review still appliesThis trigger fired on a I diffed this PR's own contribution before and after the merge:
Per this lane's skip gate ("this head already has a successful review from this lane that still applies"): the prior code-review pass on this PR (posted 2026-08-17T16:21:29Z, covering the full 29-file diff at |
|
Claude finished @kyle-sexton's task in 2m 1s —— View job Claude finished the security review
Re-review after push (
|
…-setup-xnwt0w # Conflicts: # scripts/skill-leaf-name-registry.txt
|
I'll analyze this and get back to you. |
…, and gate the parity (#3235) docs/CLOUD-SESSIONS.md claimed `enabledPlugins` turns on the whole catalog so this repo dogfoods everything it publishes. It did not: ai-slop (#2892), context-budget (#2932) and improvement (#2985) each reached main catalogued but never enabled, leaving the catalog at 70 and the enabled set at 67. The failure was silent by construction — .claude/cloud-bootstrap.sh computes its install set from that same map, so a session came up green reporting "plugins 67 enabled" with those three plugins' skills absent and nothing naming what was missing. No existing detector covered this axis for this repo: claude-config's check-plugin-drift.sh resolves marketplaces through source.repo and SKIPs one that declares none, which is exactly this repo's relative `directory` source. Adds scripts/check-plugin-catalog-enablement.sh and its contract test, wired as the plugin-catalog-enablement-gate lane inside ci-status. It holds catalog and enabledPlugins equal in both directions, checks the alphabetical layout, and checks that cloud-bootstrap.sh's hardcoded marketplace_name still names the marketplace the settings file declares. An explicit `false` passes; an absent key does not. Review rounds folded in: a literal suffix strip replacing a sed pattern that interpolated the marketplace name (security review), the bootstrap identity check (Codex P2), and relocating the .github no-suite probe in affected-tests.test.sh off a path this change gave a real dependency.
Run /code-tidying:dissolve-comments (strict) over 26 gate scripts and
suites: 12 files, every edit COMMENT-ONLY. Revision SHAs, "first
revision" narration, incident lists and review attributions removed or
put in present tense; shared vocabulary ("the #1513 shape"), usage
headers and shellcheck directives kept. Census delta: -47 comment
lines, -2862 bytes. 55 affected suites pass; the other 3 fail only on
missing htmlhint/strace in this environment.
Removed narrative, verbatim:
- scripts/check-detector-eval-coverage.test.sh: "which are being backfilled in the same change set this gate landed in, and would otherwise make this suite red for someone else's in-flight work." / "it is the shape claude-code-plugins#4149 found already merged (an eval asserting a scope two checks out of date)" / "Sections marked N1 and N2 are regression tests for two defects an independent verifier reproduced against c7f71c3, where the gate had over-corrected ... Each N case fails against c7f71c3 and passes against this revision." / "Sections marked P1..P6 below are regression tests for defects an independent verifier reproduced against the gate as first committed (0273b5b), where it verified MENTION rather than coverage, ... Each of those cases FAILS against that revision and passes against this one; that is what makes them regression tests rather than restatements of current behavior." / "P7 is the same shape one revision later: the stopping rule had an extractor of its own ... the same defect class (a fix applied at one call site and missed at its sibling) the verdict scanner beside it had already been hardened against. Its first case FAILS against adcbb77; ... and pass on both sides." / "Each case below passed the first revision at rc 0 while covering nothing" / "Parseability was the only shape check there was." / "Reproduced against c7f71c3: ... Each case below exits 1 against that revision and 0 against this one." / "so the first revision's only guard (zero ids) never fired." / "The first revision's greedy `.*` kept only the last" / "found by a defeat attempt against the merged gate." / "all four were SILENT LOSSES under one revision or another ... the revision that returned the bare head ... the revision that kept the whole tail ... the revision that tried to split the difference ... Three attempts, three silent losses, each found by the round after the one that shipped it." / "which defeated the last-unclosed-`$(` attempt, and a backtick substitution, which it could not see at all." / "these are the spellings the rewritten delimiter matcher had to keep reading." / "An earlier version of this helper asserted only that the body's `emit error P9` was not counted, and that is the assertion that let four regressions through" / "These were written as `arms nothing` cases when the matcher used a character class and refused whatever fell outside it." / "the guard passes against the revision that has the bug." / "Every heredoc defect this scanner has had" / "Reproduced against c7f71c3: each shape below made a well-formed detector un-gateable, because the site counter and the id extractor disagreed" / "this guard was a byte-identical copy of that fixture until an audit caught it, a duplicate masquerading as a second dimension of coverage." / "The discovery half kept a greedy sed of its own ... while the verdict scanner beside it read that same line as two call sites correctly. Discovery now runs that same scanner" / "a per-row exit 2 used to leave the FIRST row's report" / "P6: trailing argv was silently ignored"
- scripts/check-plugin-catalog-enablement.sh: "Nothing enforced it. Three plugins reached main with a catalog entry and no `enabledPlugins` key: ai-slop (#2892), context-budget (#2932) and improvement (#2985), while the plugin PRs on either side of them (coupling #2913, overengineering #2961) remembered the settings entry." / "WHERE ENABLEMENT LIVES NOW" / "this file no longer mirrors the whole catalog (that mirror was writing one ..." / "The class that shipped three times." / "Raised as P2 by the Codex review on #3235."
- scripts/check-plugin-catalog-enablement.test.sh: "the class that actually shipped (a catalogued plugin with no enabledPlugins key, three times: #2892, #2932, #2985)" / "--- 2. The class that shipped: catalogued, enabled nowhere." / "what settings no longer mirrors" / "This is the post-migration shape:" / "Under the original `sed -n ...`" / "Raised as informational by the security review on #3235." / "Raised as P2 by the Codex review on #3235."
- scripts/check-purged-em-dashes.sh: "WHY (#2891)." / "each landed shard ... every shard so far has needed a rationale-withheld reviewer to catch clauses the automated passes waved through. Nothing then stops ... because no lane enforces the policy." / "29,649 em-dash prose lines across 1,074 tracked markdown files remained when this gate was written" / "#3342 needs a several-hundred-file allowlist; that is only affordable". .test.sh: "the narrower config #3342 asked for".
- scripts/check-skill-precompute-compose.test.sh: "The #3377 fix hoists `git rev-parse` into the parent, so the mode dispatch ... is now load-bearing"
- scripts/check-docs-naming.sh: "docs/ carried a mix of UPPER-KEBAB, lower-kebab, and mixed-case names for years, and every reference to a doc had to remember which spelling that one file used."
- scripts/lib/fixture-tree.sh: "not the three-variable spelling that was drifting through the suites."
- scripts/sync-resolve-convention-home.sh: "plugin-quality (the ADR 0018 pilot) is the first carrier."
- scripts/test-git-helpers.sh: "The exported GIT_DIR behind the real incident came from an ad-hoc tool invocation, not from a git hook: this repository has no git hook at any scope and core.hooksPath is unset everywhere." .test.sh: "That is what happened to #2827 -> #2830."
- scripts/generate-catalog.mjs: "One-way gates existed before this: ... but nothing compared CATEGORY_ORDER back to the document"
Intentional-removal: dissolve-comments pass; removed comment text is recorded above.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UpA569K1JgnKu1jQvRqMjL

No linked issue
Summary
Ships the new
context-budgetplugin: a session's fixed startup context payload made measurable per item on the consumer's own machine — including per-tool attribution of the built-in tool pools that/contextreports only as lump sums — with a lever catalogue whose every row carries an honesty category and official citation, a stamped read-only report, a guided fix path behind an explicit override, and a per-project measure-toggle-remeasure ledger. Also carries the phase-1 repo corrections. The topic slice that produced this (plan, findings, research corpus) was pruned before merge per the topic-docs contract; see the prune record below.Fix
skills/audit/scripts/measure.mjs): SDK-primary meter over the Agent SDK'sgetContextUsage()(exact integers, live tool enumeration), degrading to a version-aware parser of headless/contextoutput and then to a structured error with a remediation — never a wrong number. Per-tool attribution by bare-name-deny A/B differencing with an additivity check; enforced comparability rules (skill-listing signature, one mode, one binary); every record stamped with the measured binary path/version, mode, precision, and session kind.reference/levers.json): data rows with honesty category, category basis, measurement-resolved conditions, posture, detection, measurement route, exact emitted config, citations, verified date, and recheck trigger; net-negative/unverified rows structurally barred from the recommendable posture; no token figure ships (mechanically enforced bylevers.test.sh).reference/report.md): smart-zone framing, incomparable rows carry reasons instead of numbers, vendor weight reported as the honest floor.fixargument only; project settings editable per approved diff, user-global print-only, one lever at a time with a mandatory re-measure + ledger loop; PreToolUse hook returnspermissionDecision: "ask"on settings-surface writes (fail-open,settings_write_ask_enabledkill switch).auditleaf-name owner set,state-key.shsync cluster..cmdshim spawning through the shell with.exepreferred in resolution.Verification
check-skill(0 errors),check-evals-quality(0 warnings), changelog parity (+ bump vs base), leaf-name registry, state-key sync, cross-plugin drift, portability, hook exec-form, markdownlint,validate-plugin-contracts, contract-slice prune gate./context; per-tool deltas reproduced (Workflow−7,900,Artifact−4,470,SendUserFile−1,066,ReportFindings−821) and exactly additive (14,257 = sum of parts); degradation rungs exercised (cli-parse fallback, exit-3 structured error).ToolSearchanti-lever, now a catalogue caveat); PreToolUseaskmeasured to fire and block underbypassPermissionsin headless mode; HTTP MCP tools measured deferred at v2.1.232 in a dedicatedMCP tools (deferred)category (upstream [BUG] Tool Search (ENABLE_TOOL_SEARCH) does not defer HTTP/Streamable HTTP MCP tools — 120K tokens loaded upfront on every session anthropics/claude-code#40314's upfront loading does not reproduce), with the engine's headline semantics widened to match.Contract (topic-docs prune record)
Pre-prune commit SHA:
0e3cbb76f048b14448dd68048c2795aac61a2560— the last commit holdingdocs/topics/context-budget/(PLAN.md, FINDINGS.md incl. the post-acceptance probe record, the nine-run research corpus with per-claim citations, the session handoff). Best-effort retrieval while GitHub retains the object:gh api "repos/melodic-software/claude-code-plugins/contents/docs/topics/context-budget/PLAN.md?ref=0e3cbb76f048b14448dd68048c2795aac61a2560".Where durable outcomes graduated (the load-bearing record):
plugins/context-budget/skills/audit/reference/(engine.md,levers.json,report.md) — shipped, permanent.count_tokensbilling, with probe designs) → context-budget: two open measurements — interactive-session deferral eligibility; deferred-tool count_tokens billing #2954 via the work-item tracker seam._typos.tomlresearch exclusion was reverted with the corpus it excluded — the managed synced copy is back to canonical.Approved PLAN digest (full text at the pre-prune SHA)
Goal: ship a
context-budgetplugin whose single skill/context-budget:auditmakes a session's fixed startup payload measurable per item, explains each contributor in operator terms, and — behind an explicit override — applies the trims the operator approves. Novel capability: per-tool attribution of the built-in tool pools, by A/B differencing (compositional, measured).Constraints held: cite-never-transcribe (no token figure/key inventory/threshold ships as skill content); every lever carries an honesty category or is not offered; settings writes gated by a PreToolUse
ask(checkpoint, not guarantee);~/.claude/settings.jsonprint-only; persistent config emitspermissions.deny(nodisallowedToolssettings key exists); measured binary pinned and stamped; report leads with reclaimed reasoning space.Phases: 0 (blockers:
getContextUsage()probed working;skillOverridesdocumented) ✔ · 1 (five repo corrections) ✔ · 2 (measurement engine, 0.1.0) ✔ · 3 (lever catalogue as data rows, 0.2.0) ✔ · 4 (report contract, 0.3.0) ✔ · 5 (guided fix path + ask hook, 0.4.0) ✔ · 6 (evals + acceptance gate, 0.5.0/0.5.1, verifier ACCEPT) ✔ · post-acceptance closeout (probes + shakedown, 0.6.0/0.6.1) ✔.Acceptance criteria (all verified): ranked measured per-item attribution with version stamp; every lever categorized + cited; state-keyed before/after ledger; zero transcribed research figures in the shipped plugin; honest structured degradation;
/doctorrouted to, never reimplemented.Related
docs/conventions/topic-docs/README.md— the contract-slice lifecycle this PR's prune commit follows.🤖 Generated with Claude Code
https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm