Skip to content

feat(context-budget): ship the startup-payload measurement plugin (0.6.0) - #2932

Merged
kyle-sexton merged 25 commits into
mainfrom
claude/context-window-setup-xnwt0w
Aug 17, 2026
Merged

kyle-sexton merged 25 commits into
mainfrom
claude/context-window-setup-xnwt0w

Conversation

@kyle-sexton

@kyle-sexton kyle-sexton commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

No linked issue

Summary

Ships the new context-budget plugin: a session's fixed startup context payload made measurable per item on the consumer's own machine — including per-tool attribution of the built-in tool pools that /context reports only as lump sums — with a lever catalogue whose every row carries an honesty category and official citation, a stamped read-only report, a guided fix path behind an explicit override, and a per-project measure-toggle-remeasure ledger. Also carries the phase-1 repo corrections. The topic slice that produced this (plan, findings, research corpus) was pruned before merge per the topic-docs contract; see the prune record below.

Fix

  • Measurement engine (skills/audit/scripts/measure.mjs): SDK-primary meter over the Agent SDK's getContextUsage() (exact integers, live tool enumeration), degrading to a version-aware parser of headless /context output and then to a structured error with a remediation — never a wrong number. Per-tool attribution by bare-name-deny A/B differencing with an additivity check; enforced comparability rules (skill-listing signature, one mode, one binary); every record stamped with the measured binary path/version, mode, precision, and session kind.
  • Lever catalogue (reference/levers.json): data rows with honesty category, category basis, measurement-resolved conditions, posture, detection, measurement route, exact emitted config, citations, verified date, and recheck trigger; net-negative/unverified rows structurally barred from the recommendable posture; no token figure ships (mechanically enforced by levers.test.sh).
  • Report contract (reference/report.md): smart-zone framing, incomparable rows carry reasons instead of numbers, vendor weight reported as the honest floor.
  • Fix path + checkpoint: explicit fix argument only; project settings editable per approved diff, user-global print-only, one lever at a time with a mandatory re-measure + ledger loop; PreToolUse hook returns permissionDecision: "ask" on settings-surface writes (fail-open, settings_write_ask_enabled kill switch).
  • Registration: marketplace entry, CATALOG row, audit leaf-name owner set, state-key.sh sync cluster.
  • Review fixes (0.6.1): collision-safe ledger run IDs; Windows .cmd shim spawning through the shell with .exe preferred in resolution.

Verification

  • Hermetic suites green: engine 25/25, catalogue 1/1 (validated against a seeded violation), hook 10/10.
  • Repo gates green: check-skill (0 errors), check-evals-quality (0 warnings), changelog parity (+ bump vs base), leaf-name registry, state-key sync, cross-plugin drift, portability, hook exec-form, markdownlint, validate-plugin-contracts, contract-slice prune gate.
  • Live verification against the pinned v2.1.232 binary: exact integers matching /context; per-tool deltas reproduced (Workflow −7,900, Artifact −4,470, SendUserFile −1,066, ReportFindings −821) and exactly additive (14,257 = sum of parts); degradation rungs exercised (cli-parse fallback, exit-3 structured error).
  • Fresh-context acceptance verifier over the six PLAN acceptance criteria: OVERALL: ACCEPT, all criteria PASS; its one actionable finding fixed in 0.5.1.
  • Post-acceptance empirical probes: end-to-end shakedown of the shipped skill (report + ledger produced per contract; surfaced the measured deny-ToolSearch anti-lever, now a catalogue caveat); PreToolUse ask measured to fire and block under bypassPermissions in headless mode; HTTP MCP tools measured deferred at v2.1.232 in a dedicated MCP tools (deferred) category (upstream [BUG] Tool Search (ENABLE_TOOL_SEARCH) does not defer HTTP/Streamable HTTP MCP tools — 120K tokens loaded upfront on every session anthropics/claude-code#40314's upfront loading does not reproduce), with the engine's headline semantics widened to match.

Contract (topic-docs prune record)

Pre-prune commit SHA: 0e3cbb76f048b14448dd68048c2795aac61a2560 — the last commit holding docs/topics/context-budget/ (PLAN.md, FINDINGS.md incl. the post-acceptance probe record, the nine-run research corpus with per-claim citations, the session handoff). Best-effort retrieval while GitHub retains the object: gh api "repos/melodic-software/claude-code-plugins/contents/docs/topics/context-budget/PLAN.md?ref=0e3cbb76f048b14448dd68048c2795aac61a2560".

Where durable outcomes graduated (the load-bearing record):

Approved PLAN digest (full text at the pre-prune SHA)

Goal: ship a context-budget plugin whose single skill /context-budget:audit makes a session's fixed startup payload measurable per item, explains each contributor in operator terms, and — behind an explicit override — applies the trims the operator approves. Novel capability: per-tool attribution of the built-in tool pools, by A/B differencing (compositional, measured).

Constraints held: cite-never-transcribe (no token figure/key inventory/threshold ships as skill content); every lever carries an honesty category or is not offered; settings writes gated by a PreToolUse ask (checkpoint, not guarantee); ~/.claude/settings.json print-only; persistent config emits permissions.deny (no disallowedTools settings key exists); measured binary pinned and stamped; report leads with reclaimed reasoning space.

Phases: 0 (blockers: getContextUsage() probed working; skillOverrides documented) ✔ · 1 (five repo corrections) ✔ · 2 (measurement engine, 0.1.0) ✔ · 3 (lever catalogue as data rows, 0.2.0) ✔ · 4 (report contract, 0.3.0) ✔ · 5 (guided fix path + ask hook, 0.4.0) ✔ · 6 (evals + acceptance gate, 0.5.0/0.5.1, verifier ACCEPT) ✔ · post-acceptance closeout (probes + shakedown, 0.6.0/0.6.1) ✔.

Acceptance criteria (all verified): ranked measured per-item attribution with version stamp; every lever categorized + cited; state-keyed before/after ledger; zero transcribed research figures in the shipped plugin; honest structured degradation; /doctor routed to, never reimplemented.

Related

🤖 Generated with Claude Code

https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm

claude added 11 commits August 17, 2026 04:28
Nine dispatched research runs plus first-hand measurement on CLI v2.1.232
establish the design for a new `context-budget` plugin whose skill measures
and trims a session's fixed startup payload.

Key results, all measured rather than asserted:

- Rule shape, not the setting, decides whether a tool schema ships. A bare
  tool name in a deny rule removes the definition from the request; a scoped
  rule is a runtime guard whose schema is still billed every turn.
- Deferral does not shrink the request, inverting the premise the source
  material's headline lever rests on.
- The skill listing is hard-capped, so disabling skills saves nothing while
  over the cap. Disabling 45 of 65 plugins moved the row by zero.
- Custom agents are not capped and do scale, but the deny form the sub-agents
  docs present for them leaves the description in the payload.
- `System tools` has skill-frontmatter tokens subtracted, so it is not
  comparable across runs whose skill listing differs.

Also records five corrections owed to this repository, including the
safe-mode clean-room claim and unhobble's CLAUDE_CODE_SIMPLE gotcha.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
…ing, split the two SIMPLE env vars

Second-eyes review over the locked brief. The /context headline excludes the
deferred row (35.19k computed vs 35.3k reported), so the Goal's '35.9k of a
35.3k-headline session' phrasing was incoherent and is restated precisely.
Binary inspection confirms CLAUDE_CODE_SIMPLE (--bare mode) and
CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT are two distinct env vars, which upgrades the
unhobble correction: its gotcha is wrong about documentation status AND
attributes prompt-stripping to the wrong var.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
Phase 0 (docs/topics/context-budget/PLAN.md): the Agent SDK's
getContextUsage() is real and probed live — exact integers matching the
CLI, but systemTools/deferredBuiltinTools arrive unpopulated, so the
engine is SDK-primary with A/B differencing for per-built-in-tool
attribution. skillOverrides is documented (settings reference + skills
page, fetched raw), with the plugin-skill carve-out; the fetch also
surfaced skillListingBudgetFraction and skillListingMaxDescChars as
documented levers.

Corrections applied:
- checks-and-sweep: safe-mode/CLAUDE_CONFIG_DIR are not clean rooms
  (dated correction with measured evidence)
- coverage-matrix S7: deferred-tool half now owned by context-budget,
  premise corrected (deferral does not shrink the request)
- permission-rule-hygiene: stale dated quote replaced with the page's
  current version-floor wording (fetched 2026-08-17); tightening gap
  recorded in the changelog
- claude-config 0.38.7: unhobble's CLAUDE_CODE_SIMPLE gotcha rewritten
  against the binary (documented, and prompt-stripping belongs to the
  sibling CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT); eval 8 regraded onto the
  scope-boundary reasoning

Filed #2895 (preload sentinel unsound as preload proof, follow-up to
#2338) and #2896 (skill-quality verb-contract mismatch check).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
… cloud handoff

The .work memory slice dies with the cloud container, and Phase 3's lever
catalogue needs the per-claim citations, so the nine research slices plus
the measurement/design/interview artifacts move to
docs/topics/context-budget/research/ (typos exclusion added: the corpus
quotes minified binary internals verbatim as evidence). The handoff file
is committed for the same reason, with a .work/handoffs/ mirror for the
standard local contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
….1.0

New plugin with single skill /context-budget:audit (read-only). The engine
(skills/audit/scripts/measure.mjs) is SDK-primary: exact structured context
usage over the Agent SDK with live tool enumeration, degrading to a
version-aware parser of headless /context output (display-rounded, refuses
loudly on format drift) and then to a structured error with a remediation -
never a wrong number. Per-tool attribution of the built-in tool pools by
bare-name-deny A/B differencing with an optional additivity verification;
enforced comparability rules (skill-listing signature, one mode, one binary
version); offline compare producing ledger rows; per-project state-keyed
ledger (one file per run plus an appended history line). Every record stamped
with the measured binary path/version, mode, precision, and session kind.

Registered in the marketplace, catalog, audit leaf-name owner set, and the
state-key sync cluster. Hermetic test suite covers the parser traps, compare
comparability, and ledger retention. Verified live against a pinned binary:
per-tool deltas reproduce and pass the additivity check.

Phase 2 of docs/topics/context-budget/PLAN.md; Phases 3-6 remain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
….2.0)

skills/audit/reference/levers.json: 19 levers, each a data row carrying its
honesty category (six-term vocabulary with an explicit request-vs-context-
window dual-ledger distinction), category basis, measurement-resolved
conditions, posture, detection, measurement route, exact emitted config,
official citations, verified date, and recheck trigger. Net-negative and
unverified rows are structurally barred from the recommendable posture; the
memory-file and hook lanes are route-outs in the catalogue meta.

levers.test.sh makes the honesty rules mechanical, including a scan that no
row ships a token figure (it caught one authoring violation, which was fixed
in the row, not the test). SKILL.md gains the lever-presentation step;
README documents the catalogue; changelog and version bump per convention.

Phase 3 of docs/topics/context-budget/PLAN.md; Phases 4-6 remain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
skills/audit/reference/report.md fixes the audit's default deliverable:
stamped header (binary, mode, precision, session kind, model, cwd, time),
smart-zone headline leading with reclaimed reasoning space (context-guard
zone framing presence-gated), measured category totals, ranked per-tool
attribution where incomparable rows carry reasons instead of numbers and
unmeasured tools are listed rather than omitted, lever findings grouped by
honesty category with citations and emitted config, route-outs, and a
degradations section. Reports persist one file per run under the keyed data
dir. The what-the-report-never-does list bars external-arithmetic
reconciliation, bucket merging, and any apply.

Phase 4 of docs/topics/context-budget/PLAN.md; Phases 5-6 remain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
…int (0.4.0)

The fix path runs only on the explicit `fix` argument (verb-contract
override): per-lever walkthrough in ranked-report order over
recommendable-on-fit catalogue rows whose conditions this audit resolved by
measurement, one lever at a time with a mandatory apply -> re-measure ->
compare -> ledger loop. Write posture splits by scope: project settings
editable after per-diff approval; user-global ~/.claude/settings.json
print-only (never written -- auto mode's classifier can approve a protected-
path write with no human, so the prompt cannot be the protection); managed
policy never targeted; env levers printed.

hooks/settings-write-ask.mjs (PreToolUse, exec-form node for Windows
safety): permissionDecision "ask" on any Write/Edit targeting a settings
surface, so auto mode prompts instead of silently approving. Documented as a
checkpoint, not a guarantee. Fail-open; settings_write_ask_enabled
userConfig kill switch via the hook mirror; hermetic contract test.

Phase 5 of docs/topics/context-budget/PLAN.md; Phase 6 remains.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
…(0.5.0)

Three fix-path eval cases join the six measurement-honesty cases: mutation
only on the explicit fix argument, user-global print-only with the
auto-mode classifier caveat as the stated reason, and one-lever-at-a-time
apply -> re-measure -> ledger with batches refused on attribution grounds.
Evals-quality gate passes with zero warnings.

Mechanical acceptance sweeps recorded in the changelog and PLAN: the
shipped plugin greps clean of every research-run figure, catalogue and hook
contract tests pass, and the skill-layout gate reports zero errors. A
fresh-context acceptance verifier over the six acceptance criteria is
running; its verdict and any fixes close the phase in a follow-up commit.

Phase 6 of docs/topics/context-budget/PLAN.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
… shipped (0.5.1)

The fresh-context acceptance verifier returned OVERALL: ACCEPT (all six
criteria PASS, all commanded runs green). Its one actionable finding is
fixed: the catalogue test's token-figure scan now fails any k-suffixed
figure in a lever row outright and catches plain integers adjacent to the
word token in either order, verified against a seeded violation. PLAN marks
Phase 6 shipped with the verdict recorded; the SKILL.md 227-line soft
warning is accepted as load-bearing fix-path posture.

Closes the build phases of docs/topics/context-budget/PLAN.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
…0.6.0)

First end-to-end shakedown of the shipped skill against this container
(state key -> baseline -> 5-tool attribution -> report -> ledger): deltas
reproduced, four-lever additivity exact, report persisted per contract. It
surfaced a measured anti-lever - denying the tool-search tool forces the
entire deferred pool upfront - now a deny-bare-tool catalogue caveat.

Two fresh-context probes resolved open questions for v2.1.232 headless:
a PreToolUse ask fires and blocks even under bypassPermissions (surfacing
as a tool error carrying the reason; interactive unmeasured, documented as
such), and HTTP MCP tools measure DEFERRED in a dedicated "MCP tools
(deferred)" category - upstream #40314's upfront loading does not
reproduce. Engine cli-parse headline now excludes every "(deferred)"
category to match; hook header, SKILL.md, engine.md, and the catalogue
updated. FINDINGS gains the post-acceptance probe record (count_tokens
billing probe blocked here - no API credential - with the two-call design
recorded); PLAN gains the closeout section.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b1e4d0a292

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/context-budget/skills/audit/scripts/measure.mjs Outdated
Comment thread plugins/context-budget/skills/audit/scripts/measure.mjs Outdated
Comment thread _typos.toml Outdated
Comment thread plugins/context-budget/skills/audit/scripts/measure.mjs Outdated
Comment thread plugins/context-budget/skills/audit/scripts/measure.mjs Outdated
claude added 2 commits August 17, 2026 15:02
…ning (0.6.1)

Two PR #2932 review findings fixed: a same-second rerun of the same lever
now collides into a numbered-suffix run file instead of overwriting the
earlier point (one-file-per-run held only by luck; test added), and binary
resolution prefers claude.exe over claude.cmd with .cmd/.bat shims executed
through the shell, since Node cannot spawn Windows command shims directly.
The Windows path is an honest manual-verification gap (no Windows hardware
in this environment).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
…-setup-xnwt0w

# Conflicts:
#	scripts/skill-leaf-name-registry.txt
claude added 2 commits August 17, 2026 15:08
…lice

Per the topic-docs contract-slice lifecycle (prune with pointer): durable
method, citations, and honesty rules live in the shipped plugin
(plugins/context-budget/skills/audit/reference/); the two remaining open
measurements graduate to the work-item tracker as #2954; the evidence
record stays retrievable via this PR's pre-prune SHA, named in the PR
body. The _typos.toml research exclusion is reverted with the corpus it
excluded, returning the managed synced copy to canonical (also resolves
the PR review's P1). References into the slice from changelogs and the
grandfathered context-engineering-claude-5 slice now point at the durable
homes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
sync-plugin-options-docs gate: the settings_write_ask_enabled option shipped
in 0.4.0 without its generated README options block; generated now, with a
changelog line under 0.6.1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
@claude

claude Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished the security review

  • Gather context (PR diff vs origin/main, custom instructions)
  • Run /review:security-review per its instructions
  • Post findings

Reviewed all 28 changed files (df2bb9e), focused on the security-relevant surface: the measurement engine's subprocess spawning (measure.mjs), the PreToolUse settings-write checkpoint hook (settings-write-ask.mjs + hooks.json), and the path-sanitizing state-key derivation (lib/state-key.sh). No GitHub Actions workflow files are touched by this PR, so that class of risk is out of scope here.

Findings

IMPORTANT — case-insensitive filesystem bypasses the settings-write ask checkpoint
plugins/context-budget/hooks/settings-write-ask.mjs#L41-L42

isSettings is matched with case-sensitive regexes. On macOS (default APFS/HFS+) and Windows (NTFS) — the two non-Linux platforms this hook is explicitly written to support (see the backslash-normalization a few lines up) — those filesystems are case-insensitive-but-preserving. A Write/Edit/MultiEdit/NotebookEdit call targeting .claude/Settings.json, .claude/settings.LOCAL.json, or Managed-Settings.json resolves to the same file on disk as the canonical lowercase name, but fails the regex, so the hook exits 0 with no output and the write proceeds with no ask prompt — silently, even in auto mode.

This matters because SKILL.md documents the checkpoint as covering "any settings-surface write" unqualified (skills/audit/SKILL.md:210), so an operator reading that has a false sense of coverage. It's also concretely triggerable, not just theoretical: a model choosing (on its own, or steered by injected content from untrusted repo/CLAUDE.md text) a differently-cased path for a settings edit defeats the entire purpose of this PR's core safety control — invisibly, with no degraded-mode signal.

Posted as an inline comment with a one-line suggestion fix (add the i flag to both regexes) and a claude.ai/code deep link to apply it plus add a case-variant regression case to settings-write-ask.test.sh.

Reviewed and judged not security findings (in scope but no issue)

  • measure.mjs's spawnBinary shell-mode branch for Windows .cmd/.bat shims — this is the correct mitigation for Node's known unsafe-implicit-shell-invocation class of issue (explicit shell: true gets Node's argument-escaping), not a regression.
  • --tools/--deny values reaching spawnSync — sourced from the trusted live SDK tool list or the invoking operator's own CLI argument, never from a lower-trust boundary.
  • lib/state-key.sh's git-remote-derived identity — already defends path traversal via allow-list regex + hash fallback, with test coverage (state-key.test.sh).
  • .claude-plugin/plugin.json / marketplace.json additions — no permission widening, no new trust grants.
  • No GitHub Actions workflow files changed in this PR.

Two P1 correctness findings and one P2 robustness finding already posted by the Codex reviewer on this PR (missing mkdir before the first snapshot write, an incomplete systemToolsComparable predicate, and ledger run-ID collisions) are logic/robustness issues rather than security issues, and the run-ID one already appears fixed in the current HEAD — left to the code-review lane per this skill's scope (security findings only).

View job run

…-out (0.6.2)

Two further PR #2932 review findings: --out now creates missing parent
directories so a fresh audit's first snapshot cannot ENOENT away an
expensive measurement, and systemToolsComparable now includes every
mismatch the row records as a reason - binary path (same version,
different install) and a moved Skills bucket under a matching listing both
poison the predicate instead of warning while the delta publishes. Tests
cover all three cases (28/28).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
Comment thread plugins/context-budget/hooks/settings-write-ask.mjs Outdated
@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Last security-reviewed head: 1e0f384bf83000b80d315af0598172c7955f25fa. On the next push, the relevance gate compares only the commits since this SHA; delete this comment to force a full re-review.

claude added 4 commits August 17, 2026 15:17
…checkpoint (0.6.3)

Security-review finding: macOS and Windows filesystems resolve
.claude/Settings.json to the same file as the lowercase name, so the
case-sensitive match let a differently-cased write bypass the checkpoint
silently on exactly the platforms it supports. Both path regexes now carry
the i flag and the user-global comparison is case-folded; case-variant
regression cases added (12/12). Also annotates a shell-portability false
positive on an embedded JavaScript regex in levers.test.sh.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
check-shell-portability reads the exemption on the hit line itself; the
line-above comment cleared nothing and splitting the expression exposed a
second hit. Both embedded-JavaScript regex lines in levers.test.sh now
carry inline portability-ok markers; gate green locally (9 files clean).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
plugin-gate, hygiene, and the review lane on the previous head all died
downloading actions from codeload.github.com (429/502/503) before any
step ran; no content changed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
Second infrastructure re-trigger; the prior head's failures (plugin-gate,
hygiene, skill-quality-gate, review lanes) all died downloading actions
from codeload.github.com before any step ran. Content unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
@claude

claude Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished the security review

  • Gather context (PR diff vs origin/main, custom instructions)
  • Run /review:security-review per its instructions
  • Post findings

Re-review after push (df2bb9e8 → aed4d06)

The previously security-reviewed head was df2bb9e8. Five commits landed since then; two touch code, two are no-op CI re-triggers, and one is a portability-annotation fix with no security relevance. I re-reviewed the delta rather than skipping, since it includes a direct fix to the finding from the prior pass.

Prior IMPORTANT finding — now fixed. plugins/context-budget/hooks/settings-write-ask.mjs#L45-L46 — commit 073435d adds the /i flag to both isSettings regexes and switches the userGlobal comparison to .toLowerCase() (L50-L51), closing the case-insensitive-filesystem bypass on macOS/Windows. A regression test (settings-write-ask.test.sh:93-103) now asserts case-variant paths (Settings.json, settings.LOCAL.json) still trigger the ask checkpoint. Verified correct — no residual gap.

New changes since then — no new security issues.

  • plugins/context-budget/skills/audit/scripts/measure.mjs (commit c756338): emit() now runs mkdirSync(dirname(resolve(outFile)), { recursive: true }) before writeFileSync (L650-L657 region). outFile originates from the operator's own --out CLI argument (same trust level as the rest of the invoking command line), so recursive directory creation here isn't attacker-influenced path traversal — it's a robustness fix, not a security-relevant change.
  • systemToolsComparable in the same file now also requires binary?.path equality and skillTokensMatch. This is a correctness/data-integrity fix (prevents an incomparable run from publishing a misleading delta), not a security control — no trust boundary involved.
  • plugins/context-budget/skills/audit/scripts/levers.test.sh (commit 932e147): repositioned portability-ok lint-exemption comments onto the flagged lines. No behavior change.
  • The two chore: re-trigger CI commits (c8f6b79, aed4d06) contain no file changes.

No GitHub Actions workflow files are touched anywhere in this PR, so that risk class remains out of scope. No new injection surfaces, credential handling, or permission-widening changes were introduced by this delta.

Overall verdict

No outstanding security findings on the current head (aed4d06). The one prior finding is verified fixed with regression coverage.

View job run

claude added 2 commits August 17, 2026 16:06
plugin-gate's cheat-sheet drift check finally ran to completion once the
codeload outage eased and caught the missing row for the new skill - the
one real failure hiding under the infrastructure noise.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
The hygiene exec-bit check (whole-repo, ungated) flags tracked shebang
files recorded 100644; the five engine/hook scripts and their tests were
written without the bit. Second real failure surfaced as the codeload
outage eased.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
@claude

claude Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished the security review

  • Gather context (PR diff vs origin/main, custom instructions)
  • Run /review:security-review per its instructions
  • Post findings

Re-review after push (aed4d06 → 6360252)

Two commits landed since the last security-reviewed head (aed4d0619dd30813c2692982f31f85332e23caec, which itself confirmed the prior settings-write-ask.mjs case-insensitivity fix with no residual gap):

  • bb56598 — docs: regenerate the skill cheat sheet with /context-budget:audit: adds one table row to docs/SKILL-CHEAT-SHEET.md. No code, no security relevance.
  • 6360252 — fix(context-budget): set the executable bit on shebang scripts: mode-only change (100644 → 100755) on five files (hooks/settings-write-ask.mjs, hooks/settings-write-ask.test.sh, skills/audit/scripts/levers.test.sh, skills/audit/scripts/measure.mjs, skills/audit/scripts/measure.test.sh) — zero content diff. I confirmed hooks/hooks.json invokes the hook explicitly via node ${CLAUDE_PLUGIN_ROOT}/hooks/settings-write-ask.mjs, not by direct execution, so this mode bit is inert for the actual hook trigger path — a portability fix for direct CLI invocation, not a security-relevant behavior change.

No new injection surfaces, credential handling, permission-widening, or trust-boundary changes in this delta. No GitHub Actions workflow files are touched anywhere in this PR.

Overall verdict

No outstanding security findings on the current head (6360252). The one substantive finding from the first pass (case-insensitive-filesystem bypass of the settings-write ask checkpoint on macOS/Windows) was fixed in 073435d and verified with regression coverage (settings-write-ask.test.sh:93-103) in the prior re-review; nothing in this delta touches that code's logic.

View job run

The hygiene shellcheck lane (whole-repo, ungated) finally ran to
completion and flagged 34 SC2015 sites (the compact `test && ok || fail`
idiom) plus one SC2181 across the two test suites. Rewritten onto explicit
assert helpers (assert_eq / assert_asks / assert_silent) and if/else -
same 28+12 cases, all green; shellcheck exits clean with the repo rcfile.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pwn8XVbDojP9o18ZohEuXm
@claude

claude Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 48s —— View job


Claude finished the security review

  • Gather context (PR diff vs origin/main, custom instructions)
  • Run /review:security-review per its instructions
  • Post findings

Re-review after push (6360252 → 382c1ae)

One commit landed since the last security-reviewed head (636025229f24bbd86093f91ceccba0f52fad91c7):

The diff is a mechanical shellcheck cleanup: &&/|| conditional chains are rewritten as if/else blocks, and repeated assertion patterns are factored into new local helpers (assert_asks, assert_silent, assert_eq). No assertions were removed, weakened, or had their expected values changed — I diffed each rewritten block against its original and confirmed the same conditions are checked (including the case-insensitive-path regression cases from the prior fix, at settings-write-ask.test.sh:93-103, which are preserved verbatim in behavior). No production code (measure.mjs, settings-write-ask.mjs, state-key.sh, levers.json, hooks config) is touched by this commit.

No new injection surfaces, credential handling, permission-widening, or trust-boundary changes. No GitHub Actions workflow files are touched anywhere in this PR.

Overall verdict

No outstanding security findings on the current head (382c1ae). The one substantive finding from the first pass (case-insensitive-filesystem bypass of the settings-write ask checkpoint on macOS/Windows) was fixed in 073435d and verified with regression coverage in the prior re-review; this delta only reformats that same test file's control flow for shellcheck compliance without altering its assertions.

@claude

claude Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 3m 0s —— View job


Claude code review

  • Gather context (PR diff vs origin/main, custom instructions)
  • Invoke /review:code-review and follow its instructions exactly
  • Post findings

Scope

This repository carries a claude-security-review.yml workflow, so per REVIEW.md's lane split this pass excludes security findings (already covered exhaustively by the security-review lane's four passes on this PR, ending with "no outstanding security findings on 6360252"). This is the code-review lane: correctness, design, conventions, error handling, observability, tests, documentation.

I read the full diff against origin/main (2828/-18 across 29 files) and reviewed both halves: the new context-budget plugin (measure.mjs, the settings-write-ask hook + tests, state-key.sh, the lever catalogue + its contract test, SKILL.md, report.md/engine.md, evals) and the phase-1 repo corrections (permission-rule-hygiene doc refresh, unhobble's CLAUDE_CODE_SIMPLE gotcha rewrite + eval update, marketplace/CATALOG/cheat-sheet/leaf-name-registry/sync-state-key registration).

Findings

None at this lane's high-signal bar. Specifically checked and found correct:

  • Registration consistency: marketplace.json, docs/CATALOG.md, docs/SKILL-CHEAT-SHEET.md, scripts/skill-leaf-name-registry.txt (audit leaf-name owner set), and scripts/sync-state-key.sh (context-budget added to the state-key.sh sync cluster) are all mutually consistent, and plugins/context-budget/lib/state-key.sh is byte-identical to the canonical claude-config and claude-memory copies.
  • measure.mjs's degradation ladder and comparability logic: the previously-flagged Codex P1s (missing mkdirSync before the first snapshot write — skills/audit/scripts/measure.mjs#L651-L657; the incomplete systemToolsComparable predicate — #L519-L522) and the P2 ledger run-ID collision (#L624-L629) are all fixed on this head, each with a corresponding regression case in measure.test.sh.
  • Test coverage: measure.test.sh (parser edge cases including the unredirected-stdin warning trap, comparability poisoning on every mismatch axis, ledger collision handling), levers.test.sh (mechanically enforces the honesty-category/citation/posture/no-token-figure contract, verified against a seeded violation per the PR description), and settings-write-ask.test.sh (case-variant and backslash-path regression cases for the fixed checkpoint bypass) all match the behavior in the code they cover.
  • Documentation accuracy: SKILL.md, README.md, reference/engine.md, and reference/report.md are internally consistent with each other and with measure.mjs's actual record shapes and CLI surface. The phase-1 corrections (permission-rule-hygiene's auto-mode citation refresh, unhobble's CLAUDE_CODE_SIMPLE vs CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT correction) are well-cited and the unhobble eval (evals.json case 8) was updated to match the corrected claim rather than left stale.
  • CHANGELOG/version parity: plugin.json (0.6.3) and CHANGELOG.md agree, with each fix-round entry matching what's actually in the diff.

The _typos.toml question the Codex reviewer raised was resolved by the prune commit (the managed file is back to byte-identical with the synced copy, per the author's reply) rather than needing a code change — nothing left to flag there.
· branch claude/context-window-setup-xnwt0w

@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Claude has reviewed this PR 2 times. The lane skips further automatic reviews after 5; deleting this comment resets the count.

…-setup-xnwt0w

# Conflicts:
#	docs/SKILL-CHEAT-SHEET.md
@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

@claude

claude Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 34s —— View job


Claude code review

  • Gather context (PR diff vs origin/main, custom instructions)
  • Invoke /review:code-review and follow its instructions exactly
  • Post findings

Skip gate: previous review still applies

This trigger fired on a synchronize event whose new head, 1e0f384, is a merge commit bringing origin/main (which had advanced with unrelated PRs, e.g. the new coupling plugin) into this branch.

I diffed this PR's own contribution before and after the merge:

  • git diff 6f864aa..1e0f384 (HEAD vs. the origin/main commit it merged in) — 29 files changed, +2828/-18, identical file list and line counts to this PR's original changeset.
  • git diff 382c1ae..1e0f384 scoped to every path this PR touches (plugins/context-budget/**, plugins/claude-config/**, docs/conventions/permission-rule-hygiene/**, docs/CATALOG.md, docs/SKILL-CHEAT-SHEET.md, docs/topics/context-engineering-claude-5/**, scripts/skill-leaf-name-registry.txt, scripts/sync-state-key.sh) — the only lines that moved are two shared list files (docs/CATALOG.md, docs/SKILL-CHEAT-SHEET.md) picking up an adjacent, unrelated row for the coupling plugin next to this PR's own row. This PR's own lines are byte-for-byte unchanged.
  • git diff 382c1ae..1e0f384 --stat -- plugins/context-budget — empty. No context-budget file changed at all.

Per this lane's skip gate ("this head already has a successful review from this lane that still applies"): the prior code-review pass on this PR (posted 2026-08-17T16:21:29Z, covering the full 29-file diff at 382c1ae with no findings at this lane's high-signal bar) covers the identical content still present at 1e0f384. Nothing new to review; no additional findings to post.
· branch claude/context-window-setup-xnwt0w

@claude

claude Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 1s —— View job


Claude finished the security review

  • Gather context (PR diff vs origin/main, custom instructions)
  • Run /review:security-review per its instructions
  • Post findings

Re-review after push (382c1ae → 1e0f384)

The previously security-reviewed head was 382c1aebbf1609051900e13ed200e98039a2c615. One commit landed since then:

  • 1e0f384 — a merge of origin/main into this branch, bringing in three unrelated upstream commits (6f864aa new coupling plugin, fc6e0be test-helper refactor for scripts/check-*-portability.test.sh, 977bf10 standards sync touching .github/standards/runner-policy/policy.json and .claude/settings.json).

I confirmed via gh pr diff 2932 that the PR's actual changed-file set is still exactly the same 29 files as the prior review, and diffed 382c1ae..1e0f384 scoped to every path this PR touches (plugins/context-budget/, plugins/claude-config/, docs/, scripts/, .claude-plugin/) — zero lines changed. The merge only pulled in commits already on main that don't intersect this PR's diff at all; nothing new to review. This applies per the skill's skip gate ("this head already has a successful security review that still applies").

Overall verdict

No outstanding security findings on the current head (1e0f384). The one substantive finding from the first pass — a case-insensitive-filesystem bypass of the settings-write ask checkpoint on macOS/Windows in plugins/context-budget/hooks/settings-write-ask.mjs — was fixed in 073435d and verified with regression coverage (settings-write-ask.test.sh:93-103) in an earlier re-review; this merge commit doesn't touch that file or any other security-relevant surface in the PR.

…-setup-xnwt0w

# Conflicts:
#	scripts/skill-leaf-name-registry.txt
@claude

claude Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Claude Code is working…

I'll analyze this and get back to you.

View job run

@kyle-sexton
kyle-sexton merged commit db32872 into main Aug 17, 2026
40 checks passed
@kyle-sexton
kyle-sexton deleted the claude/context-window-setup-xnwt0w branch August 17, 2026 17:25
kyle-sexton added a commit that referenced this pull request Aug 23, 2026
…, and gate the parity (#3235)

docs/CLOUD-SESSIONS.md claimed `enabledPlugins` turns on the whole catalog so
this repo dogfoods everything it publishes. It did not: ai-slop (#2892),
context-budget (#2932) and improvement (#2985) each reached main catalogued but
never enabled, leaving the catalog at 70 and the enabled set at 67.

The failure was silent by construction — .claude/cloud-bootstrap.sh computes its
install set from that same map, so a session came up green reporting "plugins 67
enabled" with those three plugins' skills absent and nothing naming what was
missing. No existing detector covered this axis for this repo: claude-config's
check-plugin-drift.sh resolves marketplaces through source.repo and SKIPs one
that declares none, which is exactly this repo's relative `directory` source.

Adds scripts/check-plugin-catalog-enablement.sh and its contract test, wired as
the plugin-catalog-enablement-gate lane inside ci-status. It holds catalog and
enabledPlugins equal in both directions, checks the alphabetical layout, and
checks that cloud-bootstrap.sh's hardcoded marketplace_name still names the
marketplace the settings file declares. An explicit `false` passes; an absent
key does not.

Review rounds folded in: a literal suffix strip replacing a sed pattern that
interpolated the marketplace name (security review), the bootstrap identity
check (Codex P2), and relocating the .github no-suite probe in
affected-tests.test.sh off a path this change gave a real dependency.
kyle-sexton added a commit that referenced this pull request Sep 25, 2026
Run /code-tidying:dissolve-comments (strict) over 26 gate scripts and
suites: 12 files, every edit COMMENT-ONLY. Revision SHAs, "first
revision" narration, incident lists and review attributions removed or
put in present tense; shared vocabulary ("the #1513 shape"), usage
headers and shellcheck directives kept. Census delta: -47 comment
lines, -2862 bytes. 55 affected suites pass; the other 3 fail only on
missing htmlhint/strace in this environment.

Removed narrative, verbatim:

- scripts/check-detector-eval-coverage.test.sh: "which are being backfilled in the same change set this gate landed in, and would otherwise make this suite red for someone else's in-flight work." / "it is the shape claude-code-plugins#4149 found already merged (an eval asserting a scope two checks out of date)" / "Sections marked N1 and N2 are regression tests for two defects an independent verifier reproduced against c7f71c3, where the gate had over-corrected ... Each N case fails against c7f71c3 and passes against this revision." / "Sections marked P1..P6 below are regression tests for defects an independent verifier reproduced against the gate as first committed (0273b5b), where it verified MENTION rather than coverage, ... Each of those cases FAILS against that revision and passes against this one; that is what makes them regression tests rather than restatements of current behavior." / "P7 is the same shape one revision later: the stopping rule had an extractor of its own ... the same defect class (a fix applied at one call site and missed at its sibling) the verdict scanner beside it had already been hardened against. Its first case FAILS against adcbb77; ... and pass on both sides." / "Each case below passed the first revision at rc 0 while covering nothing" / "Parseability was the only shape check there was." / "Reproduced against c7f71c3: ... Each case below exits 1 against that revision and 0 against this one." / "so the first revision's only guard (zero ids) never fired." / "The first revision's greedy `.*` kept only the last" / "found by a defeat attempt against the merged gate." / "all four were SILENT LOSSES under one revision or another ... the revision that returned the bare head ... the revision that kept the whole tail ... the revision that tried to split the difference ... Three attempts, three silent losses, each found by the round after the one that shipped it." / "which defeated the last-unclosed-`$(` attempt, and a backtick substitution, which it could not see at all." / "these are the spellings the rewritten delimiter matcher had to keep reading." / "An earlier version of this helper asserted only that the body's `emit error P9` was not counted, and that is the assertion that let four regressions through" / "These were written as `arms nothing` cases when the matcher used a character class and refused whatever fell outside it." / "the guard passes against the revision that has the bug." / "Every heredoc defect this scanner has had" / "Reproduced against c7f71c3: each shape below made a well-formed detector un-gateable, because the site counter and the id extractor disagreed" / "this guard was a byte-identical copy of that fixture until an audit caught it, a duplicate masquerading as a second dimension of coverage." / "The discovery half kept a greedy sed of its own ... while the verdict scanner beside it read that same line as two call sites correctly. Discovery now runs that same scanner" / "a per-row exit 2 used to leave the FIRST row's report" / "P6: trailing argv was silently ignored"
- scripts/check-plugin-catalog-enablement.sh: "Nothing enforced it. Three plugins reached main with a catalog entry and no `enabledPlugins` key: ai-slop (#2892), context-budget (#2932) and improvement (#2985), while the plugin PRs on either side of them (coupling #2913, overengineering #2961) remembered the settings entry." / "WHERE ENABLEMENT LIVES NOW" / "this file no longer mirrors the whole catalog (that mirror was writing one ..." / "The class that shipped three times." / "Raised as P2 by the Codex review on #3235."
- scripts/check-plugin-catalog-enablement.test.sh: "the class that actually shipped (a catalogued plugin with no enabledPlugins key, three times: #2892, #2932, #2985)" / "--- 2. The class that shipped: catalogued, enabled nowhere." / "what settings no longer mirrors" / "This is the post-migration shape:" / "Under the original `sed -n ...`" / "Raised as informational by the security review on #3235." / "Raised as P2 by the Codex review on #3235."
- scripts/check-purged-em-dashes.sh: "WHY (#2891)." / "each landed shard ... every shard so far has needed a rationale-withheld reviewer to catch clauses the automated passes waved through. Nothing then stops ... because no lane enforces the policy." / "29,649 em-dash prose lines across 1,074 tracked markdown files remained when this gate was written" / "#3342 needs a several-hundred-file allowlist; that is only affordable". .test.sh: "the narrower config #3342 asked for".
- scripts/check-skill-precompute-compose.test.sh: "The #3377 fix hoists `git rev-parse` into the parent, so the mode dispatch ... is now load-bearing"
- scripts/check-docs-naming.sh: "docs/ carried a mix of UPPER-KEBAB, lower-kebab, and mixed-case names for years, and every reference to a doc had to remember which spelling that one file used."
- scripts/lib/fixture-tree.sh: "not the three-variable spelling that was drifting through the suites."
- scripts/sync-resolve-convention-home.sh: "plugin-quality (the ADR 0018 pilot) is the first carrier."
- scripts/test-git-helpers.sh: "The exported GIT_DIR behind the real incident came from an ad-hoc tool invocation, not from a git hook: this repository has no git hook at any scope and core.hooksPath is unset everywhere." .test.sh: "That is what happened to #2827 -> #2830."
- scripts/generate-catalog.mjs: "One-way gates existed before this: ... but nothing compared CATEGORY_ORDER back to the document"

Intentional-removal: dissolve-comments pass; removed comment text is recorded above.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UpA569K1JgnKu1jQvRqMjL
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants