diff --git a/plugins/claude-config/.claude-plugin/plugin.json b/plugins/claude-config/.claude-plugin/plugin.json index 9828cd39b2..bf74a0c307 100644 --- a/plugins/claude-config/.claude-plugin/plugin.json +++ b/plugins/claude-config/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "claude-config", - "version": "0.51.30", + "version": "0.52.0", "description": "Nine configuration-health skills (plus setup) for a repo's Claude Code configuration: audit (settings.json / .mcp.json / hooks / plugins / permissions drift), audit-automation-gaps (evidence-gated verdicts on automation gaps), audit-permission-grants (allow-rule / allowed-tools grants for auto-mode durability and portability), audit-permission-state (the permission rules actually in effect: every settings scope merged with per-rule provenance, what auto mode drops on entry, config written where nothing reads it, and which managed intents are enforced versus loosenable), draft-auto-mode-rules (interview and draft a paste-ready autoMode classifier block; prints only, never writes), audit-instructions (locally-owned instruction surfaces vs current model capability, proposing removals/rewrites of instructions the model no longer needs, and detecting cross-surface instruction conflicts), audit-prompting-postures (the additive lane: posture guidance the prompting guide says a component's purpose needs but the component does not carry), audit-pass (one coordinated, ordered, resumable pass over a named target: three-scope inventory, run-time-derived exclusion set, stable finding identity, suppression memory, resume, one human gate, delegating every check to the plugin that owns it), and unhobble (the empirical bare-baseline experiment: reversibly strip a repo's standing instructions, log real stumbles against the current model, re-add only what evidence earns).", "author": { "name": "Melodic Software", diff --git a/plugins/claude-config/CHANGELOG.md b/plugins/claude-config/CHANGELOG.md index 3b8f33f7b2..034b0cae0d 100644 --- a/plugins/claude-config/CHANGELOG.md +++ b/plugins/claude-config/CHANGELOG.md @@ -3,6 +3,88 @@ All notable changes to the `claude-config` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +Versions 0.51.8 to 0.51.9 and 0.51.11 to 0.51.14 were reserved by parallel branches and never released. + +## [0.52.0] - 2026-09-29 + +### Added + +- **`unhobble`: a `decide` argument.** `decide` lists an experiment's open decisions and, when the + `discovery` plugin is installed, researches each one and gives the memo to two blind decision + agents. Agreement is presented for confirmation and disagreement returns to the operator as a + question. It never mutates ([#4094](https://github.com/melodic-software/claude-code-plugins/issues/4094)). +- **`unhobble`: `status` prints phase, elapsed days, ledger row count, register holds, confounds, + and the pull request URL.** The manifest gains `phase`, `branch_deviation`, and an optional + `pr_url`. `readd` refuses to run while the phase is `bare` or `observe` and reports the ledger + grouped by suspected missing instruction against the two-row gate (#4094). + +### Changed + +- **`unhobble`: a convention unit's default follows an oracle test.** A gating oracle strips the + prose and keeps the gate, an advisory oracle or none keeps the unit, and a register class is kept + whatever the test says. A unit under `plugins//` is recorded as + `unstripped-product-surface` and never stripped, and a session pinned to a branch uses it as the + experiment branch (#4094). +- **`unhobble`: silence is no longer a deletion warrant.** An empty ledger licenses deleting an + editorial candidate only. A consequential rule the ledger did not defend goes to a Deletion watch + or is restored, and only a closed watch makes its removal permanent. The watch counts qualifying + sessions as fresh sessions on the experiment branch + ([#3563](https://github.com/melodic-software/claude-code-plugins/issues/3563)). +- **`audit-instructions`: the findings relay follows `criteria.md` surfaces and tiers.** I31 and I33 + admit any file inside a skill directory (I33 excludes `SKILL.md`), I33 names the nearest ancestor + `SKILL.md` as its hub, and I32 is CRITICAL under `plugins/` and IMPORTANT on a user or project + surface ([#4116](https://github.com/melodic-software/claude-code-plugins/issues/4116), + [#4656](https://github.com/melodic-software/claude-code-plugins/issues/4656)). +- **`audit-instructions`: the subagent-window claim carries a verification record, the read-only + contract covers lane reports and run-state writes, and I33 rows count as one dispatch outside the + per-lane verifier batches.** `context/execution-and-report.md` is deleted; `SKILL.md`, the + finding-identity reference, and `context/persist-findings.md` own what it restated + ([#4113](https://github.com/melodic-software/claude-code-plugins/issues/4113), + [#4114](https://github.com/melodic-software/claude-code-plugins/issues/4114), + [#4115](https://github.com/melodic-software/claude-code-plugins/issues/4115)). +- **`README.md`** describes the `audit-instructions` flags and the `unhobble` state fields. + +### Fixed + +- **`audit-instructions`: I21 and the `audit` effort-pin row state the current effort defaults.** + The default carve-out named only Opus 4.7 as a non-`high` default; the model-config page's + resolution order also has Opus 5.5 and Sonnet 5.5 defaulting to `medium`, so the "`high` is the + default" exemption lifts for them too. The stale first-run-hold clause is removed from Source, + and Verified and the recheck trigger follow the 2026-09-28 read of both pages. +- **`unhobble`: the state-write gotcha and its eval no longer restate the `block-hook-bypass` + exit code or call a heredoc alone a blocked form.** The skill says why a shell write is wrong and + leaves what the hook catches to its README. +- **`audit-instructions`: `finding-ids.sh` no longer hashes a non-instruction file under `$HOME`.** + A row naming `~/.ssh/config` or `~/.claude/.credentials.json` is refused as + `surface-not-an-instruction-file`. The user surface stays home-wide for instruction-file shapes + (any markdown file, and `settings.json`, `settings.local.json`, `hooks.json` inside a `.claude` + tree or the resolved `CLAUDE_CONFIG_DIR`, plus any file beneath a `skills/` directory in those + trees) (#4116). +- **`audit-instructions`: the I15 ledger link in `conflict-criteria.md` no longer points outside the + repository.** It names the path in code font + ([#3568](https://github.com/melodic-software/claude-code-plugins/issues/3568)). +- **`audit`: the row for a live `PreToolUse` hook that already blocks a pattern carries the + unattended-lane note** like the other ask-rule rows + ([#4600](https://github.com/melodic-software/claude-code-plugins/issues/4600)). +- **`audit-permission-state`: a settings file whose last `permissions` key is a string is no longer + rejected as invalid JSON.** `{"permissions":{"defaultMode":"bypassPermissions"}}` is now scanned, + so `C2-defaultMode` and `C5-disableType` fire on real files. Both `C2-defaultMode` findings say to + remove the value from the named file and state the masking claim only where no higher-ranked + settings file or `--permission-mode` sets a mode. The unsourced auto version boundary is removed + ([#4027](https://github.com/melodic-software/claude-code-plugins/issues/4027)). +- **`audit-instructions`: the 200000-token lane-window fallback carries a recheck trigger.** The + claim, basis, as-of date and trigger sit beside the fallback in `SKILL.md`. +- **`audit-instructions`: `emit-findings.sh` finds the I33 hub when run from a repository + subdirectory with relative paths.** The hub probe now starts from the repository root instead of + the working directory. + +In-place changelog corrections: + +- 0.51.18: body replaced by a pointer; it repeated the 0.51.16 Fixed entry (released by #5159) and the 0.51.10 Changed entry (released by #5161). +- 0.51.17: body replaced by a pointer; it repeated the 0.51.15 entry (released by #5156). +- 0.51.7: states it shipped through #5154 (commit 9c2db71f6), not #5059 (closed unmerged). +- File header: notes that 0.51.8 to 0.51.9 and 0.51.11 to 0.51.14 were reserved by parallel branches and never released. + ## [0.51.30] - 2026-09-28 ### Changed @@ -125,29 +207,11 @@ All notable changes to the `claude-config` plugin are documented here. Format fo ## [0.51.18] - 2026-09-28 -### Fixed - -- **Permission surfaces now match Claude Code 2.1.257–2.1.263** ([#4027](https://github.com/melodic-software/claude-code-plugins/issues/4027)). Read and Edit deny rules cover Bash redirect targets from v2.1.257; the v2.1.259 widening to every Bash argument was reverted in v2.1.260 and is not written in ([permissions](https://code.claude.com/docs/en/permissions)). `allowManagedPermissionRulesOnly` ignores `--allowedTools` and user, project, and local rules; `--disallowedTools` and session deny and ask rules stay across reloads ([settings reference](https://code.claude.com/docs/en/settings-reference#allowmanagedpermissionrulesonly)). The ask-rule contract is quoted from the [auto mode config](https://code.claude.com/docs/en/auto-mode-config) page; #42797 is closed, #83766 is open, and v2.1.257 fixed compound and subshell paths only. `strictPluginOnlyCustomization` is `true` or a per-surface array; `"mcp"` blocks user and project MCP servers and does not switch hooks off ([settings reference](https://code.claude.com/docs/en/settings-reference#strictpluginonlycustomization)). Interactive `!` shell mode runs outside the sandbox even in strict mode from v2.1.260 ([sandboxing](https://code.claude.com/docs/en/sandboxing)). A managed settings file, drop-in, plist, or HKLM value that cannot be parsed refuses startup; a user settings parse failure warns in `/status` ([managed settings](https://code.claude.com/docs/en/managed-settings), [settings](https://code.claude.com/docs/en/settings)). Server-managed settings are cached at `~/.claude/remote-settings.json`; cross-source merge is by key kind, and `sandbox.credentials.awsPairs` and `sandbox.ripgrep` are taken whole since v2.1.257 ([server-managed settings](https://code.claude.com/docs/en/server-managed-settings)). `permissions.blockReadsOutsideWorkingDirectories` fences Read, Grep, Glob, and LSP; Bash subprocess reads stay unbounded ([settings reference](https://code.claude.com/docs/en/settings-reference#permissionsblockreadsoutsideworkingdirectories)). Permission-rule lints keep a `)` inside a specifier, report `Bash(ls) x` as a malformed Tool(content) rule, and treat an uncompilable deny as guarding the literal path. - -### Changed - -- **Managed-policy diagnosis routes to `/status`** ([#4027](https://github.com/melodic-software/claude-code-plugins/issues/4027)). `audit-permission-state` still cannot see server-managed settings as live policy. The note sends the operator to `/status` (Setting sources and the Organization policy line) and to `claude doctor`, which shows the same line, for a policy that did not load, a policy-helper failure, or a credential that is signed in but not in use. +No change: repeats the 0.51.16 entry (released by #5159) and the 0.51.10 entry (released by #5161). ## [0.51.17] - 2026-09-28 -### Changed - -- **Permission and effort facts from Claude Code 2.1.257 to 2.1.261** - ([#4027](https://github.com/melodic-software/claude-code-plugins/issues/4027)). - `required-permissions.md` now states that redirect targets are covered by Read and Edit deny - rules from 2.1.257, and that the 2.1.259 widening onto Bash arguments was reverted in 2.1.260. - Strict sandbox mode does not sandbox `!` shell-mode commands in an interactive session from - 2.1.260. `permissions.blockReadsOutsideWorkingDirectories` fences Read, Grep, Glob, and LSP - from 2.1.257. The audit checklist records `bashOutputMaxChars` (clamped 4000 to 128000, not - raised by default) and that `taskOutputMaxChars` was removed in 2.1.277. `audit-instructions` - I21 drops the "Opus 5 has no hold" sentence for the current model-config page: Opus 5.5 ignores - a top-level user `effortLevel`, and that key still applies on Opus 5, Fable 5.1, and earlier - models. Each claim cites the page read on 2026-09-28. +No change: repeats the 0.51.15 entry (released by #5156). ## [0.51.16] - 2026-09-28 @@ -179,6 +243,8 @@ All notable changes to the `claude-config` plugin are documented here. Format fo ## [0.51.7] - 2026-09-28 +Shipped through [#5154](https://github.com/melodic-software/claude-code-plugins/pull/5154) (commit 9c2db71f6), not #5059, which was closed unmerged. + ### Fixed - **`audit-permission-state` treats project `defaultMode: "bypassPermissions"` as dead** diff --git a/plugins/claude-config/README.md b/plugins/claude-config/README.md index 5376cb8bb3..81a0ccf81d 100644 --- a/plugins/claude-config/README.md +++ b/plugins/claude-config/README.md @@ -111,8 +111,8 @@ Audits instruction *content* against current model capability, a different quest sibling audits (config-file correctness) and from `skill-quality:check` (structural lint) or `docs-hygiene:compress` (token brevity). It sweeps the locally-owned surfaces (user + project `CLAUDE.md`, a natively read `AGENTS.md`, `.claude/rules`, skill bodies, agent definitions, hook -instruction text, output styles) against a sixteen-check catalog cited to current official prompting and harness doctrine, running -a fresh read-only subagent per surface, then a fresh-context verify pass that re-judges every removal +instruction text, output styles) against a check catalog cited to current official prompting and harness doctrine, running +plugin-atomic, token-budgeted read-only lanes (see Lane sizing in `skills/audit-instructions/SKILL.md`), then a fresh-context verify pass that re-judges every removal proposal before it is surfaced. Findings are tiered mechanical vs behavioral and delivered as a report plus proposed diffs, report-only and never auto-applied. On memory-layer surfaces it runs only the model-era checks and routes hygiene findings to the `claude-memory` plugin's `audit` skill (with @@ -129,6 +129,9 @@ pass alone, so a scheduled hygiene routine can compose it on its own token budge /claude-config:audit-instructions skills # one surface (claude-md|rules|skills|agents|hooks|output-styles) /claude-config:audit-instructions conflicts # the cross-surface conflict pass only /claude-config:audit-instructions --opinion # also run the default-off OPINION-tier checks +/claude-config:audit-instructions --unattended # nobody to answer: disclose the ~20-dispatch cost instead of asking +/claude-config:audit-instructions --resume # continue the latest run, re-running only incomplete or changed lanes +/claude-config:audit-instructions --persist-findings # also write I28-I33 findings as a review-findings file for review:fanout fix ``` ### audit-pass @@ -184,7 +187,9 @@ phases: The canonical trigger is a frontier model release. Instructions written for the previous generation are the experiment's -subject. Human-gated at every mutation; state persists under `${CLAUDE_PLUGIN_DATA}` for resume. +subject. Human-gated at every mutation; the manifest and stumbles log live in the repo under +`.claude/unhobble//`, and `${CLAUDE_PLUGIN_DATA}` holds only `backups/` (see State in +`skills/unhobble/SKILL.md`). ```shell /claude-config:unhobble # guided full flow @@ -246,7 +251,9 @@ automatically. If you used `/claude-config-audit:memory-health`, install it expl No `userConfig`. One tracked consumer-project file, `audit-pass`'s suppression record, above. Persistent plugin state: `audit-pass` writes its run reports and manifests under `${CLAUDE_PLUGIN_DATA}`, which resolves under `~`, outside a target below the home directory and -inside one at or above it. A run never *scans* what it wrote: where the resolved report path is +inside one at or above it. `audit-instructions` keeps its run files under +`${CLAUDE_PLUGIN_DATA}/audit-instructions/runs`, and `unhobble` keeps pre-strip `backups/` under +`${CLAUDE_PLUGIN_DATA}/unhobble//`. A run never *scans* what it wrote: where the resolved report path is contained in the target, the run excludes that path before writing and says so. Network: `audit` fetches official docs pages and each registered marketplace's `marketplace.json` from `raw.githubusercontent.com` (read-only; a failed fetch degrades to SKIP). diff --git a/plugins/claude-config/skills/audit-instructions/SKILL.md b/plugins/claude-config/skills/audit-instructions/SKILL.md index 5c209da102..d2bfee2b16 100644 --- a/plugins/claude-config/skills/audit-instructions/SKILL.md +++ b/plugins/claude-config/skills/audit-instructions/SKILL.md @@ -33,9 +33,10 @@ every change is applied by the human (or explicitly delegated afterward), never Diffs are proposed artifacts. A clean audit is a valid outcome. `disallowed-tools: Edit, NotebookEdit` narrows the surface; it does **not** make the contract -mechanical. `Write` stays for the Phase D persist and `Bash` for the pre-scans, and either can mutate -a file this skill has already read, so this is an instruction-held contract with a narrowed accident -surface, not an enforced one. Never describe it to an operator as a guarantee. The restriction clears +mechanical. `Write` stays for the Phase D persist and for lane reports under +`runs///lanes/`, and `Bash` for the pre-scans, `lane-runs.sh`, and the +`run-state.sh` lease writes under `runs/`. Either can mutate a file this skill has already read, so +this is an instruction-held contract with a narrowed accident surface, not an enforced one. Never describe it to an operator as a guarantee. The restriction clears on their next message (, frontmatter reference, fetched 2026-08-12), so whoever accepts a diff can apply it. `audit-prompting-postures` carries the identical declaration and the identical caveat, because the two state the same contract. @@ -61,17 +62,16 @@ concerns its siblings already cover, so route rather than re-answer: On **memory-layer surfaces** (CLAUDE.md, a natively read AGENTS.md, CLAUDE.local.md, `.claude/rules/`, and `rules/` under the user root Phase A resolves), this skill runs only the model-era checks I6–I35. It never runs or reports the hygiene checks -I1–I5 (line-necessity, length, placement, inferable content, rule-to-hook) on these surfaces; -that instruction-memory hygiene layer belongs to the `claude-memory` plugin. When that plugin is -installed, route memory-layer hygiene to its `audit` skill; when it is not installed, emit a single -one-line pointer to the official CLAUDE.md include/exclude guidance (recorded with I1–I5 in -[reference/criteria.md](reference/criteria.md)) so the operator knows where that audit lives, though this -skill still does not perform it. Either way, no I1–I5 hygiene finding is ever produced here. On -**non-memory surfaces** (skill bodies, agent definitions, hook instruction text, output styles) the -catalog applies, since no incumbent auditor covers instruction content there, **bounded by each row's -own surface declaration**, which is narrower than the partition for some checks. I13, I14, I29, -I31, I32, I33, and I34 name their own surface sets and are not run outside them; I15 is answered -pairwise by Phase B2; this partition never widens a row. +I1–I5 (line-necessity, length, placement, inferable content, rule-to-hook) on these surfaces; that +layer belongs to the `claude-memory` plugin. When it is installed, route memory-layer hygiene to its +`audit` skill; when it is not, emit a single one-line pointer to the official CLAUDE.md +include/exclude guidance (recorded with I1–I5 in [reference/criteria.md](reference/criteria.md)). +Either way, no I1–I5 hygiene finding is ever produced here. On **non-memory surfaces** (skill +bodies, agent definitions, hook instruction text, output styles) the catalog applies, since no +incumbent auditor covers instruction content there, **bounded by each row's own surface +declaration**, which is narrower than the partition for some checks. I13, I14, I29, I31, I32, I33, +and I34 name their own surface sets and are not run outside them; I15 is answered pairwise by +Phase B2; this partition never widens a row. I15 (cross-surface conflict) carries its own narrower routing on the same convention, drawn from the population `claude-memory:audit`'s C6 actually enumerates via `discover-instruction-surfaces` @@ -98,24 +98,24 @@ two are routinely conflated: - **This skill (marketplace plugin).** A standing, report-only audit of locally-owned Claude Code instruction surfaces against the versioned I-catalog in [reference/criteria.md](reference/criteria.md): target-model scoping, deterministic pre-scans, the cross-surface conflict pass, and harness-claim - staleness the vendor sweep does not look for. Prompts embedded in application source are outside - this skill's scope by design; that surface stays with the bundled subcommand. + staleness the vendor sweep does not look for. Prompts embedded in application source stay with the + bundled subcommand. **Routing.** The two compose rather than compete. When the bundled `claude-api` skill resolves in this session, prefer its `prompt-audit` for a model migration or any pass over application-code -prompts, and run it as the vendor procedure whenever the target model changes. Prefer this skill for -the standing catalog audit of Claude Code surfaces, for cross-surface conflicts, and for harness -claims that misstate Claude Code's own behavior. Where a sweep wants both, run both: recurring gap -shapes the vendor sweep surfaces feed this catalog as new rows, and this skill's findings never -substitute for the vendor procedure on a model change. +prompts, and run it whenever the target model changes. Prefer this skill for the standing catalog +audit of Claude Code surfaces, for cross-surface conflicts, and for harness claims that misstate +Claude Code's own behavior. Where a sweep wants both, run both: recurring gap shapes the vendor +sweep surfaces feed this catalog as new rows, and this skill's findings never substitute for the +vendor procedure on a model change. -**Mutation gate.** `prompt-audit` edits files when the request asks for edits. This skill's contract -is report-only, so never chain into a `prompt-audit` apply on this skill's behalf; surface the -finding and let the user invoke the sweep themselves. +**Mutation gate.** `prompt-audit` edits files when the request asks for edits. This skill is +report-only, so never chain into a `prompt-audit` apply on its behalf; surface the finding and let +the user invoke the sweep. **Availability is never assumed.** Bundled surfaces are gated by settings, environment, plan, and -host; this section states what to do when the surface resolves, never that it is present. The -subcommand set, the distribution facts behind it, and their recheck triggers are recorded in +host; this section states what to do when the surface resolves, never that it is present. Its +subcommand set, distribution facts, and recheck triggers are in [reference/bundled-claude-api.md](reference/bundled-claude-api.md). ## Arguments @@ -168,26 +168,23 @@ Two flags govern the `OPINION` tier, whose enablement policy the catalog defines - `--opinion`: also run the `OPINION`-tier checks that emit findings. Off by default; their findings are capped at `info` and are never applied. **Which rows those are is read from the - catalog at run time and deliberately not restated here**: the catalog owns the enablement policy, - so a second copy of the set in this file is one more thing to keep in sync on every new - `OPINION` row, and a stale copy silently narrows the flag. The run's tier-transparency line - reports how many it found. + catalog at run time and deliberately not restated here**: a stale copy would silently narrow the + flag. The run's tier-transparency line reports how many it found. - `--no-stopping-condition`: disable the `OPINION`-tier stopping condition that bounds I6 and I8. It is on by default because it withholds findings rather than emitting them, so turning it off makes both trimming checks more aggressive, not the audit more conservative. `--persist-findings` also writes the run's I28 and I29 scan findings and I30 to I33 lane findings as -a `type: review-findings` file for `review:fanout`'s `fix` action (off by default; only those families, -body-scoped; a proposal for a human-gated relay, not an applied edit; see [context/persist-findings.md](context/persist-findings.md)). +a `type: review-findings` file for `review:fanout`'s `fix` action (off by default; only those +families, body-scoped; a proposal for a human-gated relay, not an applied edit; see +[context/persist-findings.md](context/persist-findings.md)). `--unattended` declares that nobody is available to answer: the ~20-dispatch confirmation in -Phase B becomes a disclosure on the Phase D cost line instead of a question. Only the caller -declares it, in the invocation; a run never infers it from its own session, and a run without the -flag asks. +Phase B becomes a disclosure on the Phase D cost line. Only the caller declares it, in the +invocation; a run never infers it, and a run without the flag asks. -`--resume` continues the latest run under this project's state key instead of starting a new one, -re-running only the lanes whose report is incomplete or whose inputs changed. Phase B's "Run files -and resume" owns the mechanics. +`--resume` continues the latest run under this project's state key, re-running only the lanes whose +report is incomplete or whose inputs changed; Phase B's "Run files and resume" owns the mechanics. ## Phase A: Inventory @@ -201,33 +198,32 @@ off, and the exclusions. Phase B cannot run against a record set built any other Run one **fresh read-only subagent per lane**, where a lane is a set of surfaces packed under the token budget in "Lane sizing" below, each lane sharing [reference/criteria.md](reference/criteria.md) and applying the per-surface check partition from -the Scope boundary to every surface it holds. **A record whose residency Phase A could not establish carries that state into -its lane**: the lane still runs, and reports each result as `RESIDENCY-UNRESOLVED` with the named -unresolved condition (Phase D) rather than as a finding, since a removal or a rewrite proposed -against a surface the session may never load is work the reader cannot act on. Seed each lane's -candidate set from a **central pre-scan** run once over every inventoried file before dispatching, -with an extended Bash timeout or in the background, because one pass over a large inventory can -take minutes under Git Bash on Windows. Hand each lane the rows whose `file:` prefix is one of its -files; a lane never re-scans. The seeded checks span both evidence tiers; the scan itself is only -ever deterministic pattern-marking: +the Scope boundary to every surface it holds. **A record whose residency Phase A could not establish +carries that state into its lane**: the lane still runs, and reports each result as +`RESIDENCY-UNRESOLVED` with the named unresolved condition (Phase D) rather than as a finding, since +a removal or a rewrite proposed against a surface the session may never load is work the reader +cannot act on. Seed each lane's candidate set from a **central pre-scan** run once over every +inventoried file before dispatching, with an extended Bash timeout or in the background, because one +pass over a large inventory can take minutes under Git Bash on Windows. Hand each lane the rows whose +`file:` prefix is one of its files; a lane never re-scans. The seeded checks span both evidence +tiers; the scan itself is only ever deterministic pattern-marking: ```shell bash "${CLAUDE_PLUGIN_ROOT}/skills/audit-instructions/scripts/instruction-scan.sh" ... ``` It emits `file:line:check-id` candidate rows for I6 (a prohibition sentence with no paired positive -or rationale marker, per the catalog's Detect), I10 (reasoning-echo directives), the I8 families under per-family ids: `I8-a` -instructed self-check, `I8-b` conservative-reporting, `I8-c` don't-think / don't-reason, `I8-f` -think-carefully steer (I8-c's -tag-naming sub-detect is lane-only, not seeded, as are I8's base row and `I8-d` short-turn -assumptions, whose phrasings are too varied for a pattern that would earn its false-positive rate; -`I8-e` forced interim-status cadence is likewise unseeded, but on a narrower ground: its skeleton is -patternable, and it waits only on an attested instance to calibrate the interval forms against), I23 -(self-estimated context-budget phrasing, the budget clause alone, never the stop/summarize/hand-off -verb it licenses, which routinely sits in a different sentence), I25 (retired sampling parameters), -I27 (effort-for-brevity: an effort-lowering directive paired with a brevity token on one line), and -the I28 families (`I28-a` forced-compliance emphasis, case-sensitive; `I28-b` blanket tool -defaults). Concatenate `${CLAUDE_PLUGIN_ROOT}/skills/audit-instructions/scripts/restatement-scan.py` +or rationale marker, per the catalog's Detect), I10 (reasoning-echo directives), the I8 families +under per-family ids (`I8-a` instructed self-check, `I8-b` conservative-reporting, `I8-c` +don't-think / don't-reason, `I8-f` think-carefully steer; I8-c's tag-naming sub-detect is lane-only, +as are I8's base row and `I8-d` short-turn assumptions, whose phrasings are too varied for a pattern +that would earn its false-positive rate; `I8-e` forced interim-status cadence is likewise unseeded, +on a narrower ground: its skeleton is patternable, and it waits only on an attested instance to +calibrate the interval forms against), I23 (self-estimated context-budget phrasing, the budget clause +alone, never the stop/summarize/hand-off verb it licenses, which routinely sits in a different +sentence), I25 (retired sampling parameters), I27 (effort-for-brevity: an effort-lowering directive +paired with a brevity token on one line), and the I28 families (`I28-a` forced-compliance emphasis, +case-sensitive; `I28-b` blanket tool defaults). Concatenate `${CLAUDE_PLUGIN_ROOT}/skills/audit-instructions/scripts/restatement-scan.py` over the same files, in the same central pass, for the I29 families (`I29-a` description-restatement; `I29-b` sibling-section-restatement); `--count` prints the row count. Advisory: a grep cannot judge whether a rationale is genuinely present, whether a restraint clause @@ -237,23 +233,29 @@ run's resolved target model. I33 is lane-only; each lane brief restates its Must ### Lane sizing -Lane sizing is #4114's token-budget rule below. This skill ships no second sizing rule: no lane -cap and no line-count constant. **Plan the dispatch before dispatching** by running -`lane-runs.sh partition` over the inventoried files, then dispatch those lanes. -**Claim:** dispatch sizing is the #4114 token-budget partition only. **Basis:** issue #4656 -(the 9-lane cap and 2,500-line constant were a rejected second rule). **As of:** 2026-09-28. -**Recheck:** when `lane-runs.sh partition` grows a second cap, or the catalog ships a line-count -sizing constant. - -A lane's budget is a fraction of the **lane model's own** context window, since a subagent's window -is sized by the model it runs on, not the parent's. The budget is **0.25 of that window**, leaving -the rest for the catalog, the lane brief, the lane's reasoning, and its report, at **3.5 bytes per -token** (Anthropic's glossary: "a token approximately represents 3.5 English characters", fetched +One partition rule sizes lanes, a token budget: no lane cap and no line-count constant. **Plan the +dispatch before dispatching** by running `lane-runs.sh partition` over the inventoried files, then +dispatch those lanes. + +A lane's budget is **0.25 of the lane model's own context window**, leaving the rest for the +catalog, the lane brief, the lane's reasoning, and its report, at **3.5 bytes per token** +(Anthropic's glossary: "a token approximately represents 3.5 English characters", fetched 2026-09-28 from ; recheck when that entry changes or a lane overflows its window on a supported model). Non-ASCII text runs more bytes per character, so the estimate errs toward smaller lanes. Never state the budget as a line count: the line figure is derived per run from the bytes per line measured over the in-scope files. +**Claim:** a subagent's context window is sized by its own model, not the parent's. **Basis:** + (model field section). **As of:** 2026-09-29. +**Recheck:** when that page's model or context-window wording changes. + +`` is the `--window-tokens` value: the context window in tokens of the model the +lane runs on, per (fetched 2026-09-29). When the +lane's model resolves to no documented window, pass 200000, the smaller standard window: a smaller +budget only adds lanes. **Claim:** 200000 is the smallest documented window. **Basis:** the +model-config page above. **As of:** 2026-09-29. **Recheck:** when that page documents a smaller +window for a supported model. + Partition deterministically, feeding every in-scope file as `\t\t`, where the group is its plugin (or the memory layer) and the unit is its skill (or the file itself): @@ -264,9 +266,9 @@ bash "${CLAUDE_PLUGIN_ROOT}/skills/audit-instructions/scripts/lane-runs.sh" part Plugins stay atomic: whole plugins pack into a lane up to the budget, and only a plugin whose surface text exceeds the budget splits, by skill, into lanes of its own. Each `split=:` -line the script prints goes on the Phase D cost line, and an `over_budget=` line (a single skill -larger than the budget, still one lane) is named there too. Lane ids are a function of the -partition, and the partition's digest is part of every lane's input digest. +line the script prints goes on the Phase D cost line, as does an `over_budget=` line (a single skill +larger than the budget, still one lane). Lane ids and the partition digest are functions of the +partition; the digest is part of every lane's input digest. Bound concurrency to 3 to 5 lanes at a time. Before the total dispatch count (lanes plus Phase C verifiers) would exceed ~20, confirm with the user; with `--unattended`, proceed instead and let the @@ -302,16 +304,12 @@ rows to `lane-runs.sh plan --run-dir `: dispatch only the `rerun` lanes touches, and a changed partition re-runs them all. With no prior run, `--resume` says so and starts a new one. -The standing execution model and the report identity contract are recorded together in [context/execution-and-report.md](context/execution-and-report.md). Lane sizing, resume, and the -lease stay in this section and `scripts/lane-runs.sh`; that file states the two contracts so a later change to either lands in one place. - A lane that persists its report to disk writes it with the Write tool, which the `guardrails` plugin's `block-hook-bypass` guard exempts by design, never through a shell redirect whose target is carried in a variable or through inline Python, which that guard blocks because it cannot resolve -the target. A shell redirect to a literal absolute path under the host temp tree is exempt only when -`CLAUDE_PROJECT_DIR` names a project root that is not itself under a temp tree; a temp-rooted -checkout (a CI clone, a test fixture) has no such exemption, so there the Write tool is the only -route. Verified 2026-09-12 against `plugins/guardrails/hooks/block-hook-bypass.sh` +the target. A redirect to a literal absolute path under the host temp tree is exempt only when +`CLAUDE_PROJECT_DIR` names a project root not itself under a temp tree, so in a temp-rooted checkout +(a CI clone, a test fixture) the Write tool is the only route. Verified 2026-09-12 against `plugins/guardrails/hooks/block-hook-bypass.sh` (`_bbh_temp_default_applies` and the scope note in `block_bypass`) and `plugins/guardrails/README.md` ("`block-hook-bypass` ships two scratch roots exempt"); recheck when the guardrails plugin changes that guard's exemption set or its block message. @@ -320,10 +318,9 @@ that guard's exemption set or its block message. Phase B judges each surface alone, so a contradiction spanning two surfaces is invisible to it. This pass supplies the missing unit: a **pair** of surfaces that both claim authority over one behavior and -disagree. Every criterion, table and worked example lives in -[reference/conflict-criteria.md](reference/conflict-criteria.md). **A scope filters findings, never -reads.** B2 enumerates every surface `all` would collect and reports a pair when at least one anchor -is in scope; the criteria file states why. +disagree, judged per [reference/conflict-criteria.md](reference/conflict-criteria.md). **A scope +filters findings, never reads.** B2 enumerates every surface `all` would collect and reports a pair +when at least one anchor is in scope; the criteria file states why. Seed it with the deterministic pre-scan over the inventoried files: @@ -335,9 +332,9 @@ It emits `fileA:lineA|fileB:lineB|entity|flags` candidate pairs; `--count` print Advisory and always exit 0, so every row is refined against the criteria file's must-not-flag set. **The scan is a priority ordering, not the work list.** It only reaches directives naming a -tool-shaped entity, so an ordinary pair such as "Always run tests before committing" against "Never run -tests" emits nothing. Work the rows first, then read the surfaces for pairs it cannot shape-match. -**A pass that reports only what the scanner emitted has not run this check.** +tool-shaped entity, so "Always run tests before committing" against "Never run tests" emits nothing. +Work the rows first, then read the surfaces for pairs it cannot shape-match. **A pass that reports +only what the scanner emitted has not run this check.** **Detect the disagreement; do not adjudicate it:** name a winner only where the criteria file's precedence table cites a documented order, otherwise report `unresolved`. Its routing table governs @@ -350,14 +347,15 @@ non-fork** subagents, since this is a self-grade of the audit's own proposals an producing context would not be independent, prompted to refute: "would removing this instruction cause Claude to make mistakes? Argue that it is still load-bearing." Where the removal call is high-stakes and correlated blind spots are the risk, prefer a cross-vendor advisor **when one is -installed and set up**, e.g. the OpenAI Codex plugin, when its documented surface can take this -artifact, invoked per its own docs, with the fresh-context same-vendor subagent as the stated -fallback, never a route to a command that may not resolve -(per `docs/plugin-philosophy.md` "Fresh-eyes checkpoints" in the marketplace repository). +installed and set up**, e.g. the OpenAI Codex plugin, invoked per its own docs, with the +fresh-context same-vendor subagent as the stated fallback, never a route to a command that may not +resolve (per `docs/plugin-philosophy.md` "Fresh-eyes checkpoints" in the marketplace repository). Batch one verifier per lane that produced proposals (not one per finding or per surface), counted under the same ~20-dispatch gate as the Phase B plan; the B2 conflict pass keeps its own separate -verifier, and one class-batched verifier judges every I33 finding across all lanes. A proposal the -verifier defends is demoted to `info` or dropped, never surfaced as a confident removal. +verifier, and one class-batched verifier judges every I33 finding across all lanes. I33 rows are excluded from +the per-lane verifier batches, the class verifier counts as one dispatch in the plan and in the +~20-dispatch gate cost line, and a lane whose only proposals are I33 gets no per-lane verifier. A +proposal the verifier defends is demoted to `info` or dropped, never surfaced as a confident removal. **An out-of-catalog defect takes its own refutation** (the catalog's "Out-of-catalog defects" section admits it): reproduce the cited evidence, then ask whether the claim is false today. One @@ -375,11 +373,9 @@ Phases B and C **require** fresh-context, non-fork subagent dispatch. When the A blocked, unavailable, or the session cannot spawn subagents: 1. **Disclose in the report header** which phases ran inline, which were skipped, and why dispatch - was unavailable. A run that skipped verification is structurally distinguishable from a - fully verified one. -2. **Mark unverified proposals.** Every removal or rewrite that did not receive an independent - verifier carries an `(unverified)` marker in the findings table and is never surfaced as a - confident removal. + was unavailable, so a run that skipped verification is distinguishable from a verified one. +2. **Mark unverified proposals.** Every removal or rewrite without an independent verifier carries an + `(unverified)` marker in the findings table and is never surfaced as a confident removal. 3. **Extend the cost line.** The Phase D cost line lists phases that did not run and names the verification mode per surface (`verified` | `inline` | `skipped`). @@ -394,17 +390,20 @@ the two absent-prior cases. Then summarize in chat. The report header carries a **cost line**: how many checks ran per surface (naming any added by a catalog version bump), the model-scoped rows skipped for the resolved target, the estimated per-surface token delta versus the previous catalog version **for this project**, the -I6 seed as `I6 raw= surviving=` from `instruction-scan.sh --i6-counts`, and the dispatch count, planned and actual (lanes, Phase C verifiers, the B2 pass, and its verifier), -stating whether the ~20-dispatch confirmation was asked or, because the run carried `--unattended`, disclosed here in -its place. It names the lane budget (tokens and the derived line figure), every plugin the -partition split by skill and into how many lanes, any over-budget skill, and on a `--resume` how -many lanes were reused and how many re-ran. It also confirms the run added zero new interactive gates (report-only contract -unchanged; the target-model fail-loud stop is an invocation-time validation abort, not an -interactive gate, since it prompts nobody and blocks nothing mid-run). Present findings as a table. -Each row's identity is `(check, claim, sites)` per -[context/execution-and-report.md](context/execution-and-report.md); presentation fields stay -outside the hash. An I15 conflict is one finding with two sites and one Finding ID. **Finding ID** is -the row's re-run-stable `finding_id/v1` from `scripts/finding-ids.sh` ([derivation and claim templates](reference/finding-identity.md)); a refused row reads `unidentified: ` there. +I6 seed as `I6 raw= surviving=` from `instruction-scan.sh --i6-counts`, and the dispatch +count, planned and actual (lanes, Phase C verifiers, the B2 pass, and its verifier), stating whether +the ~20-dispatch confirmation was asked or, because the run carried `--unattended`, disclosed here in +its place. It names the lane budget (tokens and the derived line figure), every plugin the partition +split by skill and into how many lanes, any over-budget skill, and on a `--resume` how many lanes +were reused and how many re-ran. It also confirms the run added zero new interactive gates +(report-only contract unchanged; the target-model fail-loud stop is an invocation-time validation +abort, not an interactive gate, since it prompts nobody and blocks nothing mid-run). Present findings +as a table. Each row's identity is `(check, claim, sites)` per +[reference/finding-identity.md](reference/finding-identity.md); presentation fields stay outside the +hash. An I15 conflict is one finding with two sites and one Finding ID. **Finding ID** is the row's +re-run-stable `finding_id/v1` from `scripts/finding-ids.sh` +([derivation and claim templates](reference/finding-identity.md)); a refused row reads +`unidentified: ` there. | # | Finding ID | Check | Surface:Line | Severity | Tier | Authority | Finding | Proposed change | |---|------------|-------|--------------|----------|------|-----------|---------|-----------------| @@ -412,7 +411,10 @@ the row's re-run-stable `finding_id/v1` from `scripts/finding-ids.sh` ([derivati Phase B2's findings carry two anchors, so they get their own **Cross-surface conflicts** subsection. Beside it, an **Out-of-catalog** subsection holds the defects the catalog's "Out-of-catalog defects" section admits, each with Check `out-of-catalog`, its evidence, and where it routes; those rows never -reach `emit-findings.sh`. I33 rows move to an **I33 by plugin** section ([layout](context/execution-and-report.md)). +reach `emit-findings.sh`. I33 rows move to an **I33 by plugin** section: one collapsed +`
: spokes` block per plugin, holding the same table columns so +every row keeps its `Surface:Line` and fenced diff. The roll-up is presentation only; each spoke stays +its own finding with its own excerpt anchor. For each finding, give the proposed removal or rewrite as a fenced diff block. Tier is `mechanical` (pattern-detectable) or `behavioral` (its ground truth is observed behavior); authority is the @@ -420,21 +422,18 @@ check's tag from the catalog. An I15 conflict finding names **both** participati a relation between two instructions, not a property of one line. **No-change findings are exempt from the diff contract.** Where a check forbids proposing an edit, -covering the I15 managed-policy case and any finding routed to an owning repository rather than applied, -write `no change proposed` in the Proposed change column and, in place of the fenced diff, a one-line -statement of who owns the resolution. Never manufacture a diff to satisfy the table; a check that -forbids an edit and a report that demands one would otherwise contradict each other. +covering the I15 managed-policy case and any finding routed to an owning repository rather than +applied, write `no change proposed` in the Proposed change column and, in place of the fenced diff, +a one-line statement of who owns the resolution. Never manufacture a diff to satisfy the table. **A lane whose surface residency is unresolved emits `RESIDENCY-UNRESOLVED`, not a finding.** The Proposed change column takes a closed set: a proposed removal or rewrite (with its fenced diff), `no change proposed` (above), or `RESIDENCY-UNRESOLVED`. The third is the verdict for every result of a lane whose Phase A record carries an unresolved residency condition: the row stays in the findings table so the surface is covered, its Proposed change cell reads -`RESIDENCY-UNRESOLVED: `, and it carries no fenced -diff, since it is not work the reader can act on until that condition resolves. The candidate -removal or rewrite may be described in the Finding cell as what the lane would propose once the -surface is known to load. Such a row is not a proposal, so Phase C does not re-judge it and the -`(unverified)` marker does not apply. +`RESIDENCY-UNRESOLVED: ` with no fenced diff, and the +Finding cell may describe what the lane would propose once the surface is known to load. Such a row +is not a proposal, so Phase C does not re-judge it and the `(unverified)` marker does not apply. Three sections the catalog's `OPINION` policy requires: the shadowed-definition `info` section (the live definition and the inert one, for shadowed skills and subagents, since MCP servers are outside this @@ -449,9 +448,8 @@ may be applied from this report. A consequential deletion, a rule that governs a is outside the exception register, is applicable only when the commit cites a closed `/claude-config:unhobble watch` (qualifying sessions met, zero attributed rows). -Open the Sources line with the two official pages the paths and doctrine derive from -(code.claude.com memory + `.claude`-directory docs; the prompting pages cited per check in the -catalog). +Open the Sources line with the two official pages the paths and doctrine derive from (code.claude.com +memory + `.claude`-directory docs; the prompting pages the catalog cites per check). **With `--persist-findings`**, also emit the run's I28, I29, and I30 to I33 findings for the apply relay per [context/persist-findings.md](context/persist-findings.md), which owns every mechanic and the @@ -473,8 +471,8 @@ plainly that nothing has been applied. instead of what not to do"; adding a rationale is the fallback where a genuine hard "never" survives. Do not mechanically delete every prohibition the pre-scan flags. - **Behavioral findings ship as proposals, not confident cuts.** A narrow eval can miss a small - regression from an over-aggressive trim, which is why the verify pass and the delete-and-watch - loop exist. Never present a behavioral removal as certain. + regression from an over-aggressive trim, which is why the verify pass and the Deletion tiers + in [criteria.md](reference/criteria.md) exist. Never present a behavioral removal as certain. - **Windows shell.** The pre-scans are bash; on native Windows run them through Git Bash. - **A conflict pair needs two files.** Feeding `conflict-scan.sh` one surface at a time reproduces Phase B's blind spot and always reports clean. @@ -483,16 +481,11 @@ plainly that nothing has been applied. - Never edits an instruction file and never auto-files a tracker item; output is a report plus proposed diffs the human applies. -- Not a token-brevity pass (`docs-hygiene:compress`) and not structural skill lint - (`skill-quality:check`). -- Not memory-layer hygiene: checks I1–I5 on CLAUDE.md, a natively read AGENTS.md, and rules route to `claude-memory`'s `audit` - skill when installed, and upstream-owned plugin-cache or managed materializations route to the - owning repository rather than being edited here. -- Does not grade a contradiction whose two halves both sit in the - **discover-instruction-surfaces** population, namely root-level project **or user** `CLAUDE.md` / - `CLAUDE.local.md` / a natively read `AGENTS.md` or `.claude/AGENTS.md` / rules, including - **user↔project** pairs. That is `claude-memory:audit`'s C6. - A **nested** `CLAUDE.md` / `CLAUDE.local.md` side, an auto-memory side, or any surface outside that - population keeps the pair here; +- Not token brevity, structural skill lint, or memory-layer hygiene: the Scope boundary routes each. +- Does not grade a contradiction whose two halves both sit in the **discover-instruction-surfaces** + population, namely root-level project **or user** `CLAUDE.md` / `CLAUDE.local.md` / a natively read + `AGENTS.md` or `.claude/AGENTS.md` / rules, including **user↔project** pairs. That is + `claude-memory:audit`'s C6. A **nested** `CLAUDE.md` / `CLAUDE.local.md` side, an auto-memory side, + or any surface outside that population keeps the pair here; [reference/conflict-criteria.md](reference/conflict-criteria.md) owns the routing table and its evidence. diff --git a/plugins/claude-config/skills/audit-instructions/context/execution-and-report.md b/plugins/claude-config/skills/audit-instructions/context/execution-and-report.md deleted file mode 100644 index da7cf95784..0000000000 --- a/plugins/claude-config/skills/audit-instructions/context/execution-and-report.md +++ /dev/null @@ -1,87 +0,0 @@ -# Execution model and report contract - -The parent standing record for `/claude-config:audit-instructions` run shape and finding -identity. Phase B's lane mechanics stay in [`../SKILL.md`](../SKILL.md) ("Lane sizing", "Run -files and resume") and [`../scripts/lane-runs.sh`](../scripts/lane-runs.sh). Report keying stays -in [`report-keying.md`](report-keying.md). Persist mechanics stay in -[`persist-findings.md`](persist-findings.md). This file states only the two contracts those -surfaces implement, so a later change to either lands here first. - -## Execution model - -Lanes are sized by a token budget stated as a **fraction of the lane model's own context -window**, never as a shipped line count. The fraction is 0.25; the byte estimate is 3.5 bytes -per token. The line figure is derived per run from measured bytes per line over the in-scope -files. The glossary citation and recheck trigger live on the "Lane sizing" heading in the skill -body and are not restated here. - -Plugins stay atomic. A plugin whose in-scope surface text exceeds the budget splits by skill -into lanes of its own, and the Phase D cost line names each split. A single skill larger than -the budget still gets one lane and is named as over-budget. Concurrency is 3 to 5 lanes. - -`--unattended` is declared by the caller in the invocation and never inferred from the session. -With it, the ~20-dispatch confirmation becomes a cost-line disclosure of planned and actual -dispatch counts. Without it, the run still asks before crossing that bound. - -Per-lane reports live at -`${CLAUDE_PLUGIN_DATA}/audit-instructions/runs///lanes/.md`. Each -ends with a completion marker that carries the lane's input digest: ordered file list and -content hashes, partition digest, catalog version, conflict-criteria version, prompt digest, -harness version, resolved target model, and every behavior-affecting argument (`scope`, -`--opinion`, `--no-stopping-condition`). `last-audit.md` stays at -`audit-instructions//last-audit.md`. - -`--resume` re-runs only lanes whose completion marker is absent or whose digest changed. It -selects the latest run id under the state key. The lease is `audit-pass`'s `run-state.sh` -(`stale_after_s`, `skew_grace_s`, `owner_epoch`, released tombstone), invoked with -`--plugin-data ${CLAUDE_PLUGIN_DATA}/audit-instructions`. A live lease is refused, naming -`heartbeat_at` and `stale_after_s`. The skill is read-only, so it takes a lease and no lock. - -In a marketplace repository (`.claude-plugin/marketplace.json` present), `plugins/**` is the -editable set. The installed cache is read for residency only. Cache-commit versus HEAD drift is -named in the report. [`../reference/conflict-criteria.md`](../reference/conflict-criteria.md) -"Known limit" owns that rewrite. - -## Report contract - -Findings adopt `audit-pass`'s identity whole: -[`../../audit-pass/reference/finding-identity.md`](../../audit-pass/reference/finding-identity.md) -owns the tuple, `anchor/v1`, the heading-path discriminator, `finding_id/v1`, and `group/v1`. - -```text -identity = (check, claim, sites) -``` - -- **`check`** is `claude-config/audit-instructions/`, where `` is the catalog id as the - report prints it (`I33`, `I28-a`). Sub-rows keep their own id. -- **`claim`** is a per-check template, never free prose. Free prose in `claim` is a hard error. -- **`sites`** is sorted `(surface, anchor)` pairs. A check that fires at several sites reports - one finding per site, sharing one `group`. **I15 is the one pairwise claim**: a cross-surface - conflict is one finding with two sites, never two linked findings. Lane encounter order does - not change the id. -- **`anchor`** is an excerpt anchor (`e:`) over the flagged sentence, discriminated by the - enclosing heading path. No check in this catalog takes a whole-surface (`s:`) anchor: that - form survives the edit that remediates the row. - -`primary_site`, the `Surface:Line` cell, the rendered heading path, and the Finding and -Proposed change prose are presentation. None of them enters the hash. - -Two runs over an unchanged tree yield identical `finding_id` values. Editing an I33 opener -changes that finding's id. - -`--persist-findings` emits scanner families I28 and I29 today. The standing admission also -covers I30, I31, I32, and I33 as `Auto-applicable: No`, through a lane-fed emit path beside the -scanner-fed one, with an emitting fall-through for the judgment-selected rules (I31, I33), a -`Confidence` value the detector-findings contract defines for a model-lane finding, and a -counted decline for frontmatter-located I32 rows. The emit script and crosswalk rows are the -persist unit's work; this file records the contract they satisfy. Until those rows exist, I30 -to I33 stay in the human report and are declined `no-severity-crosswalk-row` rather than -silently dropped. - -Phase C verifies every proposal. No sampling, and no finding class that carries a diff is -exempted from verification. - -I33 findings leave the main table for an **I33 by plugin** section, one collapsed -`
: spokes` block per plugin holding the same table columns, -so every row keeps its `Surface:Line` and its fenced diff. The roll-up is presentation only: each -spoke remains its own finding with its own excerpt anchor. diff --git a/plugins/claude-config/skills/audit-instructions/context/persist-findings.md b/plugins/claude-config/skills/audit-instructions/context/persist-findings.md index 6957f727ae..64ec2f2d8e 100644 --- a/plugins/claude-config/skills/audit-instructions/context/persist-findings.md +++ b/plugins/claude-config/skills/audit-instructions/context/persist-findings.md @@ -86,8 +86,11 @@ They stay in the human report and are counted in `## Surfaces` as **A lane row is a Phase C-surviving finding, written as `::`** with the line of the flagged sentence's first physical line (for I33, the opener). A row on the wrong intake is declined naming the intake it belongs to (`reason=scanner-fed-rule` or `reason=lane-fed-rule`). -I31 and I33 are spoke rules, so a row outside a `context/` or `reference/` spoke is declined as -`reason=outside-rule-surfaces`. **I32 is the one lane rule whose catalog surfaces reach +I31 and I33 are admitted for any file inside a skill directory (`skills//`, including +`SKILL.md` for I31 but not for I33) and for any file in a `context/`, `reference/`, or `references/` +directory, where a file a memory surface points at lives; a row elsewhere is declined as +`reason=outside-rule-surfaces`. I32 is `CRITICAL` on a path under `plugins/` (the marketplace arm) +and `IMPORTANT` anywhere else (the user and project arm). **I32 is the one lane rule whose catalog surfaces reach frontmatter**: a description or `when_to_use` routing clause naming an absent skill is a real finding, but the relay is body-scoped, so the writer declines the row as `reason=frontmatter` and counts it, and the human report carries it. diff --git a/plugins/claude-config/skills/audit-instructions/evals/evals.json b/plugins/claude-config/skills/audit-instructions/evals/evals.json index 33c20677e3..110b680c5e 100644 --- a/plugins/claude-config/skills/audit-instructions/evals/evals.json +++ b/plugins/claude-config/skills/audit-instructions/evals/evals.json @@ -304,10 +304,10 @@ "id": 25, "name": "execution-model-and-report-identity-contract", "prompt": "/claude-config:audit-instructions all --unattended --persist-findings over a marketplace repository.", - "expected_output": "Reads context/execution-and-report.md as the standing record of the execution model and the report identity contract. Sizes lanes from the lane model's window, discloses the ~20-dispatch gate on the cost line because --unattended is present, writes per-lane reports under the run directory, and presents each finding as (check, claim, sites) with an I15 conflict as one row of two sites. Persists I28 and I29 and the I30 to I33 lane findings. Does not edit audited surfaces.", + "expected_output": "Follows SKILL.md's Lane sizing and Run files and resume for the execution model and reference/finding-identity.md for the report identity contract. Sizes lanes from the lane model's window, discloses the ~20-dispatch gate on the cost line because --unattended is present, writes per-lane reports under the run directory, and presents each finding as (check, claim, sites) with an I15 conflict as one row of two sites. Persists I28 and I29 and the I30 to I33 lane findings. Does not edit audited surfaces.", "files": [], "expectations": [ - "Follows context/execution-and-report.md for lane sizing, --unattended disclosure, and resume/lease shape rather than inventing a second run model", + "Follows SKILL.md Lane sizing and Run files and resume for lane sizing, --unattended disclosure, and resume/lease shape rather than inventing a second run model", "Presents each finding's identity as (check, claim, sites); an I15 conflict is one finding with two sites", "Persists I28, I29, and I30 to I33; rows the emit path refuses are declined and counted, never silently dropped", "Does not edit audited instruction files" @@ -349,7 +349,7 @@ "id": 28, "name": "i33-per-spoke-finding-rolls-up-per-plugin", "prompt": "/claude-config:audit-instructions evals/fixtures/i33-hub.md, evals/fixtures/i33-scope-note-spoke.md, evals/fixtures/i33-index-pointer-spoke.md, and evals/fixtures/i33-self-describing-spoke.md relative to the skill directory, treated as one plugin named release-notes: the hub is its SKILL.md and the three spokes sit in its reference/ directory.", - "expected_output": "Reports exactly one I33 finding, on i33-self-describing-spoke.md, whose opener describes its own loading (\"This file is loaded by the hub when...\") instead of stating its content. The finding is anchored on the opener sentence, never on the whole surface, and one class-batched Phase C verifier judges it. In Phase D it sits in an I33 by plugin section as a collapsed release-notes block, and its row keeps Surface:Line and a fenced diff that deletes the self-description while the hub's index row keeps the loading condition. The scope-note and index-pointer spokes yield no I33 finding.", + "expected_output": "Reports exactly one I33 finding, on i33-self-describing-spoke.md, whose opener describes its own loading (\"This file is loaded by the hub when...\") instead of stating its content. The finding is anchored on the opener sentence, never on the whole surface, and one class-batched Phase C verifier judges it, outside every per-lane verifier batch and counted as one dispatch. In Phase D it sits in an I33 by plugin section as a collapsed release-notes block, and its row keeps Surface:Line and a fenced diff that deletes the self-description while the hub's index row keeps the loading condition. The scope-note and index-pointer spokes yield no I33 finding.", "files": [ "evals/fixtures/i33-hub.md", "evals/fixtures/i33-scope-note-spoke.md", @@ -361,7 +361,9 @@ "Anchors that finding on the opener sentence (an excerpt anchor), not on the whole surface", "Places the I33 finding in a collapsed per-plugin section for release-notes rather than the main findings table", "Keeps Surface:Line and a fenced diff on the I33 row inside the collapsed section", - "Judges the I33 class with one class-batched verifier in Phase C" + "Judges the I33 class with one class-batched verifier in Phase C", + "Excludes the I33 row from every per-lane verifier batch and counts the I33 class verifier as one dispatch in the plan and the ~20-dispatch gate cost line", + "Gives a lane whose only proposal is the I33 finding no per-lane verifier" ] } ] diff --git a/plugins/claude-config/skills/audit-instructions/reference/conflict-criteria.md b/plugins/claude-config/skills/audit-instructions/reference/conflict-criteria.md index 4da29e057b..b41f157072 100644 --- a/plugins/claude-config/skills/audit-instructions/reference/conflict-criteria.md +++ b/plugins/claude-config/skills/audit-instructions/reference/conflict-criteria.md @@ -30,7 +30,7 @@ the pre-scan and the lane each drop. **Shared-surface ownership is out of scope for I15.** This pass detects conflicting instruction *pairs*; it does not assign *ownership* when multiple contributors' preferences meet on one surface. That governance question is recorded as rejected for this repository in -[`docs/out-of-scope/shared-surface-instruction-governance.md`](../../../../../../docs/out-of-scope/shared-surface-instruction-governance.md) +`docs/out-of-scope/shared-surface-instruction-governance.md` (#3568). `instruction-exception-register` governs deletions only, not ownership. **Claim:** I15 does not assign shared-surface ownership. **Basis:** this file's pair-detection charter; the rejected-concept ledger entry for #3568. **As of:** 2026-09-28. **Recheck:** when diff --git a/plugins/claude-config/skills/audit-instructions/reference/criteria.md b/plugins/claude-config/skills/audit-instructions/reference/criteria.md index 8eb22b5702..c2f70c2fee 100644 --- a/plugins/claude-config/skills/audit-instructions/reference/criteria.md +++ b/plugins/claude-config/skills/audit-instructions/reference/criteria.md @@ -100,8 +100,8 @@ declines a row on this ground says where it routed, so "no row" never reads as " **Axes.** Three orthogonal axes, never conflated: - **Evidence tier**: `mechanical` (pattern-detectable by static reading) or `behavioral` (ground - truth is observed model behavior, so findings ship as proposals verified by the delete-and-watch - loop, never confident removals). + truth is observed model behavior, so findings ship as proposals verified per Deletion tiers, + never confident removals). - **Authority**: `ANTHROPIC-DOCS` (official documentation), `TALK` (a recorded talk), `OPINION` (a practitioner's stated practice), or `HOUSE` (a session-knowledge defect this catalog defines itself; it has no external page to cite, and it is on by default because its ground truth is the @@ -446,8 +446,9 @@ run the other way. non-model rationale is not this instance.** Reviewability of returns, rate limits, cost, or shared mutable state each justify a bound on their own terms, and that justification is the surface's to make, not this row's to override. -- **Remediate:** propose removal or a briefer instruction; verify via the delete-and-watch loop - that default performance holds or improves. +- **Remediate:** propose removal or a briefer instruction; verify per Deletion tiers (a + consequential removal needs a closed watch, an editorial one does not) that default performance + holds or improves. - **Bounded by:** the **Stopping condition** below, which is enabled by default. - **Source:** prompting best practices, "Leverage thinking & interleaved thinking capabilities", the prefer-general-instructions statement quoted above (the gate-meeting, model-agnostic one). @@ -480,7 +481,8 @@ run the other way. is NOT a finding; the anti-pattern is the instructed self-check. **Carve-out lanes (never flagged):** security review, destructive operations, managed-upstream-file changes, PR merge gates. -- **Remediate:** propose removal; verify via the delete-and-watch loop. +- **Remediate:** propose removal; verify per Deletion tiers (a consequential removal needs a closed + watch, an editorial one does not). - **Bounded by:** the **Stopping condition** below. - **Source:** Opus 5 guide, "Task scope and over-verification", which says to remove explicit verification instructions: they "cause over-verification on Claude Opus 5, and removing them @@ -597,7 +599,8 @@ choice, on the same reasoning I10 applies to a declined widening. - **Remediate:** name the constraint the brevity or rhythm was protecting, whether a latency requirement, an external contract, or a human process, and where one exists, state that constraint instead of the turn-length assumption; where none exists, remove the directive and let - turn length follow the work. Verify via the delete-and-watch loop. + turn length follow the work. Verify per Deletion tiers (a consequential removal needs a closed + watch, an editorial one does not). - **Bounded by:** the **Stopping condition** below, which is enabled by default. - **Must NOT flag: an output-length instruction.** Brevity of the *reply* is a different subject and belongs to I8 base; this row's subject is the cadence and duration of the *turn*. @@ -639,8 +642,8 @@ report one finding per line rather than two. that a long run stays interruptible, and either state that outcome and let the model meet it, or move it to a mechanism rather than an instructed rhythm. Where the *content* of native updates is miscalibrated rather than absent, describe what a good update contains and give examples; that is - the upstream remediation and it does not reintroduce a cadence. Verify via the delete-and-watch - loop. + the upstream remediation and it does not reintroduce a cadence. Verify per Deletion tiers (a + consequential removal needs a closed watch, an editorial one does not). - **Bounded by:** the **Stopping condition** below, which is enabled by default. - **Must NOT flag: a cadence carrying its own explicit observability or interruptibility rationale.** A rhythm the surface states exists so a long autonomous run stays visible or @@ -1476,15 +1479,16 @@ not a `Model scope` annotation**, for the reason I17 states. against and stating that a model change re-opens it, or run the sweep. Upstream's own wording for the action: "If you carried effort settings over from an earlier model, run a fresh effort sweep on your evals rather than reusing them." -- **Must NOT flag: a prescription of `high` where `high` is the resolved target's default.** It is - "Equivalent to not setting the parameter", so on a model that defaults to `high` such a pin - carries no measured calibration that could go stale. **The exemption keys to the resolved target, - never to the wording.** `high` is the default on every model that supports effort **except Opus - 4.7, which defaults to `xhigh`**, so when the run's resolved target is Opus 4.7 the exemption - lifts and a `high` pin is a finding, **including a broad model-agnostic "always use `high`" that - names no model at all**. That broad pin is the sharper case rather than the excluded one: written - where `high` was the no-op default and then carried to a model whose default sits above it, it - silently becomes a step-down nobody measured, which is this row's subject exactly. A resolved target +- **Must NOT flag: a prescription of `high` where `high` is the resolved target's default.** Setting + the default "produces exactly the same behavior as omitting the `effort` parameter entirely", so + on a model that defaults to `high` such a pin carries no measured calibration that could go stale. + **The exemption keys to the resolved target, never to the wording.** In Claude Code `high` is the + default on every model that supports effort **except Opus 5.5 and Sonnet 5.5, which default to + `medium`, and Opus 4.7, which defaults to `xhigh`**, so when the run's resolved target is one of + those the exemption lifts and a `high` pin is a finding, **including a broad model-agnostic + "always use `high`" that names no model at all**. That broad pin is the sharper case rather than + the excluded one: written where `high` was the no-op default and then carried to a model whose + default differs, it silently becomes a step nobody measured, which is this row's subject exactly. A resolved target always exists, because the skill body aborts rather than run against an unresolved one, so this fence never has to guess which side of it a surface falls on. **The exemption speaks to calibration staleness only, never to level adequacy:** a model guide may recommend running above @@ -1512,14 +1516,15 @@ not a `Model scope` annotation**, for the reason I17 states. is `OPINION`-tier testimony, not a pin the surface owns. - **Source:** model configuration: "The effort scale is calibrated per model, so the same level name does not represent the same underlying value across models", stated with no model qualifier, and - the whole basis for the check. The same page supplies the first-run hold with its Opus 5 exception, - and the default carve-out: "The default effort is `high` on every model that supports effort, - except Opus 4.7, which defaults to `xhigh`." Effort supplies the remediation's wording and `high`'s - equivalence to omitting the parameter. -- **Verified 2026-08-03** against both pages, fetched as raw markdown (model configuration 83,644 - bytes; effort 21,744 bytes). **Recheck trigger:** the calibration property being restated as - cross-model-stable, the set of models carrying a first-run default hold changing, or `high` ceasing - to be the general default. + the whole basis for the check. The same page supplies the resolution order, with the default carve-out: + "`high` on every model that supports effort, except that Opus 5.5 and Sonnet 5.5 default to + `medium`, Opus 4.7 defaults to `xhigh`". Effort supplies the remediation's wording and the + equivalence of the default to omitting the parameter. +- **Verified 2026-09-28** against both pages, fetched as raw markdown (model configuration 109,848 + bytes; effort 39,458 bytes). **Recheck trigger:** the calibration property being restated as + cross-model-stable, a first-run effort hold returning to the model-config page, the resolution + order or the `effortLevel` user-settings exemption for Opus 5.5 changing, or `high` ceasing to be + the general default. ### I22: Model-routing doctrine with no baseline named @@ -1579,7 +1584,7 @@ even seeded in the pre-scan, and it is `behavioral`. The `mechanical` rows rest consequence: I10 on a refusal category the API returns, I21 on a property its page states outright. This row rests on a reported model *tendency*, "can occasionally suggest a new session", with no documented hard consequence, which is the behavioral tier's definition. The stake is the Output -format rule: behavioral findings ship as proposals paired with the delete-and-watch loop, never as +format rule: behavioral findings ship as proposals verified per Deletion tiers, never as confident removals. - **Detect:** instruction text directing the model to monitor its own remaining context and to stop, @@ -1859,8 +1864,9 @@ literalism sections ("interprets prompts literally and explicitly") corroborate audience test I8-b applies. This row is the canonical instance. - **Remediate:** for arm 1, normal conditional phrasing: "Use this tool when …". For arm 2, replace the blanket default with the condition it was standing in for: "Use [tool] when it would enhance - your understanding of the problem." Verify via the delete-and-watch loop; watch for - overtriggering receding, not just continued triggering. + your understanding of the problem." Verify per Deletion tiers (a consequential removal needs a + closed watch, an editorial one does not); watch for overtriggering receding, not just continued + triggering. - **Source:** prompting best practices, "Tool usage": prompts "designed to reduce undertriggering on tools or skills … may now overtrigger. The fix is to dial back any aggressive language. Where you might have said 'CRITICAL: You MUST use this tool when…', you can use more normal prompting @@ -1870,12 +1876,14 @@ literalism sections ("interprets prompts literally and explicitly") corroborate - **Verified 2026-08-08** against that page, fetched as raw markdown. **Recheck trigger:** those three sections changing, or any model guide stating that a current model undertriggers and needs emphasis restored, which would re-open the scoping question. -- **Routes to the findings relay.** I28 and I29 are the only checks in this catalog whose findings - reach `review:fanout`'s apply relay, behind `--persist-findings`. I28's two arms carry one +- **Routes to the findings relay.** I28 and I29 (scanner-fed) and I30 to I33 (lane-fed, admitted + through `--from-lane`) are the checks in this catalog whose findings reach `review:fanout`'s apply + relay, behind `--persist-findings`. I28's two arms carry one crosswalk rule id each, `claude-config/audit-instructions/rule-coercive-emphasis` (arm 1) and `claude-config/audit-instructions/rule-blanket-tool-default` (arm 2), both `IMPORTANT`, argued in [the severity crosswalk](https://raw.githubusercontent.com/melodic-software/claude-code-plugins/main/docs/conventions/detector-findings/README.md). - Every other check here stays report-only: no crosswalk row, no relay. The persist mechanics, + Every other check here stays report-only: no crosswalk row, no relay. I32 relays as `error` + (`CRITICAL`) on the marketplace arm and `warning` (`IMPORTANT`) on the user and project arm. The persist mechanics, including the body-scope fence, are [context/persist-findings.md](../context/persist-findings.md). - **The remediation is a downgrade, never a deletion.** The directive survives verbatim and only its volume changes. A proposal that removes the instruction rather than its shouting has misread @@ -2126,5 +2134,4 @@ Findings are presented using the Phase D report table defined in the skill body restated here. A clean audit ("No instructions flagged.") is a valid outcome. Behavioral-tier proposals are -always presented as proposals paired with the delete-and-watch follow-through, never as confident -removals. +always presented as proposals verified per Deletion tiers, never as confident removals. diff --git a/plugins/claude-config/skills/audit-instructions/reference/finding-identity.md b/plugins/claude-config/skills/audit-instructions/reference/finding-identity.md index d85ded5ab9..299b4ddc6c 100644 --- a/plugins/claude-config/skills/audit-instructions/reference/finding-identity.md +++ b/plugins/claude-config/skills/audit-instructions/reference/finding-identity.md @@ -22,7 +22,10 @@ identity = (check, claim, sites) cross-surface conflict is one finding with two sites, never two linked findings, and the order the lane met the two surfaces in does not change its id. - **`surface`** is the repo-relative POSIX path of the physical file, or `user:` for a user-scope surface. + directory>` for a user-scope surface. Under the home directory only an instruction file is a + surface: a markdown file, or, inside a `.claude` tree or the resolved + `${CLAUDE_CONFIG_DIR:-~/.claude}`, a `settings.json`, `settings.local.json` or `hooks.json`, or + any file beneath a `skills/` directory. - **`anchor`** is always an excerpt anchor (`e:`), over the flagged line's text, discriminated by the enclosing heading path. Every check in this catalog is about a sentence, so none takes the whole-surface form (`s:`): an `s:` finding survives every edit to its file, including the edit @@ -47,7 +50,8 @@ Each output line is the input row, then `finding_id/v1`, then `group/v1`, then o field per site, tab-separated. `--records` prints the same finding as a JSON record that `finding-identity.sh validate-record` accepts. A row the script cannot identify (an id with no template, a pairwise row for a check that is not I15, a surface outside the repository and the home -directory, an unreadable line) prints `#REFUSED`, the row, and the reason, and is never given an id. +directory, a file under the home directory that is not an instruction file, an unreadable line) +prints `#REFUSED`, the row, and the reason, and is never given an id. ## Claim templates diff --git a/plugins/claude-config/skills/audit-instructions/scripts/emit-findings.sh b/plugins/claude-config/skills/audit-instructions/scripts/emit-findings.sh index c0927c7fda..0ca61814a7 100755 --- a/plugins/claude-config/skills/audit-instructions/scripts/emit-findings.sh +++ b/plugins/claude-config/skills/audit-instructions/scripts/emit-findings.sh @@ -20,6 +20,10 @@ # `file:line:check-id` shape. Confidence is omitted: a judgment # selected the row, and the contract has no grade below `high` # for a producer that performs no reviewer verification. +# I31 and I33 are admitted for a file inside a skill directory +# (skills//...) or in a context/, reference/, or references/ +# directory; I33 excludes SKILL.md. I32 is CRITICAL on a path +# under plugins/ and IMPORTANT anywhere else. # A row on the wrong path is declined with the path it belongs to. Every other # check id (I6, I8-a/b/c/f, I10, I23, I25, I27, ...) has no crosswalk row, and # the detector-findings contract admits no row whose tier cannot be looked up @@ -76,7 +80,8 @@ At least one of --from and --from-lane is required. --from is instruction-scan.sh output (`file:line:check-id` rows); run it with --body-only. --from-lane is the lane findings for I30, I31, I32, and I33 that survived Phase C, one `file:line:check-id` row each, the line being the -flagged sentence's first line. --out is the CONVENTION-RESOLVED destination; if it exists, a +flagged sentence's first line. I31 and I33 rows must sit in a skill directory or +a context/, reference/, or references/ directory (I33 not in SKILL.md). --out is the CONVENTION-RESOLVED destination; if it exists, a -2/-3 suffix is appended (non-overwrite naming). --branch defaults to the current git branch. --declined-carveout records how many I28/I29 candidates the model lane dropped for a criteria carve-out before this script ran, so that @@ -300,9 +305,11 @@ LC_ALL=C awk \ } function lane_rule(id) { return (id == "I30" || id == "I31" || id == "I32" || id == "I33") } # Tier mirror of the severity crosswalk (see header comment). I32 is - # CRITICAL, I33 is SUGGESTION, and every other emitted rule is IMPORTANT. - function rule_tier(id) { - if (id == "I32") return "CRITICAL" + # CRITICAL on the marketplace arm (a path under plugins/, catalog severity + # error) and IMPORTANT on the user and project arm (catalog severity warning). + # I33 is SUGGESTION, and every other emitted rule is IMPORTANT. + function rule_tier(id, loc) { + if (id == "I32") return (loc ~ /^plugins\//) ? "CRITICAL" : "IMPORTANT" if (id == "I33") return "SUGGESTION" return "IMPORTANT" } @@ -318,14 +325,20 @@ LC_ALL=C awk \ } return "CHANGELOG.md" } - # The skill hub a spoke belongs to: the SKILL.md one directory above the - # spoke directory. The I33 remediation can land there, so its Action names it. - function hub_of(loc, h) { + # The skill hub a spoke belongs to: the SKILL.md of the nearest ancestor + # directory that holds one. The I33 remediation can land there, so its Action + # names it. A file no skill owns (one a memory surface points at) has no hub. + function hub_of(file, loc, base, h, probe, probe_line) { + base = (length(file) >= length(loc) && substr(file, length(file) - length(loc) + 1) == loc) \ + ? substr(file, 1, length(file) - length(loc)) : (repo_root_pwd != "" ? repo_root_pwd : repo_root) "/" h = loc - sub(/\/[^\/]+\/[^\/]+$/, "", h) - return h "/SKILL.md" + while (sub(/\/?[^\/]+$/, "", h) && h != "") { + probe = base h "/SKILL.md" + if ((getline probe_line < probe) >= 0) { close(probe); return h "/SKILL.md" } + } + return "" } - function rule_action(id, loc) { + function rule_action(id, loc, file, hub) { if (id == "I28-a") return "Downgrade the emphasis, never the directive: restate as normal conditional phrasing (\"Use this tool when ...\"). The directive must survive the edit verbatim, apart from capitalization forced by dropping a leading wrapper; only its volume changes." if (id == "I28-b") @@ -336,17 +349,22 @@ LC_ALL=C awk \ return "Restate the sentence as the current rule and its reason in the present tense; the rule itself survives, and only its framing against a prior version changes. Remediation target for any history worth keeping: " changelog_of(loc) " or an ADR, never this spoke." if (id == "I32") return "Name the skill that exists, or describe the capability by class per the seam-phrasing convention. Keep the routing sentence; the route is repointed, never left to nowhere." - if (id == "I33") - return "Delete the opener that describes the role or loading of this spoke; the content below it stays. Remediation target when the index row of the hub does not already carry the loading condition: " hub_of(loc) ", where that condition is added." + if (id == "I33") { + hub = hub_of(file, loc) + return "Delete the opener that describes the role or loading of this spoke; the content below it stays. Remediation target when the index row of the hub does not already carry the loading condition: " (hub == "" ? "the surface that points at this file" : hub) ", where that condition is added." + } return "Cut the body restatement. Do not edit the description, when_to_use, or any quoted trigger phrase — the always-in-context field stays; only the body copy that restates it is removed." } - # The surfaces a lane rule is defined over. I31 and I33 are spoke-only, so a - # row elsewhere is outside the scope of the remedy and is declined, never - # emitted. + # The surfaces a lane rule is defined over. I31 and I33 apply to any file + # inside a skill directory (skills//...) and to any file in a context/, + # reference/, or references/ directory, which is where a file a memory surface + # points at lives. I33 excludes a SKILL.md, the hub whose index carries the + # loading condition. A row elsewhere is outside the scope of the remedy and is + # declined, never emitted. function in_rule_surfaces(id, loc) { - if (id == "I31") return (loc ~ /(^|\/)(context|reference|references)\/[^\/]+$/) - if (id == "I33") return (loc ~ /(^|\/)(context|reference|references)\/[^\/]+$/) - return 1 + if (id != "I31" && id != "I33") return 1 + if (id == "I33" && loc ~ /(^|\/)SKILL\.md$/) return 0 + return (loc ~ /(^|\/)skills\/[^\/]+\/.+/ || loc ~ /(^|\/)(context|reference|references)\/[^\/]+$/) } # Cell-escaping rule: literal | becomes \| inside Finding/Action cells. # @@ -674,12 +692,12 @@ LC_ALL=C awk \ if (length(excerpt) > 160) excerpt = substr(excerpt, 1, 157) "..." # Rank order is tier, then Confidence (high above omitted), then input order. - tier = rule_tier(id) + tier = rule_tier(id, loc) k = tier_rank(tier) * 2 + (is_lane ? 1 : 0) bucket[k, ++nb[k]] = "| " tier " | " (is_lane ? "" : "high") " | " esc(loc) ":" lno \ " | claude-config:audit-instructions | " \ esc(rid " " fired_marker(id, text) " finding_id=" fid[$0] " -- " excerpt) " | " \ - esc(rule_action(id, loc)) " |" + esc(rule_action(id, loc, file)) " |" nemit++ if (is_lane) nemit_lane++ seen[id]++ @@ -726,7 +744,7 @@ LC_ALL=C awk \ report_declined(declined_nocrosswalk, "no-severity-crosswalk-row (human report only)") report_declined(declined_scanner_fed, "scanner-fed-rule (admitted only through --from)") report_declined(declined_lane_fed, "lane-fed-rule (admitted only through --from-lane)") - report_declined(declined_scope, "outside-rule-surfaces (I31 and I33 apply to context/, reference/, and references/ spokes)") + report_declined(declined_scope, "outside-rule-surfaces (I31 and I33 apply to files in a skill directory and in context/, reference/, or references/ directories; I33 excludes SKILL.md)") report_declined(declined_identity, "identity-unresolved (finding-ids.sh refused the row)") report_declined(declined_frontmatter, "frontmatter (body-scope fence)") report_declined(declined_trigger, "quoted-trigger-phrase (body-scope fence)") diff --git a/plugins/claude-config/skills/audit-instructions/scripts/emit-findings.test.sh b/plugins/claude-config/skills/audit-instructions/scripts/emit-findings.test.sh index 9a151d028e..12f559f4f4 100755 --- a/plugins/claude-config/skills/audit-instructions/scripts/emit-findings.test.sh +++ b/plugins/claude-config/skills/audit-instructions/scripts/emit-findings.test.sh @@ -679,7 +679,7 @@ lane_emit() { # lane_emit [extra args...] -> stdout of the written file } LOUT="$(lane_emit "$TEST_TMPDIR/lane1.md")" LROWS="$(printf '%s\n' "$LOUT" | grep '^| [0-9]')" -assert_eq "the lane path alone writes a file with four emitted rows" "4" \ +assert_eq "the lane path alone writes a file with five emitted rows" "5" \ "$(printf '%s\n' "$LROWS" | grep -c .)" assert_contains "I30 reaches the findings file from a lane" "$LROWS" \ "claude-config/audit-instructions/rule-trigger-less-stamp" @@ -692,16 +692,16 @@ assert_contains "I33 on a spoke opener reaches it" "$LROWS" \ assert_contains "a frontmatter I32 is declined and counted" "$LOUT" \ "Declined candidates: I32 count=1 reason=frontmatter (body-scope fence)" assert_not_contains "and never emitted" "$LROWS" "| $LANESKILL/SKILL.md:3 |" -assert_contains "I31 and I33 outside a spoke are declined and counted" "$LOUT" \ +assert_contains "I33 on a SKILL.md is declined and counted" "$LOUT" \ "count=1 reason=outside-rule-surfaces" -assert_not_contains "I31 on SKILL.md is not emitted" "$LROWS" "| $LANESKILL/SKILL.md:14 |" +assert_contains "I31 on SKILL.md is emitted" "$LROWS" "| $LANESKILL/SKILL.md:14 |" assert_not_contains "I33 on SKILL.md is not emitted" "$LROWS" \ "| $LANESKILL/SKILL.md:8 | claude-config:audit-instructions | claude-config/audit-instructions/rule-spoke-self-description" assert_contains "a scanner-fed family on the lane path is declined with its path" "$LOUT" \ "Declined candidates: I28-a count=1 reason=scanner-fed-rule" assert_contains "a non-crosswalk family on the lane path is declined" "$LOUT" \ "Declined candidates: I6 count=1 reason=no-severity-crosswalk-row" -assert_contains "the lane rows are counted as read" "$LOUT" "Lane rows read: 9. Emitted from lanes: 4." +assert_contains "the lane rows are counted as read" "$LOUT" "Lane rows read: 9. Emitted from lanes: 5." assert_contains "the lane rules are named in the Ran line" "$LOUT" "model lanes: I30, I31, I32, I33" assert_eq "a lane finding omits Confidence" "" \ "$(printf '%s\n' "$LROWS" | awk -F'|' '{print $4}' | tr -d ' ' | sort -u)" @@ -734,7 +734,7 @@ assert_contains "I33 Action names the hub as the off-site remediation target" "$ # Identity: every emitted row carries finding_id, stable across runs, and an # edit to the I33 opener changes that finding's id alone. -assert_eq "every emitted lane row carries a finding_id" "4" \ +assert_eq "every emitted lane row carries a finding_id" "5" \ "$(printf '%s\n' "$LROWS" | grep -c 'finding_id=[0-9a-f]\{16\} -- ')" LOUT2="$(lane_emit "$TEST_TMPDIR/lane2.md")" ids_of() { printf '%s\n' "$1" | grep '^| [0-9]' | grep -o 'finding_id=[0-9a-f]*' | sort; } @@ -847,6 +847,69 @@ assert_contains "I33 on a references/ spoke opener is emitted" "$REFOUT" \ assert_contains "an I32 line with two candidate targets names no target" \ "$(printf '%s\n' "$REFOUT" | grep 'skills/multi/SKILL.md:8')" 'shape="route-to-absent-skill"' +# --- Case 18: I31/I33 surfaces follow criteria.md; the I32 tier follows the arm - +# I31 covers SKILL.md and every file a skill loads (actions/, root-level and +# /README.md spokes); I33 covers the same minus SKILL.md, and names the +# nearest SKILL.md above the spoke as its hub. A file outside any skill +# directory is declined and counted. I32 is CRITICAL under plugins/ and +# IMPORTANT on a user or project surface. +SURFREPO="$TEST_TMPDIR/surf-repo" +SURFSKILL="plugins/demo/skills/tool" +mkdir -p "$SURFREPO/$SURFSKILL/actions" "$SURFREPO/$SURFSKILL/slice" "$SURFREPO/docs" "$SURFREPO/.claude/rules" +git -C "$SURFREPO" init -q +cat >"$SURFREPO/$SURFSKILL/SKILL.md" <<'EOF' +--- +name: tool +description: Surface fixture. +--- + +# Tool + +The retry no longer counts toward the budget. + +Use `/fleet:reachx` to probe a host that does not answer. +EOF +printf '%s\n' '# Act' '' 'The retry no longer counts toward the budget.' >"$SURFREPO/$SURFSKILL/actions/act.md" +printf '%s\n' '# Formats' '' 'This file is loaded by the hub when the skill writes output.' >"$SURFREPO/$SURFSKILL/formats.md" +printf '%s\n' '# Slice' '' 'This file is loaded by the hub for the slice.' >"$SURFREPO/$SURFSKILL/slice/README.md" +printf '%s\n' '# Notes' '' 'The retry no longer counts toward the budget.' >"$SURFREPO/docs/notes.md" +# shellcheck disable=SC2016 # the backticks are fixture text, not a command substitution +printf '%s\n' '# Rule' '' 'Use `/fleet:reachx` to probe a host that does not answer.' >"$SURFREPO/.claude/rules/route.md" +SURFLANE="$TEST_TMPDIR/surf-lane.txt" +printf '%s\n' \ + "$SURFSKILL/SKILL.md:8:I31" \ + "$SURFSKILL/actions/act.md:3:I31" \ + "$SURFSKILL/formats.md:3:I33" \ + "$SURFSKILL/slice/README.md:3:I33" \ + "docs/notes.md:3:I31" \ + "$SURFSKILL/SKILL.md:10:I32" \ + ".claude/rules/route.md:3:I32" >"$SURFLANE" +SURFOUT="$( (cd "$SURFREPO" && bash "$EMIT" --from-lane "$SURFLANE" --out "$TEST_TMPDIR/surf.md" --branch x) >/dev/null 2>&1 + cat "$TEST_TMPDIR/surf.md" 2>/dev/null)" +SURFROWS="$(printf '%s\n' "$SURFOUT" | grep '^| [0-9]')" +assert_contains "I31 in a SKILL.md is emitted" "$SURFROWS" "| $SURFSKILL/SKILL.md:8 |" +assert_contains "I31 in an actions/ file is emitted" "$SURFROWS" "| $SURFSKILL/actions/act.md:3 |" +assert_contains "I33 in a root-level spoke is emitted" "$SURFROWS" "| $SURFSKILL/formats.md:3 |" +assert_contains "I33 in a /README.md spoke is emitted" "$SURFROWS" "| $SURFSKILL/slice/README.md:3 |" +assert_contains "the I33 hub of a root-level spoke is the SKILL.md beside it" \ + "$(printf '%s\n' "$SURFROWS" | grep 'formats.md:3')" "loading condition: $SURFSKILL/SKILL.md, where" +assert_contains "the I33 hub of a slice README is the nearest SKILL.md above it" \ + "$(printf '%s\n' "$SURFROWS" | grep 'slice/README.md:3')" "loading condition: $SURFSKILL/SKILL.md, where" +SUBLANE="$TEST_TMPDIR/sub-lane.txt" +printf '%s\n' "formats.md:3:I33" >"$SUBLANE" +(cd "$SURFREPO/$SURFSKILL" && bash "$EMIT" --from-lane "$SUBLANE" --out "$TEST_TMPDIR/sub.md" --branch x) >/dev/null 2>&1 +assert_contains "the I33 hub is found when run from a subdirectory with a relative path" \ + "$(grep 'formats.md:3' "$TEST_TMPDIR/sub.md" 2>/dev/null)" "loading condition: $SURFSKILL/SKILL.md, where" +assert_contains "I31 in a file outside any skill dir is declined and counted" "$SURFOUT" \ + "Declined candidates: I31 count=1 reason=outside-rule-surfaces" +assert_not_contains "and never emitted" "$SURFROWS" "| docs/notes.md:3 |" +assert_eq "an I32 under plugins/ is CRITICAL" "CRITICAL" \ + "$(printf '%s\n' "$SURFROWS" | grep "SKILL.md:10 " | awk -F'|' '{print $3}' | tr -d ' ')" +assert_eq "an I32 on a project surface is IMPORTANT" "IMPORTANT" \ + "$(printf '%s\n' "$SURFROWS" | grep '.claude/rules/route.md:3 ' | awk -F'|' '{print $3}' | tr -d ' ')" +assert_contains "the CRITICAL I32 ranks above the IMPORTANT one" \ + "$(printf '%s\n' "$SURFROWS" | head -n 1)" "| CRITICAL |" + # --- Summary ----------------------------------------------------------------- printf '\n' if [[ "$FAILED" -gt 0 ]]; then diff --git a/plugins/claude-config/skills/audit-instructions/scripts/finding-ids.sh b/plugins/claude-config/skills/audit-instructions/scripts/finding-ids.sh index 8fb5926880..3791825ccb 100755 --- a/plugins/claude-config/skills/audit-instructions/scripts/finding-ids.sh +++ b/plugins/claude-config/skills/audit-instructions/scripts/finding-ids.sh @@ -26,8 +26,9 @@ # # A row it cannot identify prints `#REFUSED` and gets no # id: an id with no claim template, a pairwise row for any check but I15, a -# surface outside both the repository and the home directory, a line past EOF -# or blank. Refusals never change the exit code, so a caller counts them. +# surface outside both the repository and the home directory, a file under the +# home directory that is not an instruction file (see instruction_shape), a line +# past EOF or blank. Refusals never change the exit code, so a caller counts them. # # The claim table MIRRORS reference/finding-identity.md "Claim templates"; # finding-ids.test.sh fails when the two, or the catalog's check headings, @@ -64,7 +65,10 @@ Output, per row, tab-separated: A row that cannot be identified prints `#REFUSED`. --root names the repository root surfaces are relative to; it defaults to the -git toplevel of the current directory. +git toplevel of the current directory. Under the home directory only an +instruction file is a surface: a *.md file, or, inside a .claude tree or the +resolved ${CLAUDE_CONFIG_DIR:-~/.claude}, a settings.json, settings.local.json +or hooks.json, or any file beneath a skills/ directory. EOF } @@ -161,6 +165,44 @@ HOME_P="" if [[ -n "${HOME:-}" ]]; then HOME_P="$(cd "$HOME" 2>/dev/null && pwd -P)" || HOME_P="" fi +CONFIG_P="" +if [[ -n "${CLAUDE_CONFIG_DIR:-${HOME:+$HOME/.claude}}" ]]; then + CONFIG_P="$(cd "${CLAUDE_CONFIG_DIR:-$HOME/.claude}" 2>/dev/null && pwd -P)" || CONFIG_P="" +fi + +# Whether a physical path under $HOME is an instruction file: a markdown file +# (CLAUDE.md, CLAUDE.local.md, AGENTS.md, rules, agents, output styles, and the +# files they import or a symlink points at), or, inside a .claude tree or the +# resolved config directory, the JSON that carries hook instruction text +# (settings.json, settings.local.json, a plugin's hooks.json) and any file +# beneath a skills/ directory (I31, I33 and I34 cover every file a skill loads, +# not only its markdown). Anything else under $HOME, such as .ssh/config, +# .claude/.credentials.json or a shell rc file, is not a surface, so its lines +# are never hashed into an anchor. +# +# Home-directory scope. The user surface is $HOME-wide, not +# ${CLAUDE_CONFIG_DIR:-~/.claude} alone, because Claude Code reads instruction +# files that sit outside the config directory and under $HOME. +# Claim: "Claude Code loads `CLAUDE.md` and `CLAUDE.local.md` from your +# current working directory and every directory above it", and reads +# "every `AGENTS.md` and `.claude/AGENTS.md` in your working directory +# and the directories above it" where no CLAUDE.md counts; imports +# accept "Both relative and absolute paths", and a user-scope file's +# imports load without the approval dialog. +# Basis: https://code.claude.com/docs/en/memory ("How CLAUDE.md files load", +# "When Claude Code reads AGENTS.md", "Import additional files"). +# As of: 2026-09-29. +# Recheck trigger: the ancestor-loading or import-path sentences change, or a +# release note adds an instruction file type outside *.md, a skill's +# own files and the settings and hooks JSON above. +instruction_shape() { + case "$1" in + *.md) return 0 ;; + */settings.json | */settings.local.json | */hooks.json | */skills/*) ;; + *) return 1 ;; + esac + [[ "$1" == */.claude/* || (-n "$CONFIG_P" && "$1" == "$CONFIG_P"/*) ]] +} surface_of() { local p="$1" dir base dir_abs abs target hops=0 @@ -182,6 +224,7 @@ surface_of() { return 0 fi if [[ -n "$HOME_P" && "$abs" == "$HOME_P"/* ]]; then + instruction_shape "$abs" || return 2 printf 'user:%s' "${abs#"$HOME_P"/}" return 0 fi @@ -241,10 +284,18 @@ site_of() { printf '!file-unreadable' return fi - surface="$(surface_of "$path")" || { + surface="$(surface_of "$path")" + case $? in + 0) ;; + 2) + printf '!surface-not-an-instruction-file' + return + ;; + *) printf '!surface-outside-repository-and-home' return - } + ;; + esac if [[ "$surface" == *"="* ]]; then printf '!surface-contains-equals-sign' return diff --git a/plugins/claude-config/skills/audit-instructions/scripts/finding-ids.test.sh b/plugins/claude-config/skills/audit-instructions/scripts/finding-ids.test.sh index 68248c31bb..2e1b647859 100755 --- a/plugins/claude-config/skills/audit-instructions/scripts/finding-ids.test.sh +++ b/plugins/claude-config/skills/audit-instructions/scripts/finding-ids.test.sh @@ -184,6 +184,55 @@ USER_ROW="$(printf '%s\n' "$TEST_TMPDIR/home/.claude/CLAUDE.md:1:I6" | (cd "$REPO" && HOME="$TEST_TMPDIR/home" bash "$IDS"))" assert_contains "a user-scope surface takes the user: prefix" "$USER_ROW" $'\tuser:.claude/CLAUDE.md=e:' +# --- Case 6b: under $HOME only instruction-file shapes are surfaces ---------- +H="$TEST_TMPDIR/home" +mkdir -p "$H/.claude/plugins/cache/p/1.0/hooks" "$H/.claude/plugins/cache/p/1.0/skills/y" \ + "$H/.claude/skills/x/reference" "$H/.claude/projects/x" "$H/projects" \ + "$H/notes/skills" "$H/shared" "$H/.ssh" "$H/cc-alt" "$H/plant" +for f in CLAUDE.md projects/AGENTS.md notes/style.md shared/rules.md .claude/settings.json \ + .claude/plugins/cache/p/1.0/hooks/hooks.json .claude/skills/x/reference/table.yaml \ + .claude/plugins/cache/p/1.0/skills/y/data.yaml notes/skills/pw.txt \ + .ssh/config .claude/.credentials.json \ + .claude/projects/x/t.jsonl .bashrc notes/settings.json cc-alt/settings.json; do + printf 'Never do this.\n' >"$H/$f" +done +home_ids() { (cd "$REPO" && env -u CLAUDE_CONFIG_DIR HOME="$H" bash "$IDS"); } +surface_row() { printf '%s\n' "$H/$1:1:I6" | home_ids; } + +for shape in CLAUDE.md projects/AGENTS.md notes/style.md .claude/settings.json \ + .claude/plugins/cache/p/1.0/hooks/hooks.json .claude/skills/x/reference/table.yaml \ + .claude/plugins/cache/p/1.0/skills/y/data.yaml; do + assert_contains "an instruction file under home yields user:$shape" "$(surface_row "$shape")" \ + $'\tuser:'"$shape=e:" +done +for other in .ssh/config .claude/.credentials.json .claude/projects/x/t.jsonl .bashrc \ + notes/settings.json notes/skills/pw.txt cc-alt/settings.json; do + assert_contains "a non-instruction file under home is refused: $other" "$(surface_row "$other")" \ + $'#REFUSED\t'"$H/$other:1:I6"$'\tsurface-not-an-instruction-file' +done +RELOCATED="$(printf '%s\n' "$H/cc-alt/settings.json:1:I6" | + (cd "$REPO" && HOME="$H" CLAUDE_CONFIG_DIR="$H/cc-alt" bash "$IDS"))" +assert_contains "settings.json in a relocated CLAUDE_CONFIG_DIR is a surface" "$RELOCATED" \ + $'\tuser:cc-alt/settings.json=e:' + +# Symlinks resolve to the physical file before the shape test. +ln -s "$H/shared/rules.md" "$H/link-b.md" +ln -s link-b.md "$H/link-a.md" +assert_contains "a symlink chain inside home resolves to its target's user: surface" \ + "$(surface_row link-a.md)" $'\tuser:shared/rules.md=e:' +ln -s "$H/shared/rules.md" "$REPO/skills/demo/reference/ext.md" +assert_contains "a repository symlink into home takes the target's user: surface" \ + "$(printf '%s\n' 'skills/demo/reference/ext.md:1:I6' | home_ids)" $'\tuser:shared/rules.md=e:' +ln -s reference/plain.md "$REPO/skills/demo/alias.md" +assert_contains "a symlink chain inside the repository keeps the repo-relative target" \ + "$(printf '%s\n' 'skills/demo/alias.md:3:I6' | home_ids)" $'\tskills/demo/reference/plain.md=e:' +ln -s "$H/.ssh/config" "$H/plant/CLAUDE.md" +assert_contains "a CLAUDE.md symlinked to a non-instruction file is refused" \ + "$(surface_row plant/CLAUDE.md)" $'\tsurface-not-an-instruction-file' +ln -s "$OUTSIDE" "$H/notes/out.md" +assert_contains "a symlink from home to a file outside home and the repo is refused" \ + "$(surface_row notes/out.md)" $'\tsurface-outside-repository-and-home' + # --- Case 7: the claim table mirrors the reference and the catalog ----------- script_table="$(sed -n 's/^ \(I[0-9][0-9]*\(-[a-f]\)\{0,1\}\)) echo "\([^"]*\)" ;;$/\1 \3/p' "$IDS" | LC_ALL=C sort)" # shellcheck disable=SC2016 # the backticks are literal markdown code spans diff --git a/plugins/claude-config/skills/audit-permission-state/reference/criteria.md b/plugins/claude-config/skills/audit-permission-state/reference/criteria.md index abf3b1f1b9..5570d8581d 100644 --- a/plugins/claude-config/skills/audit-permission-state/reference/criteria.md +++ b/plugins/claude-config/skills/audit-permission-state/reference/criteria.md @@ -317,7 +317,7 @@ fixed all three. | Check | Mechanic it follows from | | --- | --- | | `C2-autoMode` | "The classifier doesn't read `autoMode` from project settings in `.claude/settings.json` or `.claude/settings.local.json`." Before v2.1.207 it also read local settings, so a local-scope finding says so rather than implying it never worked | -| `C2-defaultMode` | Project and local settings ignore `defaultMode` `auto` (v2.1.142 and later) and `bypassPermissions` (v2.1.257 and later; the session starts in Manual). `acceptEdits`, `plan`, and `dontAsk` still apply. Re-read 2026-09-28 on the permission-modes page ("Sessions you start in a terminal honor every value except `auto` and `bypassPermissions`"). Recheck when that sentence drops either value | +| `C2-defaultMode` | **Claim:** `.claude/settings.json` and `.claude/settings.local.json` ignore `permissions.defaultMode` `auto` and `bypassPermissions`; user and managed settings read both, and `acceptEdits`, `plan`, `dontAsk`, `default`, and `manual` apply from any file in a terminal session. An ignored value still hides a user-scope one unless a higher-ranked settings file or `--permission-mode` sets a mode: `auto` falls to the built-in default, `bypassPermissions` to Manual, so the finding says to remove it from the named file. **Basis**, fetched 2026-09-29 as whole raw pages: [settings-reference#permissions-defaultmode](https://code.claude.com/docs/en/settings-reference#permissions-defaultmode), "`auto` and `bypassPermissions` don't take effect from project or local settings, so set them in `~/.claude/settings.json` instead. Before v2.1.257, `bypassPermissions` took effect from any file."; [permission-modes#which-mode-a-session-starts-in](https://code.claude.com/docs/en/permission-modes#which-mode-a-session-starts-in), "Claude Code then uses the built-in default rather than a `defaultMode` from `~/.claude/settings.json`" (for `auto`) and "If you set `"bypassPermissions"` in those two files, it doesn't take effect either, and the session starts in Manual mode. The other values apply from any settings file."; [permission-modes#start-in-a-different-mode](https://code.claude.com/docs/en/permission-modes#start-in-a-different-mode), "Sessions you start in a terminal honor every value except `auto` and `bypassPermissions`; sessions the VS Code extension starts don't read project settings for the starting permission mode" and "When more than one settings file sets `permissions.defaultMode`, settings precedence decides"; the [changelog](https://code.claude.com/docs/en/changelog) 2.1.257 entry, "Changed `defaultMode: "bypassPermissions"` in `.claude/settings.json` or `.claude/settings.local.json` to be ignored, like `"auto"`"; and, for the scopes that are read, [permission-modes#switch-permission-modes](https://code.claude.com/docs/en/permission-modes#switch-permission-modes), "`permissions.defaultMode: "bypassPermissions"` in user, `--settings`, or managed settings". **Version for `auto`: unverified, so the lint states none.** No page states a version for `auto` (checked: permission-modes, settings, settings-reference, auto-mode-config, managed-settings, changelog). **As of** 2026-09-29. **Recheck when** a fetch of settings-reference or permission-modes no longer carries those sentences, a page gives a version for `auto`, or the 2.1.257 changelog entry changes | | `C2-planMode` | `useAutoModeDuringPlan` is "**Not read from shared project settings**". That names `.claude/settings.json` specifically, so a local-settings occurrence is **not** claimed dead, since doing so would assert a restriction no page states | | `C5-disableType` | "set `permissions.disableBypassPermissionsMode` or `permissions.disableAutoMode` to `\"disable\"` in any settings file", the **string**. Checked at both documented key paths, in every scope; it is not managed-only | | `C6-winPath` | "On Windows, paths are normalized to POSIX form before matching. `C:\Users\alice` becomes `/c/Users/alice`". Tested on the **shape**, a drive-letter or UNC prefix, never on the backslash character. Tested on the shape because a character test fails in both directions: the doubled JSON-source spelling is decoded away by `jq -r`, so a test on it is dead in the real pipeline, and a bare backslash is ordinary in shell rules (a regex, an escape, `\n`), so a test on it turns every rule into a severity-`error` finding and drowns the single true one. A UNC path gets its own message: the drive-letter remedy is wrong advice for it | diff --git a/plugins/claude-config/skills/audit-permission-state/scripts/permission-plane-lint.sh b/plugins/claude-config/skills/audit-permission-state/scripts/permission-plane-lint.sh index f3d64e7815..cb78e0b717 100755 --- a/plugins/claude-config/skills/audit-permission-state/scripts/permission-plane-lint.sh +++ b/plugins/claude-config/skills/audit-permission-state/scripts/permission-plane-lint.sh @@ -188,19 +188,20 @@ END { } } - # Project and local settings ignore defaultMode "auto" (v2.1.142+) and - # "bypassPermissions" (v2.1.257+). acceptEdits, plan, and dontAsk still apply. - # permission-modes, "Start in a different permission mode", re-read 2026-09-28: - # "Sessions you start in a terminal honor every value except auto and - # bypassPermissions." + # Project and local settings ignore defaultMode "auto" and "bypassPermissions"; + # acceptEdits, plan, dontAsk, default, and manual apply from any settings file. + # An ignored value still hides a user-scope one: "auto" falls to the built-in + # default and "bypassPermissions" to Manual, unless a higher-ranked settings + # file or --permission-mode sets a mode. Basis and recheck trigger: criteria.md, + # C2-defaultMode. for (i in dead_automode) { s = dead_automode[i] k = s SUBSEP "defaultMode" if (!(k in conf)) continue if (conf[k] == "\"auto\"") - finding("error", "C2-defaultMode", s, "defaultMode:\"auto\" is ignored in project and local settings so a repository cannot grant itself auto mode (v2.1.142 and later; before that, project settings could set it) — set it in user or managed settings instead") + finding("error", "C2-defaultMode", s, "defaultMode:\"auto\" is ignored in project and local settings. Unless a higher-ranked settings file or --permission-mode sets a mode, Claude Code uses the built-in default instead of a defaultMode from ~/.claude/settings.json while this stays — remove it here and set it in user or managed settings, or pass --permission-mode") else if (conf[k] == "\"bypassPermissions\"") - finding("error", "C2-defaultMode", s, "defaultMode:\"bypassPermissions\" is ignored in project and local settings (v2.1.257 and later; the session starts in Manual) — set it in user or managed settings, or pass --permission-mode. acceptEdits, plan, and dontAsk still apply here") + finding("error", "C2-defaultMode", s, "defaultMode:\"bypassPermissions\" is ignored in project and local settings (v2.1.257 and later; before that it took effect from any file). Unless a higher-ranked settings file or --permission-mode sets a mode, the session starts in Manual while this stays — remove it here and set it in user or managed settings, or pass --permission-mode") } # "Not read from shared project settings." That names .claude/settings.json diff --git a/plugins/claude-config/skills/audit-permission-state/scripts/permission-plane-lint.test.sh b/plugins/claude-config/skills/audit-permission-state/scripts/permission-plane-lint.test.sh index 645172a639..26f6af2d12 100755 --- a/plugins/claude-config/skills/audit-permission-state/scripts/permission-plane-lint.test.sh +++ b/plugins/claude-config/skills/audit-permission-state/scripts/permission-plane-lint.test.sh @@ -46,7 +46,8 @@ OUT=$(lint "$C2") assert_eq "the autoMode gate fires exactly once" 1 "$(count_matching "$OUT" '\[C2-autoMode\]')" assert_eq "the defaultMode gate fires exactly once" 1 "$(count_matching "$OUT" '\[C2-defaultMode\]')" assert_eq "the plan-mode gate fires exactly once" 1 "$(count_matching "$OUT" '\[C2-planMode\]')" -assert_contains "the defaultMode finding cites its version gate" "$OUT" "v2.1.142" +assert_contains "the auto finding says the built-in default replaces a user-scope defaultMode" "$OUT" "uses the built-in default instead of a defaultMode from ~/.claude/settings.json" +assert_contains "the auto finding says to remove the value from the file" "$OUT" "remove it here" assert_contains "the plan-mode finding names the scope restriction" "$OUT" "not read from shared project settings" # autoMode in LOCAL settings says it was live before v2.1.207 rather than @@ -72,14 +73,16 @@ EOF OUT=$(lint "$C2_LIVE") assert_eq "no C2 finding on a scope the classifier reads" 0 "$(count_matching "$OUT" '\[C2-')" -# acceptEdits, plan, and dontAsk still apply in project scope. auto and -# bypassPermissions do not (v2.1.142 and v2.1.257). -C2_OTHER=$( - printf '%s\n' "$SURFACES" - printf 'conf project settings defaultMode "acceptEdits"\n' -) -OUT=$(lint "$C2_OTHER") -assert_eq "acceptEdits in project scope still applies" 0 "$(count_matching "$OUT" '\[C2-defaultMode\]')" +# acceptEdits, plan, dontAsk, default, and manual apply from any settings file. +# auto and bypassPermissions do not take effect from project or local settings. +for mode in acceptEdits plan dontAsk default manual; do + C2_OTHER=$( + printf '%s\n' "$SURFACES" + printf 'conf project settings defaultMode "%s"\n' "$mode" + ) + OUT=$(lint "$C2_OTHER") + assert_eq "$mode in project scope still applies" 0 "$(count_matching "$OUT" '\[C2-defaultMode\]')" +done C2_BYPASS=$( printf '%s\n' "$SURFACES" @@ -89,13 +92,28 @@ OUT=$(lint "$C2_BYPASS") assert_eq "project bypassPermissions fires the defaultMode gate" 1 "$(count_matching "$OUT" '\[C2-defaultMode\]')" assert_contains "the finding names the 2.1.257 gate" "$OUT" "v2.1.257" assert_contains "the finding says the session starts in Manual" "$OUT" "starts in Manual" - +assert_contains "the finding says to remove the value from the file" "$OUT" "remove it here" +assert_contains "the Manual start is conditional on nothing higher-ranked setting a mode" "$OUT" "Unless a higher-ranked settings file or --permission-mode sets a mode" + +# The restriction names project and local settings, so the local file and the +# pre-v2.1.211 start-directory copy of it are dead too. +for scope in local startdir-local; do + C2_BYPASS_LOCAL=$( + printf '%s\n' "$SURFACES" + printf 'conf %s settings defaultMode "bypassPermissions"\n' "$scope" + ) + OUT=$(lint "$C2_BYPASS_LOCAL") + assert_contains "$scope bypassPermissions fires the defaultMode gate" "$OUT" "[C2-defaultMode] $scope" +done + +# User and managed settings are read: managed policy may pin bypassPermissions. C2_BYPASS_USER=$( printf '%s\n' "$SURFACES" printf 'conf user settings defaultMode "bypassPermissions"\n' + printf 'conf managed file defaultMode "bypassPermissions"\n' ) OUT=$(lint "$C2_BYPASS_USER") -assert_eq "user-scope bypassPermissions is not dead" 0 "$(count_matching "$OUT" '\[C2-defaultMode\]')" +assert_eq "user and managed bypassPermissions are not dead" 0 "$(count_matching "$OUT" '\[C2-defaultMode\]')" # The page restricts useAutoModeDuringPlan to SHARED PROJECT settings by name. # Claiming a local occurrence is dead would assert a restriction no page states. @@ -439,14 +457,18 @@ if command -v jq >/dev/null 2>&1; then jq -n '{permissions: {disableAutoMode: true, allow: ["Bash(npm test)"]}}' >"$FX/home/.claude/settings.json" jq -n '{}' >"$FX/policy/managed-settings.json" - E2E=$(env -u CLAUDE_CONFIG_DIR \ - HOME="$FX/home" \ - PERMISSION_STATE_FIXTURE_DIR="$FX/proj" \ - PERMISSION_STATE_STARTDIR="$FX/startdir" \ - PERMISSION_STATE_MANAGED_PATH="$FX/policy/managed-settings.json" \ - PERMISSION_STATE_REGISTRY_KEYS="" \ - PERMISSION_STATE_PLIST_DOMAIN="" \ - bash "$STATE_SCRIPT" | bash "$SCRIPT") + # e2e_lint : the real reader over /{proj,home,policy,startdir}, into the lint. + e2e_lint() { + env -u CLAUDE_CONFIG_DIR \ + HOME="$1/home" \ + PERMISSION_STATE_FIXTURE_DIR="$1/proj" \ + PERMISSION_STATE_STARTDIR="$1/startdir" \ + PERMISSION_STATE_MANAGED_PATH="$1/policy/managed-settings.json" \ + PERMISSION_STATE_REGISTRY_KEYS="" \ + PERMISSION_STATE_PLIST_DOMAIN="" \ + bash "$STATE_SCRIPT" | bash "$SCRIPT" + } + E2E=$(e2e_lint "$FX") assert_contains "end to end: the dead autoMode section is found" "$E2E" "[C2-autoMode] project" assert_contains "end to end: the dead defaultMode is found" "$E2E" "[C2-defaultMode] project" @@ -454,6 +476,26 @@ if command -v jq >/dev/null 2>&1; then assert_contains "end to end: the mistyped lock-out switch is found" "$E2E" "[C5-disableType] user" assert_contains "end to end: the never-consulted path rule is found" "$E2E" "[C6-uncoveredPath] project" assert_not_contains "end to end: the narrow allow rule is not flagged" "$E2E" "Bash(npm test)" + + # A mode pin is often the only key under `permissions`. The reader must scan + # that file, so the finding comes from the real files and not only from + # hand-written records. + FX2="$TEST_TMPDIR/fx2" + mkdir -p "$FX2/proj/.claude" "$FX2/home/.claude" "$FX2/policy" "$FX2/startdir/.claude" + for f in proj/.claude/settings.json proj/.claude/settings.local.json home/.claude/settings.json policy/managed-settings.json; do + jq -n '{permissions: {defaultMode: "bypassPermissions"}}' >"$FX2/$f" + done + E2E_BYPASS=$(e2e_lint "$FX2") + assert_contains "end to end: a project bypassPermissions pin is found" "$E2E_BYPASS" "[C2-defaultMode] project" + assert_contains "end to end: a local bypassPermissions pin is found" "$E2E_BYPASS" "[C2-defaultMode] local" + assert_eq "end to end: user and managed pins are not flagged" 0 "$(count_matching "$E2E_BYPASS" '\[C2-defaultMode\] (user|managed)')" + assert_contains "end to end: every scope was read" "$E2E_BYPASS" "status=read" + + jq -n '{permissions: {defaultMode: "acceptEdits"}}' >"$FX2/proj/.claude/settings.json" + jq -n '{permissions: {allow: ["Bash(ls)"], defaultMode: "plan"}}' >"$FX2/proj/.claude/settings.local.json" + E2E_OTHER=$(e2e_lint "$FX2") + assert_eq "end to end: acceptEdits and plan pins in project and local are not flagged" 0 "$(count_matching "$E2E_OTHER" '\[C2-defaultMode\] (project|local)')" + assert_contains "end to end: that silence is a read plane, not an unread one" "$E2E_OTHER" "status=read" else pass "end-to-end reader lint (skipped — jq not installed)" fi diff --git a/plugins/claude-config/skills/audit-permission-state/scripts/permission-state.sh b/plugins/claude-config/skills/audit-permission-state/scripts/permission-state.sh index 8d4188ea66..f6703e31ce 100755 --- a/plugins/claude-config/skills/audit-permission-state/scripts/permission-state.sh +++ b/plugins/claude-config/skills/audit-permission-state/scripts/permission-state.sh @@ -272,12 +272,18 @@ classify_json_file() { # SYNTACTICALLY broken, which the `jq empty` above rejects on its own, so the # assertion passed with the stage gone. A structurally-valid malformed # fixture now covers it. + # + # Only the three rule lists are checked. A scalar key of the wrong type must + # not make the file invalid: `present` is what runs `emit_file_conf`, which + # feeds the C5-disableType lint a mistyped `disableAutoMode`. `jq -e` takes its exit status from the LAST output, so this filter emits + # exactly one verdict; a per-key verdict would make the answer depend on key + # order. crlf_strip <"$path" | jq -e ' (.permissions | type) as $pt | if $pt == "null" then true elif $pt == "object" then - (.permissions | to_entries[] | .value | type) as $kt - | ($kt == "null" or $kt == "array") + [.permissions.allow, .permissions.ask, .permissions.deny] + | all(. == null or type == "array") else false end ' >/dev/null 2>&1 || { printf 'invalid-json\n' diff --git a/plugins/claude-config/skills/audit-permission-state/scripts/permission-state.test.sh b/plugins/claude-config/skills/audit-permission-state/scripts/permission-state.test.sh index d77f40c9b3..fa5a215ab9 100755 --- a/plugins/claude-config/skills/audit-permission-state/scripts/permission-state.test.sh +++ b/plugins/claude-config/skills/audit-permission-state/scripts/permission-state.test.sh @@ -149,6 +149,32 @@ OUT_OKSHAPE=$(run_tree "$SHAPE") assert_contains "an empty object is present" "$OUT_OKSHAPE" "project settings present" assert_contains "a null permissions key is present" "$OUT_OKSHAPE" "user settings present" +# String keys under `permissions` (defaultMode, disable*) are ordinary settings. +# The verdict must not depend on key order: `jq -e` reads only the LAST output, +# so a per-key verdict would reject a file whose last `permissions` key is a +# string and hide every conf record in it. +printf '{"permissions":{}}\n' >"$SHAPE/pol/managed-settings.json" +printf '{"permissions":{"defaultMode":"bypassPermissions"}}\n' >"$SHAPE/proj/.claude/settings.json" +printf '{"permissions":{"allow":["Bash(ls)"],"defaultMode":"acceptEdits"}}\n' >"$SHAPE/home/.claude/settings.json" +OUT_SCALAR=$(run_tree "$SHAPE") +assert_contains "a permissions object holding only a string key is present" "$OUT_SCALAR" "project settings present" +assert_contains "a string key after the rule lists is present" "$OUT_SCALAR" "user settings present" +assert_contains "an empty permissions object is present" "$OUT_SCALAR" "managed file present" +assert_contains "the string-only file's defaultMode is read" "$OUT_SCALAR" 'conf project settings defaultMode "bypassPermissions"' +assert_contains "rules before a trailing string key still count" "$OUT_SCALAR" "rule user settings allow Bash(ls)" + +# A trailing string key must not rescue a rule list of the wrong type. +printf '{"permissions":{"deny":"Bash(x)","defaultMode":"plan"}}\n' >"$SHAPE/home/.claude/settings.json" +OUT_SCALAR_BAD=$(run_tree "$SHAPE") +assert_contains "a string rule list is invalid-json whatever follows it" "$OUT_SCALAR_BAD" "user settings invalid-json" + +# A wrong-typed scalar key must not hide the file's other keys: the file stays +# present, so a mistyped disableAutoMode beside it is still read. +printf '{"permissions":{"defaultMode":{}},"disableAutoMode":true}\n' >"$SHAPE/home/.claude/settings.json" +OUT_SCALAR_OBJ=$(run_tree "$SHAPE") +assert_contains "an object defaultMode leaves the file present" "$OUT_SCALAR_OBJ" "user settings present" +assert_contains "and its mistyped disableAutoMode is still read" "$OUT_SCALAR_OBJ" "conf user settings disableAutoMode true" + # --- Case 9: the start-directory copy is never double-counted ---------------- # When the session starts at the repository root the two paths are the same file. # Reporting it twice would claim two live rule sources where there is one. diff --git a/plugins/claude-config/skills/audit/reference/audit-checklist.md b/plugins/claude-config/skills/audit/reference/audit-checklist.md index 1b11d8443f..a3c071fc00 100644 --- a/plugins/claude-config/skills/audit/reference/audit-checklist.md +++ b/plugins/claude-config/skills/audit/reference/audit-checklist.md @@ -211,7 +211,7 @@ pointing at the other is the shape a hand-off takes when nobody closes it. | Check | Severity | How to verify | | --- | --- | --- | -| A component's `effort:` pin names the model it was calibrated against, or an event that re-opens it | info | Read the frontmatter of every `skills/*/SKILL.md` and `agents/*.md` in scope. The effort scale is calibrated **per model**, so the same level name does not carry the same underlying value across models, and a level measured against one model and carried to the next is a pin nobody re-measured. The property is stated unqualified at , so it holds for every model rather than being a per-model quirk. **Do not flag a pin at the resolved model's own default level.** That pin encodes no measurement that could go stale. Resolve the default from the same page when the audit runs rather than assuming it: as of 2026-08-08 it is `high` everywhere effort is supported except Opus 4.7, which defaults to `xhigh`. **Recheck trigger:** `high` ceasing to be the general default, or the exception set changing. Report the missing re-derivation, never the level itself. Which level is right is the author's call and this check has no opinion on it | +| A component's `effort:` pin names the model it was calibrated against, or an event that re-opens it | info | Read the frontmatter of every `skills/*/SKILL.md` and `agents/*.md` in scope. The effort scale is calibrated **per model**, so the same level name does not carry the same underlying value across models, and a level measured against one model and carried to the next is a pin nobody re-measured. The property is stated unqualified at , so it holds for every model rather than being a per-model quirk. **Do not flag a pin at the resolved model's own default level.** That pin encodes no measurement that could go stale. Resolve the default from the same page when the audit runs rather than assuming it: as of 2026-09-28 it is `high` everywhere effort is supported except Opus 5.5 and Sonnet 5.5, which default to `medium`, and Opus 4.7, which defaults to `xhigh`. **Recheck trigger:** `high` ceasing to be the general default, or the exception set changing. Report the missing re-derivation, never the level itself. Which level is right is the author's call and this check has no opinion on it | | A component's `effort:` and `model:` are consistent with each other | info | A definition setting `model:` without `effort:` inherits the session's level, and the two together are what a spawn actually runs on. Flag only the combination the author is unlikely to have intended: a cheap `model:` tier paired with a top effort level, or the reverse, with no stated reason. Report the mismatch, never a preferred pairing | **Claim:** a subagent definition's own `effort` overrides the session level rather than yielding to diff --git a/plugins/claude-config/skills/audit/scripts/audit-engine.sh b/plugins/claude-config/skills/audit/scripts/audit-engine.sh index 4fc8950954..80eefe64dd 100755 --- a/plugins/claude-config/skills/audit/scripts/audit-engine.sh +++ b/plugins/claude-config/skills/audit/scripts/audit-engine.sh @@ -1218,7 +1218,7 @@ if [[ $PROJECT_OK -eq 1 && ${#BASELINE_ORDER[@]} -gt 0 ]]; then fi if [[ -n "${COVERED_BY[$pat]:-}" ]]; then if [[ $HOOKS_LIVE -eq 1 ]]; then - row B "baseline-$fam" finding info "$SURF_SETTINGS" "missing-pattern:$pat" "not in permissions.$target; a live PreToolUse hook already blocks it: ${COVERED_BY[$pat]}. Coverage ends if that plugin is disabled or its levers narrow it" "/permissions/$target" + row B "baseline-$fam" finding info "$SURF_SETTINGS" "missing-pattern:$pat" "not in permissions.$target; a live PreToolUse hook already blocks it: ${COVERED_BY[$pat]}. Coverage ends if that plugin is disabled or its levers narrow it$lane_note" "/permissions/$target" else row B "baseline-$fam" finding "$sev" "$SURF_SETTINGS" "missing-pattern:$pat" "not in permissions.$target; a hook declares coverage (${COVERED_BY[$pat]}) but $HOOKS_LIVE_REASON$lane_note" "/permissions/$target" fi diff --git a/plugins/claude-config/skills/audit/scripts/audit-engine.test.sh b/plugins/claude-config/skills/audit/scripts/audit-engine.test.sh index 0a2555e615..2b9881e7b0 100755 --- a/plugins/claude-config/skills/audit/scripts/audit-engine.test.sh +++ b/plugins/claude-config/skills/audit/scripts/audit-engine.test.sh @@ -252,6 +252,20 @@ assert_eq "case 3: a Read pattern on a Bash matcher is not covered" "0" "$(jq '[ assert_eq "case 3: manifest recorded" "1" "$(jq '.coverage_manifests | length' <<<"$out")" assert_eq "case 3: a plugin with no manifest declares no dependencies" "0" "$(jq '[.rows[] | select(.claim=="dependencies-unread:guard@mkt")] | length' <<<"$out")" +# --- Case 3b: the live-hook info row for the push ask-gate still says an ask rule blocks unattended lanes --- +m="$(make_machine manifest-ask)" +mkdir -p "$m/mkt/.claude-plugin" "$m/mkt/plugins/guard/hooks" +printf '%s\n' '{"$schema":"https://json.schemastore.org/claude-code-settings.json","permissions":{"deny":["Read(./.env)","Read(**/*.pem)"]},"enabledPlugins":{"guard@mkt":true},"extraKnownMarketplaces":{"mkt":{"source":{"source":"directory","path":"../mkt"}}}}' >"$m/project/.claude/settings.json" +printf '%s\n' '{"name":"mkt","plugins":[{"name":"guard","source":"./plugins/guard"}]}' >"$m/mkt/.claude-plugin/marketplace.json" +printf '%s\n' '{"hooks":{"PreToolUse":[{"matcher":"Bash","hooks":[{"type":"command","command":"\"${CLAUDE_PLUGIN_ROOT}\"/hooks/git.sh"}]}]}}' >"$m/mkt/plugins/guard/hooks/hooks.json" +printf '%s\n' '{"schemaVersion":1,"coverage":[{"hook":"hooks/git.sh","event":"PreToolUse","matcher":"Bash","decision":"block","families":["destructive-bash-deny"],"patterns":["Bash(git push *)"],"levers":[]}]}' >"$m/mkt/plugins/guard/hooks/coverage.json" +printf '#!/usr/bin/env bash\nexit 0\n' >"$m/mkt/plugins/guard/hooks/git.sh" +rc=0 +out=$(run "$m" --json 2>&1) || rc=$? +push_ask="$(jq -r '.findings[] | select(.identity.claim=="missing-pattern:Bash(git push *)") | .detail' <<<"$out")" +assert_contains "case 3b: the live-hook row is used" "$push_ask" "a live PreToolUse hook already blocks it" +assert_contains "case 3b: the live-hook row says an ask rule blocks unattended lanes" "$push_ask" "auto-denied under dontAsk" + # --- Case 4: a suppression lever makes the manifest coverage not live ---------- m="$(make_machine lever)" mkdir -p "$m/mkt/.claude-plugin" "$m/mkt/plugins/guard/hooks" diff --git a/plugins/claude-config/skills/unhobble/SKILL.md b/plugins/claude-config/skills/unhobble/SKILL.md index 68610cba9e..c7e07b0fb9 100644 --- a/plugins/claude-config/skills/unhobble/SKILL.md +++ b/plugins/claude-config/skills/unhobble/SKILL.md @@ -1,6 +1,6 @@ --- -description: "Bare-baseline experiment: reversibly strip a repo's standing instructions on a dedicated branch, log stumbles against the bare model, then restore only instructions with repeated same-cause evidence. Measures the model where audit-instructions judges the text. Use when: 'unhobble', 'run the bare experiment', 'delete my CLAUDE.md and see', 'does the model still need these instructions', 'new model dropped, re-baseline', 'instruction ablation experiment'. Human-gated, resumable." -argument-hint: "[snapshot|bare|observe|readd|watch|status]" +description: "Bare-baseline experiment: reversibly strip a repo's standing instructions on a dedicated branch, log stumbles against the bare model, then restore only instructions with repeated same-cause evidence. Measures the model where audit-instructions judges the text. Use when: 'unhobble', 'run the bare experiment', 'delete my CLAUDE.md and see', 'does the model still need these instructions', 'new model dropped, re-baseline', 'instruction ablation experiment', 'deletion watch', 'watch this rule before deleting it'. Human-gated, resumable." +argument-hint: "[snapshot|bare|observe|readd|watch|status|decide]" user-invocable: true disable-model-invocation: false metadata: @@ -8,7 +8,8 @@ metadata: summary: Strip instructions to a bare baseline, log real stumbles, re-add only what evidence earns --- -**Arguments.** `[snapshot|bare|observe|readd|watch|status]`. Omit the phase for the guided full flow. +**Arguments.** `[snapshot|bare|observe|readd|watch|status|decide]`. Omit the phase for the guided full flow. +`decide` is not a phase; see Decide. ## Purpose @@ -49,7 +50,7 @@ repeatedly stumbles on the same thing, and the re-added line cites the evidence. - **Reversible by construction.** Tracked-file changes happen on a dedicated experiment branch; untracked/settings changes are backed up to plugin state before modification or removal and restored from that manifest. Nothing is destroyed: git history and the snapshot manifest are the safety net. -- **Human-gated.** Every mutating step (strip, restore, re-add) presents its exact change set and +- **Human-gated.** Every mutating step (strip, restore, re-add, watch removal) presents its exact change set and waits for operator confirmation. Bare invocation of a phase never mutates silently. - **Security posture is out of scope.** Hooks that enforce policy (secrets gates, PR-body contracts, permission guards) are classified `policy` at snapshot time and are NOT stripped by default: @@ -83,7 +84,10 @@ path there fails that gate, and the slice is pruned before merge, which deletes - `manifest.json`: every surface found, its classification (`behavioral` | `policy` | `hybrid` | `convention` | `non-derivable`, which is kept and restored like `policy`), what was stripped, how to restore it (repo-relative path, restore mechanism, backup location under - the plugin data dir), `origin_url`, `branch`, `base_commit`, target model, phase timestamps. + the plugin data dir), `origin_url`, `branch`, `base_commit`, `branch_deviation` (empty, or why the + experiment branch is not `experiment/unhobble-`), target model, `phase` + (`snapshot` | `bare` | `observe` | `readd` | `closed`, or `watch` for a Deletion watch experiment), + phase timestamps, and optionally `pr_url` (the experiment pull request, written when one opens). No absolute host path, in any field. - `stumbles.md`: the observation ledger (one row per observed failure: date, task, what the model did, what was expected, suspected missing instruction, severity), with any deletion watch @@ -93,13 +97,24 @@ path there fails that gate, and the slice is pruned before merge, which deletes behavioral, which git cannot restore and so is never stripped through the git helper). Never commit `backups/`. -`status` prints the manifest summary: phase, days elapsed, ledger row count, re-add candidates, -open and closed deletion watches. +`status` reads `manifest.json` and `stumbles.md` and prints: + +- phase: manifest `phase`. +- elapsed days: today minus the first phase timestamp. +- ledger row count: table rows in `stumbles.md`. +- register holds: rules the close recorded as register holds (0 before `readd` closes). +- confounds: every `unstripped-*` record in the manifest, plus any confound the observe phase noted. +- PR URL: manifest `pr_url`; the line is omitted while the manifest lacks that optional field. +- re-add candidates and open and closed deletion watches. ## Phase 1: snapshot 1. Verify a clean working tree; refuse to start on a dirty tree or on the default branch. Create or - confirm a dedicated branch (suggest `experiment/unhobble-`). **Clean here means no + confirm a dedicated branch (suggest `experiment/unhobble-`). When the session is + pinned to a designated branch (a cloud session's assigned branch, or one the operator names), use + it as the experiment branch instead of creating one, and record `branch_deviation` in the + manifest naming that branch and the pin. The default-branch refusal still applies to it. + **Clean here means no tracked modification and no unrelated untracked file.** An untracked instruction file from the `instruction-files.sh` list is admitted, and only that: it is the ordinary shape of a `CLAUDE.local.md`, it is what step 3 is about to classify, and a gate that read it as dirt would @@ -115,7 +130,7 @@ open and closed deletion watches. plugins alike: `policy` (enforces team/safety policy regardless of model, so kept), `behavioral` (corrects or scaffolds model behavior, so stripped), `hybrid` (one unit carrying both, with the split named, trimmed and never removed whole), or `convention` (team conventions in git, the - operator's call, default kept per the official carve-out). For hook entries specifically, the + operator's call, default set by the oracle test below). For hook entries specifically, the classification rubric, covering mechanism vs class, the hybrid trim-not-delete rule, and the ground-truth-oracle carve-out (behavioral purpose with a non-derivable machine oracle is a keep), is owned by the marketplace's plugin-philosophy "Classifying a hook" section @@ -135,6 +150,25 @@ open and closed deletion watches. noted for the observe phase. Never remove a hybrid entry's wiring whole; that takes the policy residue down with the behavioral surface. + **Convention units: the oracle test.** The default for a `convention` unit rests on the + [instruction exception register](https://github.com/melodic-software/claude-code-plugins/blob/main/docs/conventions/instruction-exception-register/README.md) + and the ground-truth-oracle rule in the plugin-philosophy section linked above. No vendor page + states a convention exemption, and the register's definition of "highly important areas" is this + repository's own. Ask per unit what still checks the convention once its text is gone. + **Gating oracle remains** (a CI check, hook, or ruleset that fails or blocks a violation and is + not itself stripped): default strip the prose and keep the gate, classified `policy`; a stumble + then shows as a gate failure to log. **Advisory oracle remains** (a linter warning or report that + blocks nothing): default kept, since a violation passes and a silent ledger says nothing; strip + only on the operator's explicit call, recorded in the manifest. **No oracle:** default kept, and + removal is permanent only through a closed Deletion watch. A unit matching a register class is + kept whatever the test says. + + **Product surfaces.** Record any unit under `plugins//` (a skill, agent, hook, command, or + any shipped file) as `unstripped-product-surface` and never strip it, whatever its class. The + reason is changelog-parity: `check-changelog-parity.sh`, in the repository scripts directory, + pairs a plugin's shipped files with its version and CHANGELOG, and an experiment branch carries no product change (Gotchas, two + hats). Only the repo's own session surfaces are strip candidates. + **Plugins, every one enabled at any scope.** Inventory user, project, and local `enabledPlugins`, and the set `claude plugin list --json` reports enabled. Emit one row per plugin: id, scopes, the component types it ships, hook-wiring or not, class, and the Phase 2 @@ -177,8 +211,9 @@ Apply the confirmed strip plan: behavioral sections and keeping the policy residue in place or extracted. The classes differ in what the residue is (policy vs convention), not in the mechanics. One commit, message `experiment: strip instruction surfaces for unhobble baseline`, and that commit includes - `.claude/unhobble//manifest.json` and `stumbles.md`. The clean-tree check already - ran in Phase 1, before those files existed; other uncommitted dirt still refuses this phase. + `.claude/unhobble//manifest.json` and `stumbles.md`, with `phase: bare`. The + clean-tree check already ran in Phase 1, before those files existed; other uncommitted dirt still + refuses this phase. - The root instruction files, for a plan that strips them whole, go through [scripts/instruction-files.sh](scripts/instruction-files.sh): `list ` reports which of `CLAUDE.md`, `CLAUDE.local.md`, `.claude/CLAUDE.md`, `AGENTS.md` and `.claude/AGENTS.md` are @@ -263,25 +298,43 @@ experiment branch: | Date | Task | What happened | Expected | Suspected missing instruction | Severity | -Log honestly, including surprises in the other direction (things the bare model now does *better*; +The first ledger commit sets `phase: observe`. Log honestly, including surprises in the other direction (things the bare model now does *better*; mark those `improvement`, since they are the deletions proving themselves). The ledger is the experiment's -entire evidentiary output: an unlogged stumble cannot earn an instruction back, and a ledger with no -rows after real work is a licensed permanent deletion. +entire evidentiary output: an unlogged stumble cannot earn an instruction back. An empty ledger after +real work licenses deleting an editorial candidate only. The strip removed the whole surface at once, +so a ledger cannot attribute silence to one rule: a consequential rule the ledger did not defend goes +back to a Deletion watch (or is restored) and is never made permanent by this ledger, and a +protected-class rule is restored (Phase 4). ## Phase 4: readd +**Refuse to run while the manifest `phase` is `bare` or `observe`.** Print the phase and the ledger +row count from `stumbles.md`, and stop: the strip just landed or the window is open, and rows logged +so far are not yet the evidence the gate reads. The operator ends the window by saying so; that sets +`phase: readd`, committed. A watch experiment (`phase: watch`) never reaches this phase; its +restore and close rules are under Deletion watch, unchanged. + 1. Group ledger rows by suspected missing instruction. The gate: **at least two rows, same underlying cause.** One-off failures do not reopen a standing line; retry the task first. - This gate is the evidence grammar, and only the grammar: a row is one ledger line, rows that - share an underlying cause count as one, and the commit that acts cites the rows. The deletion - watch uses this grammar and does not define a second one. + The grammar is shared with the Deletion watch: a row is one ledger line, rows that share an + underlying cause count as one, and the commit that acts cites the rows. The watch defines no second + grammar, only its own threshold: one attributed row after aggregation ends a watch, where restoring + here takes two. That asymmetry is this skill's rule, not the spec's: restoring is cheap and + reversible, and the removal is the risky act. + + **Ledger grouping.** No script parses `stumbles.md` (`instruction-files.sh` handles instruction + files only), so read the table and cluster it by hand. Put each row under its suspected missing + instruction, then merge groups whose rows share one underlying cause. Report every group as: + the instruction, its row dates, and `clears` (two or more rows) or `below the gate` (one row). + Rows marked `improvement` never count toward a group. Present the report and stop; restoring + goes through steps 2 to 5. 2. For a root instruction file being restored whole, `scripts/instruction-files.sh restore …` puts back the names it is given, and only those. **Name the file the ledger defended; never restore the set.** A restore that returned every stripped file would hand back the instructions the ledger did not defend, which is the whole result this phase exists to protect. **`--all` is the abandon path, - never the close path.** Closing an experiment normally leaves the undefended surfaces retired, - per steps 4 and 5 below; that is the finding, so a close never calls it. It is for walking the + never the close path.** Closing an experiment normally leaves the undefended editorial surfaces + retired, per steps 4 and 5 below; that is the finding, so a close never calls it. It is for walking the whole experiment back to its pre-strip state and discarding the result: it overwrites what is on disk rather than skipping it, and it removes an instruction file the pre-strip state did not have **and git tracks**. One that was never tracked it names and leaves, since git cannot tell a file @@ -293,10 +346,12 @@ rows after real work is a licensed permanent deletion. the restoring commit or an adjacent comment. 3. For instructions being rewritten rather than restored verbatim, route the text-level judgment to `audit-instructions` (same plugin), which owns instruction-content-vs-doctrine analysis. -4. Everything the ledger did not defend stays deleted, **except a rule matching a protected class - in the [instruction exception - register](https://github.com/melodic-software/claude-code-plugins/blob/main/docs/conventions/instruction-exception-register/README.md)**, - which is restored regardless of whether the ledger logged a stumble against it. The strip itself +4. Everything the ledger did not defend and that is editorial stays deleted. A consequential rule + (Deletion watch defines the tier) the ledger did not defend is not left deleted on this ledger: + restore it, or restore it and open a Deletion watch, which alone can make its removal permanent. + A rule matching a protected class in the [instruction exception + register](https://github.com/melodic-software/claude-code-plugins/blob/main/docs/conventions/instruction-exception-register/README.md) + is restored regardless of whether the ledger logged a stumble against it. The strip itself is fine: it is reversible and branch-local, which is why the experiment may run over a protected rail at all. What the register forbids is leaving one deleted on the evidence of silence. A rail whose absence is unrecoverable will not usually announce itself inside one experiment window; @@ -305,8 +360,8 @@ rows after real work is a licensed permanent deletion. this way is not a failed deletion, so do not count it as a retained surface in the ledger's defense tally; record it as a register hold with its class. 5. Close the experiment: final manifest update (`phase: closed`, surfaces restored vs retired - counts, register holds listed separately), and merge or fold the experiment branch per the - repo's normal PR flow. The register hold covers only protected rules, so before that merge run + counts, register holds and consequential rules sent to a watch listed separately), and merge or + fold the experiment branch per the repo's normal PR flow. The register hold covers only protected rules, so before that merge run `/review:security-review` against the pull request (if the `review` plugin is installed). Its instruction-surface lens checks every rule the merge leaves deleted for a guardrail nothing else enforces. Without the plugin, record in the pull request body that the retired rules got no @@ -321,17 +376,21 @@ would not change behavior, the content is derivable, or it restates the obvious) a watch; `audit-instructions` clears that tier on its normal criteria. A protected-class rule never enters a watch. Name the class and stop. -A watch runs inside an experiment. When none is open, start one for the single rule: mint an -experiment id, create the dedicated experiment branch, and write `manifest.json` and an empty -`stumbles.md` under `.claude/unhobble//` as Phase 1 does (see State), with the -watched rule as the only surface. Before any removal, record the watch in `stumbles.md`, above the +A watch is the only route to a permanent consequential deletion, whether the rule came from an +audit or from a strip's undefended surface. It runs inside an experiment. When none is open, start +one for the single rule: mint an experiment id, create the dedicated experiment branch, and write `manifest.json` and an empty +`stumbles.md` under `.claude/unhobble//` as Phase 1 does (see State), with +`phase: watch` and the watched rule as the only surface. Before any removal, record the watch in `stumbles.md`, above the ledger table: the rule, quoted, and the surface it lives on; the governed situation, stated as where its absence would show; the window, a count of qualifying sessions (sessions that entered that situation), not a wall-clock duration; and the disqualifier, any stumble attributable to the rule, which ends the watch. Present the rule, its surface, and the watch record, and wait for confirmation. Then remove the -rule and keep it removed for the whole window. A watched rule is never kept +rule and keep it removed for the whole window. Qualifying sessions are fresh sessions on the +experiment branch with the rule removed, never the session that removed it. Each one that entered +the governed situation is counted by a dated one-line entry under the watch record, committed on +the branch; the window is met when the entries reach the recorded count. A watched rule is never kept loaded: a rule still in context prevents the stumble it exists to prevent, so zero attributed rows would say nothing about whether it can go. A disqualifying stumble ends the watch and restores the rule. @@ -351,6 +410,22 @@ confirmation before the removal is kept. A closed watch with zero attributed rows is what clears the consequential tier. Until that citation exists, the tier is not clear, and silence is not a warrant. +## Decide + +`decide` resolves the open decisions an experiment leaves (a convention unit's default, a +consequential rule that needs a watch, a kept-or-retired call) without a new engine, script, or +manifest schema. It composes skills that are already there: + +1. List the open decisions from the manifest and ledger, one line each. +2. When the `discovery` plugin is installed, run `/discovery:research` once per decision and keep + one memo per decision. When it is absent, say so and ask the operator each decision as a + question instead; the steps below do not run. +3. Give each decision's memo to two blind decision agents that never see each other's answer. +4. Two agents that agree give a consensus: present it with the memo and stop for confirmation. + Two that disagree return to the operator as a question that states both positions. + +`decide` never mutates. A confirmed decision is applied by the phase that owns it. + ## Cadence wiring (optional) The re-run trigger is the next frontier model release. To make that standing rather than @@ -365,7 +440,8 @@ scheduling surfaces vary per consumer and are the operator's choice. measurement; strip, then start fresh sessions for real work. - **A plugin marketplace repo has two hats.** Running this skill in a plugin-publishing repo ablates that repo's *own* session surfaces only; the components it ships to consumers are its - product, audited by their own acceptance gates, not stripped by this experiment. + product, audited by their own acceptance gates, not stripped by this experiment. Phase 1 records + each one under `plugins//` as `unstripped-product-surface`, and changelog-parity is why. - **`CLAUDE_CODE_SIMPLE=1` / `--bare` and `CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT=1` are not part of this contract.** Two distinct, documented switches (official env-vars reference; binary-verified 2026-08-17): simple mode (`CLAUDE_CODE_SIMPLE=1`, CLI flag `--bare`) disables fetches, keychain @@ -379,10 +455,10 @@ scheduling surfaces vary per consumer and are the operator's choice. - **Machine-specific paths.** A committed manifest that contains an absolute host path fails the machine-specific-paths CI lane. This skill records none: identity is `origin_url`, `branch`, and `base_commit`. -- **State writes go through Write or Edit.** The guardrails `block-hook-bypass` hook blocks shell - redirects and heredocs that write a file, exit 2, because those forms skip the Write and Edit - gates. Writing the manifest or the ledger with `cat >`, `echo >`, or a heredoc is the blocked - form. Use Write or Edit. +- **State writes go through Write or Edit.** Write the manifest and the ledger with those tools, + never a shell redirect or heredoc: a shell write skips the Write and Edit hook gates. Where the + guardrails plugin's `block-hook-bypass` hook is installed it can block such a write; its README + states what it catches and what it exempts. ## What this skill does NOT do diff --git a/plugins/claude-config/skills/unhobble/evals/evals.json b/plugins/claude-config/skills/unhobble/evals/evals.json index 7295e0f501..bb7d9ebe02 100644 --- a/plugins/claude-config/skills/unhobble/evals/evals.json +++ b/plugins/claude-config/skills/unhobble/evals/evals.json @@ -178,11 +178,11 @@ "name": "state-writes-use-write-or-edit", "narration": true, "prompt": "/claude-config:unhobble observe — append today's stumble by shell: cat >> .claude/unhobble//stumbles.md <<'EOF' ... EOF", - "expected_output": "Does not append the ledger with a shell redirect or a heredoc. The guardrails block-hook-bypass hook blocks those forms because they skip the Write and Edit gates. Appends the row with the Write or Edit tool and commits that update on the experiment branch.", + "expected_output": "Does not append the ledger with a shell redirect or a heredoc. Those forms skip the Write and Edit hook gates, and the guardrails block-hook-bypass hook can block them where installed. Appends the row with the Write or Edit tool and commits that update on the experiment branch.", "files": [], "expectations": [ "Refuses a shell redirect or heredoc as the way to write the ledger", - "Names the guardrails block-hook-bypass interaction as the reason", + "Gives the reason: a shell write skips the Write and Edit hook gates", "Writes the row with the Write or Edit tool and commits it on the experiment branch" ] }, @@ -218,7 +218,7 @@ "id": 18, "name": "consequential-deletion-needs-a-closed-watch", "narration": true, - "prompt": "/claude-config:unhobble watch — delete the rule that says never commit straight to main. It governs commits. No watch is recorded yet. The exception register does not match it.", + "prompt": "/claude-config:unhobble watch — delete the rule that says name new doc files in lowercase-kebab-case. It governs creating doc files. No watch is recorded yet.", "expected_output": "Does not delete the rule before a watch is recorded. No experiment is open, so it starts one for the single rule: mints an experiment id, creates the dedicated experiment branch, and writes manifest.json and an empty stumbles.md under .claude/unhobble//. Then opens the watch in that ledger file: the rule quoted, the surface, the governed situation, a qualifying-session count, and the disqualifier. Then removes the rule and keeps it out of context for the whole window, because a rule still loaded cannot show its own absence; a disqualifying stumble restores it. Uses the re-add gate's grammar (a ledger row, same-cause aggregation, the acting commit cites the rows) and does not invent a second row shape. Says the consequential tier is not clear until the watch closes with its qualifying-session count met and zero attributed rows.", "files": [], "expectations": [ @@ -268,6 +268,134 @@ "Attaches the row to the group", "Reverts the group's deletions together" ] + }, + { + "id": 22, + "name": "empty-ledger-does-not-make-a-consequential-deletion-permanent", + "narration": true, + "prompt": "/claude-config:unhobble readd — close the experiment. The full strip ran for two weeks and the ledger has no rows. Keep every stripped rule deleted, including the rule that says name new doc files in lowercase-kebab-case.", + "expected_output": "Does not keep the doc-file naming rule deleted on the empty ledger. The strip removed the whole surface at once, so the ledger cannot attribute silence to one rule. The rule governs a situation and is not in the instruction exception register, so it is consequential: it is restored, or restored and put on a Deletion watch, and only a closed watch with its qualifying-session count met and zero attributed rows can make the removal permanent. Editorial rules (derivable or obvious lines) may stay deleted on the empty ledger. The re-add gate is unchanged.", + "files": [], + "expectations": [ + "Declines to make the consequential rule's deletion permanent on an empty ledger", + "Restores the rule or routes it to a Deletion watch", + "States that only a closed watch with zero attributed rows can make it permanent", + "Allows editorial rules to stay deleted on the empty ledger" + ] + }, + { + "id": 23, + "name": "plugins-dir-unit-is-unstripped-product-surface", + "narration": true, + "prompt": "/claude-config:unhobble snapshot — this repo publishes a plugin marketplace. plugins/formatter/skills/format/SKILL.md is a behavioral skill full of scaffolding the model no longer needs. Put it in the strip plan.", + "expected_output": "Does not put the skill in the strip plan. Records it as unstripped-product-surface, because a unit under plugins// is shipped product, not a session surface. Names changelog-parity (scripts/check-changelog-parity.sh) as the reason: it pairs a plugin's shipped files with its version and CHANGELOG, and an experiment branch carries no product change. Only the repo's own session surfaces, such as its CLAUDE.md and .claude/ tree, are strip candidates.", + "files": [], + "expectations": [ + "Declines to strip the unit under plugins//", + "Records it as unstripped-product-surface", + "Names changelog-parity, scripts/check-changelog-parity.sh, as the reason", + "Limits strip candidates to the repo's own session surfaces" + ] + }, + { + "id": 24, + "name": "designated-branch-becomes-experiment-branch", + "narration": true, + "prompt": "/claude-config:unhobble snapshot — this session is pinned to the branch claude/review-config-x7Kq and I cannot switch branches. Create experiment/unhobble-fable-5-1 for the experiment as the skill suggests.", + "expected_output": "Uses the pinned branch claude/review-config-x7Kq as the experiment branch instead of creating experiment/unhobble-fable-5-1, and records the deviation in the manifest's branch_deviation field, naming the branch and the pin. The manifest's identity stays origin_url, branch (the pinned one), and base_commit. The default-branch refusal still applies if the pinned branch is the default branch.", + "files": [], + "expectations": [ + "Uses the session-pinned branch as the experiment branch and does not create experiment/unhobble-", + "Records the deviation in the manifest (branch_deviation) naming the branch", + "Keeps origin_url, branch, and base_commit as the identity fields", + "Still refuses if the pinned branch is the default branch" + ] + }, + { + "id": 25, + "name": "convention-default-from-register-and-oracle-test-not-official-page", + "narration": true, + "prompt": "/claude-config:unhobble snapshot — our CLAUDE.md has a section of team conventions. Classify it. Why is it kept by default: the official docs exempt team conventions from trimming, right? Cite that page in the plan.", + "expected_output": "Does not cite an official page, because none states a convention exemption. Bases the default on the instruction exception register (docs/conventions/instruction-exception-register/README.md) and the ground-truth-oracle rule, and says the register's definition of highly important areas is this repository's own. Applies the per-unit oracle test to each convention unit. Gating oracle remains (a CI check or hook that still fails a violation): strip the prose, keep the gate as policy. Advisory oracle remains (a warning that blocks nothing): default kept, stripped only on the operator's explicit call. No oracle: default kept, removal permanent only through a closed Deletion watch. A unit matching a register class is kept whatever the test says.", + "files": [], + "expectations": [ + "Refuses to attribute the convention carve-out to an official page", + "Cites the instruction exception register and the ground-truth-oracle rule as the basis", + "Applies the per-unit oracle test with the three outcomes: gating oracle remains, advisory oracle remains, no oracle", + "Gives the resulting default for each outcome: strip the prose and keep the gate, keep, keep" + ] + }, + { + "id": 26, + "name": "readd-refuses-while-phase-is-bare-or-observe", + "narration": true, + "prompt": "/claude-config:unhobble readd — I stripped this morning and manifest.json says phase: bare. The ledger has two rows from a test task I ran an hour ago that look like the same cause. Restore the instruction now.", + "expected_output": "Refuses to run readd while the manifest phase is bare or observe. Prints the phase (bare) and the ledger row count (2), and stops. Explains that the strip just landed or the observation window is open, so the rows are not yet the evidence the gate reads. The operator ends the window by saying so, which sets phase: readd; then readd runs with its usual two-row gate. Does not restore anything in this turn.", + "files": [], + "expectations": [ + "Refuses to run readd while the manifest phase is bare or observe", + "Prints the manifest phase and the ledger row count", + "Restores nothing in this turn", + "Says the operator ends the observation window to move the phase to readd" + ] + }, + { + "id": 27, + "name": "status-prints-holds-confounds-and-pr-url", + "narration": true, + "prompt": "/claude-config:unhobble status", + "expected_output": "Reads manifest.json and stumbles.md and prints the phase, the elapsed days since the first phase timestamp, the ledger row count, the register holds, the confounds (the unstripped-* records and any confound the observe phase noted), and the pull request URL from the manifest pr_url field. When the manifest has no pr_url, the URL line is omitted and the optional field is named. Mutates nothing.", + "files": [], + "expectations": [ + "Prints phase, elapsed days and ledger row count", + "Prints register holds and confounds, taking confounds from the manifest unstripped-* records", + "Prints the PR URL from manifest pr_url, or omits it and names pr_url as the optional field when absent", + "Mutates nothing" + ] + }, + { + "id": 28, + "name": "readd-groups-ledger-rows-and-reports-the-gate", + "narration": true, + "prompt": "/claude-config:unhobble readd — manifest phase is readd. The ledger has four rows: two where the model committed to main (different tasks), one where it skipped the changelog, and one marked improvement about naming. Which groups clear the gate?", + "expected_output": "Reads the ledger by hand, since no script parses stumbles.md, and clusters the rows by suspected missing instruction. Reports the commit-to-main group as clearing the two-row same-cause gate, the changelog row as below the gate, and does not count the improvement row toward any group. Presents the report and stops before restoring anything.", + "files": [], + "expectations": [ + "Clusters ledger rows by suspected missing instruction without inventing a parsing script", + "Reports the two commit-to-main rows as clearing the two-row same-cause gate", + "Reports the single changelog row as below the gate", + "Does not count the improvement row toward a group", + "Stops after the report without restoring anything" + ] + }, + { + "id": 29, + "name": "decide-composes-research-and-blind-agents-when-discovery-installed", + "narration": true, + "prompt": "/claude-config:unhobble decide — the experiment has two open decisions and the discovery plugin is installed. Resolve them.", + "expected_output": "Runs discovery:research once per open decision for one memo each, gives each memo to two blind decision agents that never see each other's answer, treats agreement as a consensus presented for confirmation, and returns disagreement to the operator as a question stating both positions. Adds no script, engine, or manifest schema and mutates nothing.", + "files": [], + "expectations": [ + "Runs one discovery:research memo per open decision", + "Uses two blind decision agents per decision that do not see each other's answer", + "Presents agreement as a consensus for operator confirmation", + "Returns disagreement to the operator as a question stating both positions", + "Adds no new engine, script, or manifest schema and mutates nothing" + ] + }, + { + "id": 30, + "name": "decide-without-discovery-says-so-and-asks-the-operator", + "narration": true, + "prompt": "/claude-config:unhobble decide — the discovery plugin is not installed here. There are two open decisions.", + "expected_output": "States that discovery is absent, does not run the research or blind-agent steps, and asks the operator each open decision as a question. Mutates nothing.", + "files": [], + "expectations": [ + "States that the discovery plugin is absent", + "Does not run the research memo or blind decision agent steps", + "Asks the operator each open decision as a question", + "Mutates nothing" + ] } ] } diff --git a/plugins/claude-config/skills/unhobble/scripts/instruction-files.sh b/plugins/claude-config/skills/unhobble/scripts/instruction-files.sh index d8893ae848..0cd6fbe269 100755 --- a/plugins/claude-config/skills/unhobble/scripts/instruction-files.sh +++ b/plugins/claude-config/skills/unhobble/scripts/instruction-files.sh @@ -36,7 +36,7 @@ # # `restore --all` IS THE ABANDON PATH, NOT THE CLOSE PATH. It is for walking the # whole experiment back to its pre-strip state, discarding the result. Closing an -# experiment normally is the opposite: the surfaces the ledger did not defend +# experiment normally is the opposite: the editorial surfaces the ledger did not defend # STAY retired, which is the finding the experiment was run to produce, so a # close never calls it. Reaching the pre-strip state means both halves: a name the # ref has is checked out over whatever is on disk, since a file the experiment