Skip to content

fix(claude-config): finish audit-instructions catalog gaps and dispatch plan #4656

Description

@kyle-sexton

Problem

One full /claude-config:audit-instructions run over a user root (~/.claude: CLAUDE.md, 5
reference files, 5 agents, 8 user skills; 9 Phase B lanes, 1 B2 pass, 7 Phase C verifiers, catalog
1.21.1) hit catalog rows that do not fit real defects, a dispatch shape that exceeds the skill's own
gate, and a slow pre-scan. All of it is still present on origin/main (946caf11), catalog
reference/criteria.md version 1.22.0. Paths below are under
plugins/claude-config/skills/audit-instructions/.

  1. Dangling citation. reference/criteria.md:1921-1923 (I31): "The skill-body row for this shape
    stops at SKILL.md". No such row exists in the catalog, .claude/rules/, docs/conventions/
    or skill-quality.
  2. Literal surface lists. I31 (criteria.md:1921-1922) and I33 (criteria.md:1955-1956) name
    reference/, context/, references/ with no "such as". Surfaces loaded on the same terms but
    named otherwise (a skill's actions/*.md, root-level spokes like formats.md, per-slice
    <slice>/README.md spokes, ~/.claude/references/*.md a CLAUDE.md points at) got split rulings:
    two lanes applied the rows, two skipped them, one verifier dropped, one demoted.
  3. I32 Detect (criteria.md:1942-1944) only knows plugins/<plugin>/skills/<skill>/SKILL.md. A
    user-scope skill routing to a skill that exists nowhere matched the defect but not the layout;
    the lane filed error, the verifier recalibrated to warning because the row gives no severity
    for that arm.
  4. I30 (criteria.md:1901) was stretched twice to cover an undated moving ref
    (**Commit**: \main`) and an undatedStable vXliteral that was really stale against NuGet; both were dropped because I30's Detect needs an as-of date. I12 (criteria.md:741) was stretched twice to cover a skill's false claim about what the operator's ownsettings.json` chains, which
    is config content, not harness behavior. Nothing routes a real defect that no row fits, so lanes
    stretch the nearest row.
  5. The surface-partition paragraph (criteria.md:140-146) accounts for I6 to I12, I15 to I28 and
    I13/I14 only; I29 to I35 are unaccounted.
  6. Dispatch shape. SKILL.md:225-226 fans out one lane per skill and gates at about 20 dispatches;
    SKILL.md:282-283 batches one verifier per surface. On the 8-skill root that is about 22; the run
    had to improvise batching to land at 17. The body prescribes no way to stay under its own gate.
  7. scripts/instruction-scan.sh:243-247 forks a printf | grep per raw hit for the I6 reject and
    I27 require filters, plus 11 greps per file, so process count grows with file and hit count. Over
    78 files (about 25k lines) it exceeded the 120 s Bash timeout on Windows Git Bash.

Evidence

Verified this pass (read-only):

  • Every citation above re-read with git show origin/main:<path>.
  • Pushed branch audit-instructions-catalog-surface-gaps is 8 commits over merge base 3175e45ca;
    origin/main is 28 commits ahead of that base. git merge-tree --write-tree origin/main origin/audit-instructions-catalog-surface-gaps merges clean.
  • Main moved claude-config to 0.49.1 and claude-ops to 0.62.6 since the branch forked; the branch
    bumps neither and its criteria.md frontmatter is still 1.22.0.

Carried from the lane that built the branch (its notes live in a gitignored local .work/ slice, so
the facts are restated here):

  • Scan timing on this Windows host (Git Bash, GNU grep 3.0, gawk 5.4.0), 78-file / 19,056-line
    fixture tree: base 2492 s, new 15 s, both 1052 rows. Host fork cost was about 2.6 s per bare grep
    under load, so wall time is recorded, not gated.
  • Process count (drift-immune, counted with a PATH shim that logs one line per grep): new scan is
    11 greps whatever the file count (78 files or 3, default or --body-only), plus one group
    subshell, one tr, one awk. Base by construction: 858 greps (11 per file) plus 1065 filter
    pipelines (1059 I6 hits + 6 I27 hits).
  • Equivalence, all cmp-identical against the base script: fixture tree default (1052 rows) and
    --body-only (1043); a real corpus of 104 files (every 10th of plugins/*/skills/**/*.md)
    default (919) and --body-only (908).
  • instruction-scan.test.sh: 102 of 102 passing after dc82f3134. Commit c3fdb9838 added cases
    afterwards; no recorded run covers them.
  • Design choices came from an unattended interview with one fresh-context validator per open
    question, not from Kyle.

Not reproduced: the original run's report
(~/.claude/plugins/data/claude-config-melodic-software/audit-instructions/.../last-audit.md) is
local-only and was not re-read; re-running the audit against the real ~/.claude is a live probe
and was not done.

Proposed approach

The work is on pushed branch audit-instructions-catalog-surface-gaps. That branch also carries
the sibling skillOverrides issue's three commits, which touch disjoint files and are agent-ready on
their own. Create a fresh branch from origin/main and cherry-pick this issue's commits in branch
order: e1789ec91, 697eda512, dc82f3134, c3fdb9838, 63970450d. What each commit does:

Commit What it does State
e1789ec91 criteria.md: I31 retitled to cover skill bodies and every file they load, dangling sentence removed, I31/I33 Surfaces stated by class with directory names as examples; I32 retitled "names a skill that does not resolve" with a user/project arm at warning; I30 and I12 gain Must NOT flag lines routing the stretched cases; new ## Out-of-catalog defects section; partition covers I6 to I35. SKILL.md: dispatch plan before dispatch (cap 9 lanes, about 2,500 lines per lane, one verifier per lane that produced proposals), central pre-scan sliced by file:, Phase C refutation for out-of-catalog rows, Phase D Out-of-catalog subsection that never reaches emit-findings.sh done, unreviewed
dc82f3134 instruction-scan.sh: each family greps all files once (-nHZ, chunked under the Windows command-line limit); one awk applies the I6/I27 filters, the --body-only fence and argument-order regrouping; 42 test lines done, equivalence-proven, unreviewed
c3fdb9838 grep --null instead of -Z (BSD grep reads -Z as decompress); order test covers a literal colon in a file name done, tests not re-run
63970450d SKILL.md: states why the central pre-scan gets an extended timeout done
697eda512 claude-ops setup/SKILL.md: corrects the --config scope consequence (a rerun at the wrong scope adds a second install record and enables the plugin there; the value always lands in user settings; a rejected value warns and exits 0). Grounded in docs/conventions/hook-config-delivery/README.md facts 9 and 10 (sandbox probe, 2.1.283). Unrelated to either item; keep it and give it its own CHANGELOG line done, unreviewed
d8240bfc9, 703652060, 8b47d6d17 inert skillOverrides detection tracked by the sibling issue (item 20260912-012851)

Remaining work for this issue:

  1. Resolve the dispatch-sizing conflict (below), then adjust the e1789ec91 Phase B text.
  2. Version bumps from main at PR time (at 2026-09-27: claude-config 0.49.1, claude-ops 0.62.6),
    the level chosen for this change alone; one CHANGELOG entry each; criteria.md frontmatter to
    1.23.0 with the PR date as last-updated.
  3. Re-run the suites and gates listed in the acceptance criteria.
  4. Fresh-context review of the whole branch diff.

Decision needed (why needs-decision): #4113/#4114 recorded, after two validation rounds, that
lanes are sized by a token budget derived at run time, plugins stay atomic and split by skill only
when over budget, and no line constant ships. The branch ships a lane cap of 9 and "about 2,500
lines per lane", and packs one skill per lane with siblings merged. Recommended: keep
plan-before-dispatch, one verifier per lane that produced proposals, the separate B2 verifier and
the central pre-scan (none of which #4114 contradicts); drop the 9-lane cap and the 2,500-line
constant and state that lane sizing is #4114's. Alternative: land the constants as an interim rule
that #4114 explicitly replaces. Either way, never two sizing rules.

Alternatives considered by the lane and declined: making I31/I33 list more directory names (still
literal); widening I30 and I12 instead of adding the out-of-catalog route (turns them into catch-all
rows).

Acceptance criteria

  • criteria.md I31 cites no row that does not exist and covers SKILL.md as well as spokes.
  • I31 and I33 Surfaces state the class, with directory names as examples, naming actions/,
    root-level spokes and per-slice README spokes.
  • I32 Detect has a user/project arm with a stated severity (warning); its Must NOT flag
    exempts names that could be bundled, claude.ai-synced (anthropic-skills:,
    ~/.claude/skills/synced/) or gated.
  • I30 and I12 each carry a Must NOT flag for the stretched case pointing to the out-of-catalog
    route; the route is admitted in criteria.md, refuted in Phase C, reported in its own Phase D
    subsection, and never reaches emit-findings.sh.
  • The surface-partition paragraph accounts for every row I6 to I35.
  • SKILL.md Phase B sizes lanes by exactly one rule, consistent with audit-instructions: token-budgeted lanes, --unattended, and --resume over per-lane run files #4114 (or names itself as
    the interim rule audit-instructions: token-budgeted lanes, --unattended, and --resume over per-lane run files #4114 replaces).
  • instruction-scan.sh output is byte-identical to origin/main's version over the test
    fixtures and a real corpus sample, with and without --body-only, including space- and
    colon-bearing paths; grep process count is constant in file count.
  • bash plugins/claude-config/skills/audit-instructions/scripts/instruction-scan.test.sh passes
    with the count reconciled against its case list.
  • scripts/check-changed-skills.sh <merge-base> and scripts/check-changelog-parity.sh --check-bump <merge-base> pass; shellcheck clean on touched shell; no U+2014 in added lines.
  • claude-ops CHANGELOG names the setup --config correction.

Constraints and gotchas

Context

Source: local handoff item 20260920-053728-audit-instructions-catalog-surface-gaps-and-dispatch-shape.md
(retired into this issue). Branch: audit-instructions-catalog-surface-gaps (pushed). Sibling:
the inert skillOverrides issue drafted from item 20260912-012851.
Related: #4113, #4114, #4115, #4116, #4119, PR #4628, PR #4121.

Commit 697eda512 also covers the claude-ops part of #4651; land it once and note it on both issues.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority: lowNice-to-have, cosmetic, or speculative; opportunistic.work-class: scopedA briefed fix or small feature; blast radius bounded by the brief, tests exist.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions