Skip to content

claude-config/audit-automation-gaps: inventory undercounts, findings are not persisted, and the enforcement posture needs guarding #4146

Description

@kyle-sexton

Context

Companion to the factual-correction issue. These are design changes to
/claude-config:audit-automation-gaps, each surfaced by dogfooding the skill on this repository and
then put through two independent fresh-context reviews with the proposing rationale withheld. Three
rows are marked needs sign-off: reviewers disputed them, and they should not be ratified by
whoever picks this up.

The headline defect: the bundled scripts/inventory.sh reported Hook scripts: 3 / Skills: 0 / Agents: 0 / MCP servers: 0 / Plugins enabled: 1 for a repository containing 143 hook scripts, 258
skills, 13 agents and 77 plugins. That output is injected as pre-computed context at SKILL.md:14,
so the model anchors on it before doing anything. The script reads only the project .claude tree,
while SKILL.md Phase 1.1 already instructs the model to read user, project and local settings,
managed policy, every enabled plugin's hooks/hooks.json, and skill or agent frontmatter. The script
is behind its own skill's spec.

Proposed work

1. Inventory correctness

The hooks reference documents seven hook locations: user, project and local settings, managed policy,
plugin hooks/hooks.json, skill frontmatter, subagent frontmatter. Two mechanics must not be
conflated: hooks merge across settings levels, while enabledPlugins follows precedence
(project over user, local to opt out, managed can force) with defaultEnabled as the fallback.

  • Default output stays a per-scope count table, not a row dump. An all-scope enumerator emits 70+
    hook rows into pre-computed context before the skill starts.
  • Distinguish always-on from skill-scoped and subagent-scoped rows. The latter two are conditional,
    not part of the standing set, and reporting them undifferentiated repeats the
    present-versus-active error claude-ops:inventory already names.
  • Report an unreadable scope as unreadable, never as zero.
  • Emit enablement inputs plus a pointer rather than a computed verdict, matching what
    claude-ops:inventory already declines to compute.
  • Report managed policy as not-probed absent a sourced per-OS path.
  • Note that a cloud session does not read local user settings, so its effective set differs.

2. Hook enumeration (needs sign-off)

A --hooks mode emitting machine-readable hook rows was proposed. One reviewer showed the native
/hooks browser already displays event, matcher, type, source file and command, rendered from the
harness's own resolved set. docs/NATIVE-SURFACES.md has no /hooks entry, and the
native-references convention wants that verdict recorded before a component duplicates a native
surface. Scope any script to what /hooks does not provide, and record the registry verdict first.

3. Incident-count discipline

git log --grep counts drove two live gates directly. Measured here: 1182 of 2144 commits (55.1%)
matched markdownlint|lint, and 211 (9.8%) matched secret|gitleaks with zero real secret incidents
among them. A keyword count is a ceiling, not a frequency. Cite a frequency only after reading a
sample. Note the asymmetry: a ceiling already below 5% settles the YAGNI gate with no sample needed,
so only a PASS or an affirmative frequency claim requires one. Both gate rows need editing, not just
prose.

4. Shift-left carve-out on the Already enforced gate (needs sign-off)

Six formatter plugins here deliberately duplicate a CI check at edit time;
plugins/markdown-format/README.md says the hook is advisory and tells you to make CI your hard
gate. Read literally, the Already enforced gate rejects all of them.

Both reviewers rejected the first carve-out as drafted. Its test, that the higher rung does not run
on every edit, is true of every rung in the hierarchy, so it would reopen nearly every rejection and
invert the skill's stated REJECT default. No official documentation addresses CI-versus-hooks
placement, so any latency argument is a house rule. What the docs do support is determinism: hooks
are for what must happen every time with zero exceptions.

If a carve-out lands at all, the reviewers converged on three required conjuncts: the concern clears
the incident and YAGNI gates on a sampled count; the hook is advisory and non-blocking; and its
measured cost fits the consumer's remaining hook-budget headroom as a hard gate. Where a consumer
documents a budget and has no headroom, the verdict stays REJECT.

5. Findings persistence

--implement currently has nothing to read in a later session. Persist the verdict table, evidence
and implementation plans under the plugin data directory, following claude-config:audit-pass, which
already persists run state there via scripts/run-state.sh. Two implementation constraints:
docs/conventions/plugin-data-report-keying/README.md Rule 1 is [SPEC] and requires the
<component>/<state-key>/<filename> keying through the shared library, since an unkeyed artifact
would serve one repository's verdicts to another; and ${CLAUDE_PLUGIN_DATA} does not expand in the
Bash tool environment, so a resolved path must be passed as an argument. This also changes the
skill's tool surface and should be named as such.

6. Research trigger

Every candidate in the dogfood run was settled locally, and the skill's unconditional parallel
research fan-out did no work. Making it purely conditional was challenged: this very session's docs
pass changed the shape of two candidates after local evidence had already "settled" them. Make a
single batched docs fetch mandatory for any candidate whose mechanism is a Claude Code surface (hook
event semantics, async and timeout, MCP scope, frontmatter), and keep conditionality only for
non-harness external facts such as tool performance. Record how an external fact was settled when
research is skipped, so "not needed" stays distinguishable from "not done".

7. Sibling routing (needs sign-off)

overengineering:audit says it is "Not for proposing NEW automation", which is this skill's job, but
neither names the other. Add a presence-gated boundary line here, with the fallback and ownership
framing the seam-phrasing convention requires. On the overengineering side, body only: its
description is already 1217 codepoints against a 1024 spec maximum and sits in
scripts/skill-description-cap-baseline.txt, whose header states rows may only shrink.

8. Declined: a ci category

Considered and rejected. It would expand the skill's charter beyond Claude Code automation, collide
with overengineering:audit's ci-lanes layer and review:audit-enforceability's rung crosswalk,
and gate a ci candidate against a hierarchy in which CI is the denominator. Add a routing line
naming those two owners instead.

9. Authoring-rule debt surfaced along the way

plugins/skill-quality/scripts/check-skill.sh on this skill reports: no Gotchas surface, 210 lines
against a 200-line soft target, and same-context judgment language with no fresh-context delegation
at SKILL.md:194. There is also no ## Next section, which
.claude/rules/skill-bodies-state-current-rules.md requires. The third is the interesting one: a
skill whose whole value is refusing proposals runs its refusal gate in the context that generated
them. Recorded here, deliberately not solved, since it changes the skill's execution shape.

Also unaddressed: the skill's hook guidance reaches only PreToolUse and PostToolUse of 33
documented events, so a whole class of candidates cannot surface; and PermissionRequest ignores
exit code 2 entirely, denying through a JSON decision object instead, so a gate written as exit 2
there is silently inert.

Acceptance criteria

  • The inventory reports per-scope counts across all seven hook locations, marks conditional scopes,
    and reports an unreadable scope as unreadable.
  • Pre-computed context does not grow into a row dump.
  • The findings artifact is written under a keyed plugin-data path and --implement reads it back.
  • Any carve-out that ships states three falsifiable conjuncts and keeps the budget a hard gate.
  • Eval cases cover the persisted artifact and any revised gate.
  • The repo's skill gate reports no new warnings; a ## Next section and a Gotchas surface exist.

References

  • Hook locations, merge semantics, enabledPlugins precedence: https://code.claude.com/docs/en/hooks, https://code.claude.com/docs/en/settings-reference
  • docs/conventions/plugin-data-report-keying/README.md, docs/conventions/seam-phrasing/README.md, docs/conventions/native-references/README.md
  • plugins/claude-config/skills/audit-pass/scripts/run-state.sh

Metadata

Found by dogfooding /claude-config:audit-automation-gaps on this repository. Filed as raw intake;
not self-triaged. Items 2, 4 and 7 are marked needs sign-off because independent reviewers disputed
them; they warrant their own verdict rather than ratification as checklist rows.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-triageNot yet classified. Floor until a type and one priority tier are set.

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions