Skip to content

fix(claude-config): repair the inert Too slow gate and drop the unsourced hook threshold - #4152

Merged
kyle-sexton merged 4 commits into
mainfrom
claude/audit-automation-gaps-5jho2o
Sep 13, 2026
Merged

kyle-sexton merged 4 commits into
mainfrom
claude/audit-automation-gaps-5jho2o

Conversation

@kyle-sexton

@kyle-sexton kyle-sexton commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Closes #4145

Summary

The audit-automation-gaps skill's Too slow quality gate could not fire for the most common hook
class, and the timing threshold it cited had no upstream basis and contradicted this repository's
own hook budget by more than an order of magnitude. Both were found by running the skill against
this repository and then verifying every claim against the official Claude Code documentation.

Fix

The gate was structurally inert for PostToolUse. It required Tool exceeds 30s AND the hook must block. The hooks reference states PostToolUse cannot block, because the tool already ran, so
the conjunct was never satisfiable and a PostToolUse formatter cleared the gate at any latency.
This repository has measured an in-repo PostToolUse Write at 13,225 ms. The gate now judges each
event against its actual cost: a blocked call where the event can block, turn latency where it
cannot.

The threshold was unsourced. No official page documents a hook latency budget; there is no
performance section in the hooks documentation at all. Meanwhile docs/conventions/hook-budget/README.md
sets 1 s typical and 2 s worst case per tool call and states the budget never relaxes. Timing now
resolves the consuming repository's documented budget first, and labels any fallback this skill
supplies as a house rule rather than guidance.

The stale figure lived in three files, not one. All are updated, including the eval that would
otherwise grade against a rule the skill no longer states:

File Was
SKILL.md "must complete in <15-30s", "exceeds 30s"
context/gap-analysis.md "fast enough for a per-edit hook (<15s)"
evals/evals.json eval 1 "commonly exceeds the 15-30s budget"

New spoke. context/hook-timing.md carries the budget-resolution ladder, the per-event cost
table, the documented levers (narrow the matcher, add an if rule, async: true, explicit
timeout), and dated upstream-fact records per the upstream-drift convention for the 600-second
command-hook timeout default and the PostToolUse no-block behavior. Putting the detail in a
spoke keeps SKILL.md from growing against its 200-line soft target.

Verification

Re-run after merging origin/main (head f34310b7):

check-skill.sh audit-automation-gaps    PASS, 0 errors, 2 warnings (both pre-existing)
check-changed-skills.sh origin/main     PASS for this skill (full-tree run; the one unrelated
                                        failure it reports is in a skill this diff does not touch)
check-changelog-parity.sh --check       rc=0
  ... --check-order                     rc=0  (93 changelogs, newest-first, no duplicates)
check-jsonschema (evals.schema.json)    ok
check-evals-quality.sh                  PASS, 0 warnings
check-purged-em-dashes.sh               rc=0
markdownlint-cli2                       0 issues in 4 files
typos / editorconfig-checker            clean
affected-tests.sh --run                 7 shell suites PASS; 12 Python/Node suites NOT RUN
                                        (documented exit-3 ecosystem path, not a failure)
grep for the stale figure               no match anywhere under the skill

The two remaining check-skill.sh warnings (no Gotchas surface, 200-line soft target) predate this
change and are recorded in #4146 rather than folded in here.

Merge note

origin/main had advanced 57 commits, leaving this branch un-mergeable. Both conflicts were in
claude-config and are resolved in f34310b7:

  • plugin.json: main rewrote the description and advanced to 0.44.0. Main's description is taken
    as-is and this branch's bump re-based on top as 0.44.1.
  • CHANGELOG.md: main shipped its own 0.40.41 (the resolve-convention-home BOM fix), colliding
    with the number this branch had used. Main's 0.40.41 is restored unchanged and this branch's
    entry moved to 0.44.1.

Related

Refs #4146 (the design batch, deliberately separated so this correction does not wait on contested
work; it also tracks the pre-existing warnings above)

🤖 Generated with Claude Code

https://claude.ai/code/session_0114NEwLXgKoQfaPCEcKU8R3

…rced hook threshold (#4145)

The `audit-automation-gaps` quality gate required a tool to exceed 30s AND the
hook to block. `PostToolUse` can never block, so a `PostToolUse` candidate
cleared the timing gate at any latency, including this repository's own measured
13,225 ms in-repo `PostToolUse` Write. The gate now judges each event against
its actual cost: a blocked call where the event can block, turn latency where it
cannot.

The 15-30 second threshold it cited has no upstream basis. No official Claude
Code page documents a hook latency budget, and the figure was 15-30x this
repository's own ceiling of 1s typical / 2s worst case per tool call. Timing now
resolves the consuming repository's documented budget first and labels any
fallback this skill supplies as a house rule. The stale figure is removed from
all three files that carried it, including the eval that would otherwise grade
against a rule the skill no longer states.

The new `context/hook-timing.md` spoke carries the budget-resolution ladder, the
per-event cost table, the documented levers, and dated upstream-fact records for
the 600-second command-hook `timeout` default and the `PostToolUse` no-block
behavior, per the upstream-drift convention.

Closes #4145

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0114NEwLXgKoQfaPCEcKU8R3
Base moved 57 commits. Two conflicts, both in claude-config:

- plugin.json: main rewrote the description (em-dash purge) and advanced the
  version to 0.44.0. Took main's description and re-based this branch's bump on
  top as 0.44.1.
- CHANGELOG.md: main shipped its own 0.40.41 (the resolve-convention-home BOM
  fix), colliding with the entry this branch had added at that number. Main's
  0.40.41 is restored unchanged and this branch's entry moved to 0.44.1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0114NEwLXgKoQfaPCEcKU8R3
@kyle-sexton
kyle-sexton marked this pull request as ready for review September 13, 2026 02:26
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-13T02:31:19.904270Z f34310b Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@claude

claude Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 5m 0s —— View job


/review:code-review — claude-config audit-automation-gaps timing fix

  • Gather context (PR description, changed files)
  • Run /review:code-review criteria against the diff
  • Post findings

Scope: reviewed all 6 changed files against origin/main per the CI code-review lane's criteria (correctness + maintainability; security out of scope for this lane).

Verification performed: independently re-fetched https://code.claude.com/docs/en/hooks.md (live, not from memory) to check every upstream claim this PR adds in context/hook-timing.md and the eval:

  • PostToolUse: "Exit code 2 isn't honored for this event. The tool already ran successfully, so the action can't be blocked" — matches the PR's claim exactly. The doc even has the matching escape hatch: "To surface a warning to Claude from a PostToolUse or PostToolUseFailure hook, exit 2 instead so Claude sees the stderr even though the tool already ran" — this is the exact basis for the row's "shows stderr to Claude" wording.
  • PermissionRequest: "Exit code 2 isn't honored for this event ... Deny through the decision object instead" — matches.
  • Timeout defaults: verbatim doc text is "Claude Code lowers the command, http, and mcp_tool default to 30 on UserPromptSubmit, PreModelSwitch, and PostModelSwitch, and to 10 on MessageDisplay" and "SessionEnd hooks share a 1.5-second budget ... up to 60 seconds" — matches the table in context/hook-timing.md row for row.
  • No performance/latency-budget section exists on the hooks page — consistent with the PR's "no upstream basis" claim.

All restated upstream facts in context/hook-timing.md check out against a live, verbatim read of the cited basis — no drift found.

Consistency checks:

  • No leftover references to the old "15-30s" / "exceeds 30s" threshold anywhere under the skill (SKILL.md, gap-analysis.md, evals.json, templates/checklist.md, scripts/*.sh all clean).
  • plugin.json version (0.44.1) and CHANGELOG.md's ## [0.44.1] entry are consistent with each other and with the described merge-conflict resolution against origin/main's 0.44.0/0.40.41.
  • Markdown anchors in hook-timing.md's Contents section all resolve to real headings.
  • The gate rewrite in SKILL.md (Too slow) is logically sound: it no longer requires a PostToolUse hook to block (which the docs confirm it structurally cannot), and it correctly routes to the resolved-budget ladder instead of a bare number.

No blocking or high-signal findings. This is a well-scoped, well-sourced fix — the claims that motivated it check out, and the diff doesn't introduce new inconsistencies. Two non-blocking, pre-existing check-skill.sh warnings (no Gotchas surface, 200-line soft target — SKILL.md was already at 210 lines before this diff and is now 214) are correctly left untouched here and are tracked in #4146 per the PR description.

@claude

claude Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 55s —— View job


Security review complete

  • Gather context (PR diff, changed files)
  • Apply skip gate (open PR? security-relevant surface? prior review?)
  • Review changed files for security issues
  • Post findings (or "no issues found")

Skip gate: PR is open, but the change has no security-relevant surface. All six changed files are documentation/config for the audit-automation-gaps skill:

  • plugin.json — version bump string only.
  • CHANGELOG.md — prose changelog entry.
  • SKILL.md — prose guidance on hook-timing quality-gate wording, no executable logic.
  • context/gap-analysis.md — prose checklist wording.
  • context/hook-timing.md — new reference doc (tables and prose describing budget-resolution logic and upstream facts); no scripts, no shell/command execution, no secrets or credentials referenced.
  • evals/evals.json — updated expected-output string for an eval prompt, plain text.

None of these introduce or touch command execution, authorization/access control, injection surfaces (command/SQL/path/template), token/secret handling, or GitHub Actions workflow configuration. This is purely documentation content advising a skill's own guidance text — no logic, trust-boundary, or data-flow change to review.

No security issues found. Nothing to report.

@github-actions

Copy link
Copy Markdown
Contributor

Last security-reviewed head: f34310b7a3027a8a6f35548e15c74d4f6a014365. On the next push, the relevance gate compares only the commits since this SHA; delete this comment to force a full re-review.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f34310b7a3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/claude-config/skills/audit-automation-gaps/context/hook-timing.md Outdated
Comment thread plugins/claude-config/skills/audit-automation-gaps/context/hook-timing.md Outdated
@github-actions

Copy link
Copy Markdown
Contributor

Claude has reviewed this PR 1 time. The lane skips further automatic reviews after 5; deleting this comment resets the count.

…he absence-claim sources

Two findings from the Codex review on #4152, both confirmed against the file.

The three-rung ladder could not produce a verdict in its most common case.
Phase 2.1 always measures, so a repository with no documented budget landed on
rung 2, which supplied a measurement but no threshold; rung 3's one-second
ceiling was reachable only when BOTH the documentation and the measurement were
absent, which never happens after a successful timing run. The `Too slow` gate
compares a measured cost against a resolved budget, so it had nothing to compare
against. The ladder is now two rungs: the measurement is named as the left side
of the comparison rather than a rung, and the house-rule fallback applies
whenever the repository documents no ceiling.

The absence claim named four documentation pages in prose with no pointers, so a
reader could not re-fetch the corpus that backs the central premise. Each page is
now linked, the check performed is stated, and the two nearest event-scoped
statements are named so the claim's boundary is visible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0114NEwLXgKoQfaPCEcKU8R3

Copy link
Copy Markdown
Contributor Author

Note on the red ci-status against f34310b7: it is a stale-head race, not a lane failure.

That run was the contract-only ci-status triggered by the draft-to-ready flip. It waited 360s for the full lane run on the same SHA, then failed with no successful ci-lanes status on f34310b7.... The reason is visible in its log: the full run it was waiting on (34733053973) was still in flight when the review-fix push landed, and the concurrency group cancelled it. A cancelled run posts no successful ci-lanes status, so the aggregator had nothing to read.

No lane failed an assertion. Nothing was skipped or disabled to get past it.

The current head 6cd2cfb3 has its own full run (34733309469) in flight, which supersedes it: hook-utils, changes, managed-files-guard and GitGuardian are green, lint and the four test-linux shards are running, and test-windows is skipped as usual. ci-status re-aggregates on that head once the lanes finish, so no re-run of the stale one is needed. If it fails again on the current head, that one is real and I will treat it as such.


Generated by Claude Code

The lint lane's purged-em-dashes gate declares
plugins/claude-config/skills/*/context/*.md already purged, so the new
hook-timing spoke's H1 was a regression the moment it landed. Rewrite it
to the colon form its sibling gap-analysis.md already uses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0114NEwLXgKoQfaPCEcKU8R3
@kyle-sexton
kyle-sexton merged commit 49912c6 into main Sep 13, 2026
12 checks passed
@kyle-sexton
kyle-sexton deleted the claude/audit-automation-gaps-5jho2o branch September 13, 2026 03:44
kyle-sexton added a commit that referenced this pull request Sep 13, 2026
…ngs, and guard the posture (#4159)

Closes #4146

## Summary

Implements the design batch for `/claude-config:audit-automation-gaps`:
an inventory that undercounted by two orders of magnitude, findings that
nothing persisted, gates driven by keyword counts mistaken for
frequencies, and a refusal gate that ran in the context which generated
the proposals. All nine items are addressed. The three rows marked needs
sign-off got researched verdicts rather than ratification, and one ships
as a refusal.

## Fix

**Inventory.** `scripts/inventory.sh` read only the project `.claude`
tree, reporting `Hook scripts: 3 / Skills: 0 / Agents: 0 / MCP servers: 0
/ Plugins enabled: 1` for a repository carrying 77 plugin roots, 271
skills and 93 wired handlers. That output is injected as pre-computed
context, so the model anchored on it before the audit began. It now emits
a per-scope table over all seven documented hook locations: a status
vocabulary where a scope it could not read never reads as zero, a
standing versus conditional split, enablement inputs plus a pointer
rather than a computed verdict, managed policy probed rather than
guessed, and a cloud-session note.

**Findings persistence.** New `scripts/findings-state.sh` stores the
verdict table, evidence and plans so `--implement` has something to read
later, keyed per project through the shared `lib/state-key.sh` so one
checkout never reads another's verdicts. The audit writes twice, once
when verdicts are presented and once after the human selects, so
`--implement` acts on what was chosen rather than on every candidate the
skill happened to pass.

**Incident-count discipline.** A `git log --grep` count drove two gates
directly. Both rows now treat it as a ceiling: a frequency claim needs a
sample reported with its denominator and sample size, while a ceiling
already under 5 percent still settles YAGNI without one.

**Shift-left carve-out** ships with three falsifiable conjuncts and the
hook budget as a hard gate, so a consumer documenting a budget with no
headroom keeps the REJECT.

**Hook enumerator: does not ship.** Recording the `/hooks` registry
verdict is the stated prerequisite, and this repository's
`native-references` convention puts that verdict in the human-gated
category. The body records the deliberate absence instead of a registry
row this change is not entitled to write.

Also: presence-gated routing to `overengineering:audit` both ways, a
routing line for the declined `ci` category, a Gotchas surface, a `## Next`
section, and a fresh-context judge for the refusal gate.

## Verification

Three independent review passes found sixteen findings, all closed. Four
were blocking and every repository gate passed them: a symlinked
`.claude/skills` reported `absent` and `0`; an unusable project root
printed this repository's counts under the requested root's name; five
concurrent writers on one run id produced five successes and one
surviving file; and a multi-document payload validated in full while
persisting only its first document.

The most useful finding was about the tests. Mutation testing showed
`inventory.test.sh` staying green against a script pinned to print 7, 999
or 0, so a suite that could not detect a confident wrong number was
guarding a script whose entire purpose is not producing one.

    inventory.test.sh        34 -> 125 checks, 32 mutations, all caught
    findings-state.test.sh  101 -> 156 checks, 16 mutations, all caught
    check-skill              PASS, 0 errors, 0 warnings (was 2 warnings)
    em-dash, changelog parity/order/bump, evals schema and quality  clean
    shellcheck, shfmt, portability                                  clean
    CI on the merged head                                           green

Keying is proved rather than asserted: two repositories writing the same
run id into one plugin-data root land in different keys and each reads
back its own, and a mutation removing the key from the path makes one
repository serve the other's artifact, so the suite fails if the keying
ever stops carrying weight.

Three figures in the issue did not survive checking: the
`overengineering` description is 1202 codepoints rather than 1217, the
200-line soft target it cites is not what the gate enforces, and it names
a documentation path renamed before this work began.

## Related

Refs #4145 (the factual-correction half, merged as #4152)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_0114NEwLXgKoQfaPCEcKU8R3
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

claude-config/audit-automation-gaps: hook-timing facts are unsourced and the Too slow gate cannot fire for PostToolUse

2 participants