-
Notifications
You must be signed in to change notification settings - Fork 2
docs(adr): record that verification guards earn default-on by measured precision #1357
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
5 commits
Select commit
Hold shift + click to select a range
9f7e2fb
docs(adr): record that verification guards earn default-on by measure…
kyle-sexton a94af10
docs(adr): ground the duplicate-gate rule in one-mechanism-per-concern
kyle-sexton 4d4d1ed
docs(adr): disqualify guards on observed false positives, not absent …
kyle-sexton e20d3c5
docs(adr): require acceptable precision, and stop claiming a suite pr…
kyle-sexton a699f49
docs(adr): scope the no-firing exemption to zero findings, not low vo…
kyle-sexton File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
136 changes: 136 additions & 0 deletions
136
docs/adr/0003-verification-guards-earn-default-on-by-measured-precision.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,136 @@ | ||
| # Verification guards earn default-on by measured precision, not by plausible oracle | ||
|
|
||
| - Status: accepted | ||
| - Date: 2026-07-25 | ||
|
|
||
| ## Context | ||
|
|
||
| A program to add claim-verification guards to `guardrails` (#1270) scoped three, on the | ||
| reasoning that each had a mechanical oracle and no existing coverage: | ||
|
|
||
| | Candidate | Oracle | Tier claimed at scoping | | ||
| |---|---|---| | ||
| | Asserted repo-relative path that does not exist | filesystem test | Deterministic | | ||
| | Version string disagreeing with its owning manifest | manifest compare | Deterministic | | ||
| | `/plugin:skill` reference that does not resolve | glob the plugins tree | Detect-then-judge | | ||
|
|
||
| All three reasoned soundly from the oracle. Two did not survive contact with the repository. | ||
|
|
||
| The **version guard** was dropped before implementation. Its only in-repo surface — a | ||
| `## [<version>]` heading versus the manifest — is already covered deterministically by | ||
| `scripts/check-changelog-parity.sh --check-bump`, a required CI gate. An enumeration of the | ||
| residual prose surface (all `*.md` under `plugins/` and `docs/`, every string in every | ||
| `plugin.json` and `marketplace.json` outside the `version` field, `plugins/*/hooks/*.sh`, | ||
| `.github/**`, `scripts/`, `.claude/`) found only third-party versions no local manifest can | ||
| adjudicate, plus claim shapes a manifest compare reads *wrong*: historical | ||
| ("before the `0.6.0` split"), minimum floors ("implementation `0.9.0`+", true and documented | ||
| as explicitly *not* a version dependency), and planned values ("CREATE: `0.1.0`"). | ||
|
|
||
| The **path guard** was built, tested to 61 contract cases, reviewed, and then withdrawn on | ||
| measurement. Swept across all 975 tracked markdown files, each fed to the hook as a real | ||
| `PostToolUse` payload: | ||
|
|
||
| | | | | ||
| |---|---| | ||
| | Files firing | **231 / 975 = 23.7%** | | ||
| | Findings | **389** | | ||
| | True positives | **0** | | ||
|
|
||
| The oracle never misfired — no candidate resolved at the repo root. Every finding was a | ||
| scoping failure. 72% were consumer-project config paths (`.claude/**` and similar) that a | ||
| doc describes for a *consuming* repo and that a marketplace correctly lacks; its | ||
| first-segment gate passed only because this repo happens to carry same-named top-level | ||
| directories. Fixing the three dominant causes still left ~4% firing at zero true positives. | ||
|
|
||
| The **reference guard** shipped. Same corpus: 4 true positives, verified individually. Its | ||
| noise was 89% CHANGELOG rename entries — content the hook's own advisory calls correct as | ||
| written, because a CHANGELOG is append-only by contract and a rename entry must keep naming | ||
| the old command. Excluding CHANGELOGs moved it from 3.4% firing at 6% precision to **0.51% | ||
| firing at 57% precision**. | ||
|
|
||
| ## Decision | ||
|
|
||
| **A verification guard does not ship default-on until its firing rate and precision have | ||
| been measured against a real corpus.** A sound oracle is a necessary condition, never a | ||
| sufficient one. Specifically: | ||
|
|
||
| 1. **Measure before shipping, on the real corpus, at real scale.** Not a sample, not the | ||
| contract suite. A passing contract suite is evidence for the cases it contains — nothing | ||
| more; an input shape absent from the suite can still be misclassified. A corpus sweep is | ||
| what exposes the scoping. The path guard had 61 green cases and a 23.7% real-world firing | ||
| rate. | ||
| 2. **Report the number in the PR.** "It looks quieter now" is not a measurement. The | ||
| before/after firing rate and precision are the artifact that justifies default-on. | ||
| 3. **A guard that fires and is never right disqualifies itself**, however sound its oracle. | ||
| The disqualifying condition is **observed false positives with no true positive**, not the | ||
| absence of true positives on its own. A guard that fires on a quarter of writes and is | ||
| never right trains readers to ignore every advisory, including the ones that are right — | ||
| negative value, not low value. | ||
|
|
||
| **A guard that does not fire at all is not thereby disqualified.** A corpus may simply | ||
| contain no violation of the class, which leaves precision *undefined* rather than zero, and | ||
| a guard for a rare defect is exactly the kind worth having. The exemption is for **zero | ||
| findings only** — it is not a low-volume allowance, because precision does not improve with | ||
| scarcity: | ||
|
|
||
| | Sweep result | Reading | | ||
| |---|---| | ||
| | **zero findings** — did not fire at all | precision *undefined*, not zero. Inconclusive: establish detection by seeding known defects, then ship if the unseeded firing rate stays acceptable | | ||
| | fires at all, no true positive | disqualified. Precision is 0% whether the firing count is 389 or 1 — the path guard is the loud case, but a single false finding with no real one fails identically | | ||
| | fires AND precision is acceptable | shippable — the reference guard, 0.51% firing, 57% precision | | ||
|
|
||
| For a zero-finding sweep the burden shifts to **seeded defects**: plant instances the guard | ||
| should catch, confirm it catches them, and report the unseeded firing rate — which is zero | ||
| by construction in this branch — as the noise figure. Absence of evidence is not the | ||
| evidence of absence that rule 1 asks for. | ||
|
|
||
| **"Acceptable precision" is a judgment, and the ADR deliberately does not fix a | ||
| threshold** — the tolerable ratio depends on how costly a false advisory is in the surface | ||
| being guarded, and a single number would be false precision. But rarity is not itself | ||
| evidence: one real finding among a hundred false ones is a 1%-precision guard however | ||
| rarely it fires, and it fails for the same reason the path guard did at 23.7%. Scarcity | ||
| changes the volume of noise, never its ratio. State the measured precision, name the ratio you consider | ||
| acceptable for that surface, and justify it. The reference guard shipped at 57% with the | ||
| remaining noise attributed to a single identified cause — that shape of argument, not the | ||
| bare number, is what earns default-on. | ||
| 4. **Distinguish "wrong oracle" from "wrong scope" — by checking, not by assuming.** That | ||
| distinction decides whether a guard is deleted or re-filed for rescoping, so it is worth | ||
| establishing rather than inferring from a green suite. For the path guard it was | ||
| established per finding: each candidate was re-tested against the repo root and confirmed | ||
| genuinely absent, which is what made "the oracle is exact, the scope is wrong" a finding | ||
| instead of an assumption. It was re-filed (#1314). Had any candidate turned out to resolve, | ||
| that would have been an oracle defect the contract suite simply had no case for. | ||
| 5. **A guard whose only surface an existing gate owns does not ship at all.** A guard | ||
| duplicating `check-changelog-parity --check-bump` would fire correctly; it would simply | ||
| add no signal the required gate does not already produce, while adding a second place the | ||
| rule can drift. That is the **one mechanism per concern** principle | ||
| (`PLUGIN-PHILOSOPHY.md`'s validation section, owned upstream in `standards`) — redundancy, | ||
| not a silent skip. The deletion criterion for a future guard author is "supplies no | ||
| additional signal," which is a coverage test, not a correctness one. | ||
|
|
||
| ## Consequences | ||
|
|
||
| - Withdrawal is a normal outcome of the build, not a failure of it. Two of three candidates | ||
| in #1270 were withdrawn, one before implementation and one after full review. Both | ||
| withdrawals are recorded with their evidence (#1314 carries the sweep) so the ideas can be | ||
| rescoped rather than rediscovered. | ||
| - The cost is real: the path guard reached 61 contract cases and a full review round before | ||
| the measurement ran. Measuring earlier — right after the first working version, before | ||
| review — would have saved that. **The sweep belongs immediately after the guard first | ||
| works, not after it is polished.** | ||
| - This extends the verification-promotion discipline in | ||
| [ADR 0002](0002-default-on-ai-review-advisory-with-earned-promotion.md) one step earlier in | ||
| the lifecycle. That ADR governs promoting an advisory gate to blocking on demonstrated | ||
| precision; this one governs whether an advisory gate ships default-on at all, on the same | ||
| evidentiary basis. | ||
| - `docs/conventions/hook-precision/README.md` owns the over-fire discipline for a guard | ||
| already in the tree. This ADR is the pre-ship counterpart and defers to it thereafter. | ||
|
|
||
| ## Sources | ||
|
|
||
| - #1270 (scoping, amended twice), #1284 (the two-guard PR, closed), #1319 (the shipped | ||
| guard), #1314 (the withdrawn guard's measurement and rescope) | ||
| - `docs/conventions/hook-precision/README.md` — over-fire discipline | ||
| - `docs/PLUGIN-PHILOSOPHY.md` — one mechanism per concern (validation section) | ||
| - `melodic-software/standards`, `conventions/engineering/enforceability-tiers.md` — | ||
| classify the tier first, justify automation second | ||
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
For ordinary
PostToolUse:Editevents, the repository's precision convention requires scanning only the changed hunk (docs/conventions/hook-precision/README.md, lines 20–22), whereas feeding every complete tracked file through the hook measures static file prevalence under synthetic whole-file writes. Consequently, the reported 23.7% cannot be treated as a real-world firing rate: files containing an old candidate do not fire when unrelated hunks are edited, and actual edit frequency is absent. Replay representative Write/Edit payloads, or label this metric as whole-file prevalence rather than using it as the firing-rate evidence required by the ADR.Useful? React with 👍 / 👎.