Skip to content

docs(coverage-report): re-measure REPORT.md, and run the instrument it reports on - #8

Merged
lywinged merged 1 commit into
mainfrom
fix/report-figures-stale
Aug 27, 2026
Merged

lywinged merged 1 commit into
mainfrom
fix/report-figures-stale

Conversation

@lywinged

Copy link
Copy Markdown
Owner

Found by a repo-wide sweep for figures stated in prose against what the code measures, which is the instrument that found the 3 of 8 drift in the separation module. It has a sibling, and the sibling is worse.

Every figure in the report is unreproducible

coverage-report/scripts/ is referenced by no test and no workflow. The instrument was run by hand, once, and the document has reported what it said that day ever since.

Between then and now the verifier's obligations moved from failures.append(...) into an explicit RULES registry. mutation_report.py did not know the registry form, so it saw one site out of twenty-three and reported "every obligation is held by at least one vector" over it, exiting zero. That was fixed in #5. What is fixed here is the document:

stated in REPORT.md re-measured
obligations 21 23
held 21 23
pairs 210 253
triples 1330 1771
margin median 1 2

210 and 1330 are C(21,2) and C(21,3), which is how the whole set of figures can be seen to descend from one stale count.

The correction is not only arithmetic

§5 was titled "the one thing that should be fixed: there is no margin", on a distribution of twenty obligations at margin 1. The current distribution is twenty-two at 2 and one at 3. The section's own request had been carried out and the section was still asking for it.

It now carries both measurements and says what closed the gap, and it keeps its three failure modes as written, because the first of them is what then happened. §5 said:

A new rule is added with no vector ... provided the rule is written in the append("literal") shape the inventory recognises. Written as extend, +=, an f-string or a named constant, it joins the suite invisibly, and the suite then reports complete coverage of a rule set that no longer includes it.

That is a description of the registry refactor, written before it. Predicting a failure mode is not the same as being guarded against it, and nothing was.

§3.2 had the same shape. Its subject, obligations outside the mutation criterion for reasons of code shape, is still real; receipt_missing is no longer one of them, and the recommendation it made has since been carried out.

§3.1 and §3.3 needed their margins updated and their findings stand.

The guard, which is the report's own first recommendation

Adopt scripts/inventory_guard.py into CI. It finds nothing today ... Its value is prospective and its cost is one test.

It was never done, and not doing it is why the document drifted. Correcting the numbers alone leaves them free to drift again, which is precisely what happened the first time.

tests/test_report_figures_are_measured.py runs the measurement and pins both sides: the figures the tool produces, and the figures the document states. A figure in prose and a figure from a tool are two artifacts, and only something that reads both keeps them one.

Each half was shown load-bearing by causing the condition it exists to catch:

probe result
revert 23 of 23 in the document to 21 of 21 1 failed
revert all 253 pairs to all 210 pairs 2 failed
revert vectors: 48 to vectors: 24 1 failed
delete one entry from the verifier's RULES registry 2 failed

The pair and triple counts are asserted as C(sites, 2) and C(sites, 3) rather than transcribed, since transcribing them is how they went stale.

Checks

976 passed, 1 skipped. Ruff clean across src, tests and examples. A sweep for residual 21 / 210 / 1330 / 24 in the report returns nothing.


Generated by Claude Code

…t reports on

Every figure in `coverage-report/REPORT.md` was produced by a tool that nothing
runs. `coverage-report/scripts/` was referenced by no test and no workflow, so the
instrument was run by hand, once, and the document kept reporting what it said
that day.

Between then and now the verifier's obligations moved from `failures.append(...)`
into an explicit `RULES` registry. `mutation_report.py` did not know the registry
form, so it saw one site out of twenty-three and reported "every obligation is
held by at least one vector" over it, exiting zero. That was fixed separately.
What is fixed here is the document, which had been describing a measurement the
tool no longer produces:

    stated in REPORT.md      re-measured
    obligations       21              23
    held              21              23
    pairs            210             253
    triples         1330            1771
    margin median      1               2

210 and 1330 are C(21,2) and C(21,3), which is how the whole set of figures can be
seen to descend from one stale count.

The correction is not only arithmetic. Section 5 was titled "the one thing that
should be fixed: there is no margin", on a distribution of twenty obligations at
margin 1. The current distribution is twenty-two at 2 and one at 3, so the
section's own request had been carried out and the section still asked for it.
It now carries both measurements and says what closed the gap, and it keeps the
three failure modes as written, because the first of them is what then happened:
"a rule written in a shape the inventory does not recognise joins the suite
invisibly" is a description of the registry refactor, written before it.

Section 3.2 had the same shape. Its subject, two obligations outside the mutation
criterion for reasons of code shape, is still real; `receipt_missing` is no longer
one of them, and the recommendation it made has since been carried out.

`tests/test_report_figures_are_measured.py` runs the measurement and pins both
sides: the figures the tool produces, and the figures the document states. The
report's own first recommendation was to put this instrument into CI at a stated
cost of "one test". It was never done, and not doing it is the reason the
document drifted. Correcting the numbers alone would have left them free to drift
again.

Each half of the guard was shown load-bearing by causing what it exists to catch:
reverting any figure in the document to its previous value fails it, and deleting
one entry from the verifier's `RULES` registry fails it too.

976 passed, 1 skipped. Ruff clean.

Signed-off-by: Louielunz <48041247+lywinged@users.noreply.github.com>
@github-actions

github-actions Bot commented Aug 27, 2026 •

Copy link
Copy Markdown

❔ Contributor Check: UNKNOWN

Check Result
Profile UNKNOWN
Credential LOW
Overall UNKNOWN

Automated check by AgenTrust Contributor Check.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs-review:UNKNOWN Contributor check flagged UNKNOWN risk

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant