Repository navigation
unreferenced_by_name: state the census contract, and stop asking the one language that cannot answer it - #2825
Merged
Conversation
`orphaned_logic` reported every function in the file as unused for five languages, and #2806 proposed suppressing it for all five. Measurement kept one of them. This states the contract the census has always had, declares the one language that cannot be asked it, and renames the signal to what it measures. The contract (docs/unreferenced_by_name_contract.md, six corollaries, one test each in tests/core_engine/test_unreferenced_by_name_contract_2806.py): one hit is one extracted callable unit whose name occurs nowhere in the file outside its own definition. - jcl declares the top-level `invocation_model: "positional"`: a job's EXEC steps run in the order they are written and no JCL syntax can reference one, so the census is not computed there -- absent, not maximal. 334 of 376 crucible steps read unreferenced, and the 42 that cleared did so by colliding with a DSN fragment, inline SQL or a repeated step name. Both halves were noise. Every other language keeps the `by_name` default and is unchanged by construction. The key is a top-level language property, beside `lexical_family`, NOT a rule: language_lens.py compiles every string value inside `rules` into a regex, so a declaration there arrives as re.compile("positional") and reads as "not positional" -- with every unit test still green, since they build the extractor from LANGUAGE_DEFINITIONS directly. - abap, m4, makefile and objective-c stay in the census: each has an ordinary invoke-by-name form (real-world ABAP reads a 6% unreferenced rate through PERFORM), and their 100% reading was keyword-rosetta's `main` never dispatching its probes -- a corpus plant gap, fixed in the companion PR with no engine change. - the rename: equations `orphaned_logic` -> `unreferenced_by_name`, columns `state_slop_orphans` -> `state_unreferenced` and `raw_state_slop_orphans` -> `raw_state_unreferenced`. Both old names asserted the count was dead weight, which corollary 3 says the measurement cannot claim. record_keeper renames a pre-#2806 database's columns in place; the values were never wrong, only the label. Found on the way and filed rather than fixed here: #2823 (haskell and scheme read 0.25/file -- a module export list and a type signature are declarations, not references), #2824 (abap args counts call-site parameter keywords), and a comment on #2822 (cobol branch counts a bare out-of-line PERFORM, which is why cobol's plant is deliberately not in the companion PR). Closes #2806 Part of #2812 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…absence Scoped with tests/tools/bless_scope.py plus a full uncapped diff (the runner's own output is capped at ~28 of 6,200 lines): - 5,636 key renames, one pair per parsed file: section 7's label `Orphaned Logic` -> `Unreferenced By Name`. No attached value moved. - 564 value differences, every one of them jcl: 150 Tech Debt Exposure (`cics-java-jcics-samples` 73.11 -> 0.0), 34 API Exposure and 34 Documentation Exposure (the orphan -> api conversion has nothing to convert), 33 Structural Mass, 267 topological X/Y/Z (the layout re-solves corpus-wide when any node's mass changes), and three global aggregates that are the same change summed -- avg_tech_debt 31.243 -> 27.332, avg_documentation 7.942 -> 7.753, jcl impact 916.94 -> 877.94. No other language moved in either fixture. Part of #2806 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This was referenced Sep 6, 2026
Closed
Contributor
Same rationale as #2765 committing rule_probe.py/bless_scope.py: this session rebuilt a throwaway script and then lost calls to two traps the workflow docs did not name. - `tests/tools/census_probe.py` -- rule_probe.py's sibling for the signals that are NOT rules (`unreferenced_by_name`, `duplicate_logic`, `api_declared_orphans`, per-function `usage_status`), same snapshot/`--compare` shape so an audit table comes out of it directly. - The pipeline-vs-unit-test trap, in CLAUDE.md's debugging section, the ci-push-checklist (before the bless step, where it would have been caught) and the probe's own docstring: a real scan goes through `language_lens.py`, which compiles every string value inside `rules` into a regex; tests and probes do not. A wrong golden-master bless is what that costs. - The measurement ORDER, in rule-contract-audit Phase 0: read the crucible rate before the control corpus for every language an issue accuses. One number ended #2806's premise for four of its five languages. A corpus cell reading 100% defective is not evidence the language cannot answer -- check whether the plant asks. - Two smaller ones: format the whole changed set BEFORE regenerating the ruff baseline (a rename touches files you never opened, and CI reads the line shift as a new finding -- two wasted pushes here), and re-run rosetta_audit.py to 46/46 before opening the PRs, because a signal change moves the cells DERIVED from it (jcl's api_orphan_credit). Part of #2806 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #2806. Part of #2812 (contract roadmap, Phase 3).
What #2806 asked for, and what the measurement changed
orphaned_logicreported every function in the file as unused for five languages — abap, jcl, m4, makefile, objective-c, all at 3.25 per file against a 2.50 corpus median. The issue proposed suppressing the metric for all five, on the reading that none of them reach a callable unit by name.That premise held for one of them.
abapPERFORM probe_branchis one. Real-world ABAP (language-crucible, abapGit) reads 8 unreferenced subroutines in 124 — a 6% rate, entirely throughPERFORMm4makefilemain.mk's one wired probe already cleared)objective-cjclWhat the four had in common was the corpus, not the engine: every median language's
maindispatches its three probes (python'sentry()callsprobe_branch/probe_io/probe_risk), and in those fourmaindeclared adispatch/entryunit that called nothing. Planting each language's own idiom moves all four to the median with no engine change — that is the companion corpus PR.What this PR does
1. The contract.
docs/unreferenced_by_name_contract.md: one hit is one extracted callable unit whose name occurs nowhere in the file outside its own definition, six corollaries, and the 46-language audit. One test per corollary intests/core_engine/test_unreferenced_by_name_contract_2806.py, plus engine rule 19 inhow_to_add_a_language.md.2. jcl declares
invocation_model: "positional"and the census is not computed there — absent, not maximal. Before it, 334 of 376 crucible steps (89%) read unreferenced, and the 42 that cleared did so by accident, not invocation: a step namedCREATEbeside inline SQL in aSYSIN DD *block,COBOLbeside the same word inside a DSN,CRTABSrepeated seven times in one job. Both tails were noise. Every other language keeps theby_namedefault and is unchanged by construction.3. The rename (#2806 shape (c)).
orphaned_logic→unreferenced_by_name; columnsstate_slop_orphans→state_unreferenced,raw_state_slop_orphans→raw_state_unreferenced. Both old names asserted the count was dead weight, which corollary 3 says the measurement cannot claim — any other occurrence of the name clears the flag, including a mention in a string literal.record_keeper.pyrenames a pre-#2806 database's columns in place; the values were never wrong, only the label.One trap worth reading if you review nothing else
The first version of this put
_invocation_model: "positional"inrules, beside_scope_filters.language_lens.py's_calibrate_lookup_mapscompiles every string value insiderulesinto a regex (a defensive guard for definitions loaded from external JSON), so the detector receivedre.compile("positional")and compared unequal to"positional". Every unit test passed — they build the extractor fromLANGUAGE_DEFINITIONSdirectly, which the lens has not touched — and the first golden-master bless was silently wrong. Only a realgalaxyscoperun on one JCL file showed it. The declaration is now a top-level language property besidelexical_family.Verification
pytest tests/green;signal_contract_audit.py --ciclean;audit_check.pyclean (ruff baseline regenerated for pure line-shifts, mypy/dead-key/ast-accuracy OK).bless_scope.pyand a full uncapped diff: 5,636 key renames (section 7's label), and 564 value differences, every one of them jcl — 150 Tech Debt Exposure (cics-java-jcics-samples73.11 → 0.0), 34 API Exposure and 34 Documentation Exposure (the orphan → api conversion has nothing to convert), 33 Structural Mass, 267 topological X/Y/Z, and three global aggregates that are the same change summed. No other language moved in either fixture.rosetta_audit.pyagainst the companion corpus branch: 44/46 pass; the two that move are the plants' own cost, both re-blessed and ledgered there (abapargs4 → 7, m4args6 → 8).Found on the way, filed rather than fixed here
_visibility_exportto yield more than one name per match, plus its own corpus plant.argscounts call-site parameter keywords (PERFORM ... CHANGING,CALL FUNCTION ... EXPORTING) as declared parameters — corollary 1 of the statedargscontract. This PR's corpus plant is what exposed it.branchcounts a bare out-of-linePERFORM(2,191 of 21,549 crucible hits), which is why cobol's plant is deliberately not in the companion PR: it would move a gated cell onto that defect to fix a cell already in band.Cross-repo
Companion corpus PR: keyword-rosetta#82 — the four plants, the two owed
argscells, three ledger entries and the regenerated bias report. This PR merges first; the corpus PR goes green against engine main once it lands. Label:rosetta:rebless-owed.🤖 Generated with Claude Code