Skip to content

unreferenced_by_name: state the census contract, and stop asking the one language that cannot answer it - #2825

Merged
squid-protocol merged 4 commits into
mainfrom
fix/2806-orphan-invocation-model
Sep 6, 2026
Merged

squid-protocol merged 4 commits into
mainfrom
fix/2806-orphan-invocation-model

Conversation

@squid-protocol

@squid-protocol squid-protocol commented Sep 6, 2026 •

Copy link
Copy Markdown
Owner

Closes #2806. Part of #2812 (contract roadmap, Phase 3).

What #2806 asked for, and what the measurement changed

orphaned_logic reported every function in the file as unused for five languages — abap, jcl, m4, makefile, objective-c, all at 3.25 per file against a 2.50 corpus median. The issue proposed suppressing the metric for all five, on the reading that none of them reach a callable unit by name.

That premise held for one of them.

claim measurement
abap no invoke-by-name form PERFORM probe_branch is one. Real-world ABAP (language-crucible, abapGit) reads 8 unreferenced subroutines in 124 — a 6% rate, entirely through PERFORM
m4 no invoke-by-name form bare-name macro expansion is one
makefile no invoke-by-name form a prerequisite is one (main.mk's one wired probe already cleared)
objective-c no invoke-by-name form a selector send is one
jcl no invoke-by-name form confirmed. A job's EXEC steps run in the order they are written; nothing in JCL references a step to make it run

What the four had in common was the corpus, not the engine: every median language's main dispatches its three probes (python's entry() calls probe_branch/probe_io/probe_risk), and in those four main declared a dispatch/entry unit that called nothing. Planting each language's own idiom moves all four to the median with no engine change — that is the companion corpus PR.

What this PR does

1. The contract. docs/unreferenced_by_name_contract.md: one hit is one extracted callable unit whose name occurs nowhere in the file outside its own definition, six corollaries, and the 46-language audit. One test per corollary in tests/core_engine/test_unreferenced_by_name_contract_2806.py, plus engine rule 19 in how_to_add_a_language.md.

2. jcl declares invocation_model: "positional" and the census is not computed there — absent, not maximal. Before it, 334 of 376 crucible steps (89%) read unreferenced, and the 42 that cleared did so by accident, not invocation: a step named CREATE beside inline SQL in a SYSIN DD * block, COBOL beside the same word inside a DSN, CRTABS repeated seven times in one job. Both tails were noise. Every other language keeps the by_name default and is unchanged by construction.

3. The rename (#2806 shape (c)). orphaned_logic → unreferenced_by_name; columns state_slop_orphans → state_unreferenced, raw_state_slop_orphans → raw_state_unreferenced. Both old names asserted the count was dead weight, which corollary 3 says the measurement cannot claim — any other occurrence of the name clears the flag, including a mention in a string literal. record_keeper.py renames a pre-#2806 database's columns in place; the values were never wrong, only the label.

One trap worth reading if you review nothing else

The first version of this put _invocation_model: "positional" in rules, beside _scope_filters. language_lens.py's _calibrate_lookup_maps compiles every string value inside rules into a regex (a defensive guard for definitions loaded from external JSON), so the detector received re.compile("positional") and compared unequal to "positional". Every unit test passed — they build the extractor from LANGUAGE_DEFINITIONS directly, which the lens has not touched — and the first golden-master bless was silently wrong. Only a real galaxyscope run on one JCL file showed it. The declaration is now a top-level language property beside lexical_family.

Verification

  • pytest tests/ green; signal_contract_audit.py --ci clean; audit_check.py clean (ruff baseline regenerated for pure line-shifts, mypy/dead-key/ast-accuracy OK).
  • Golden masters re-blessed, scoped with bless_scope.py and a full uncapped diff: 5,636 key renames (section 7's label), and 564 value differences, every one of them jcl — 150 Tech Debt Exposure (cics-java-jcics-samples 73.11 → 0.0), 34 API Exposure and 34 Documentation Exposure (the orphan → api conversion has nothing to convert), 33 Structural Mass, 267 topological X/Y/Z, and three global aggregates that are the same change summed. No other language moved in either fixture.
  • rosetta_audit.py against the companion corpus branch: 44/46 pass; the two that move are the plants' own cost, both re-blessed and ledgered there (abap args 4 → 7, m4 args 6 → 8).

Found on the way, filed rather than fixed here

Cross-repo

Companion corpus PR: keyword-rosetta#82 — the four plants, the two owed args cells, three ledger entries and the regenerated bias report. This PR merges first; the corpus PR goes green against engine main once it lands. Label: rosetta:rebless-owed.

🤖 Generated with Claude Code

squid-protocol and others added 2 commits September 6, 2026 16:20
`orphaned_logic` reported every function in the file as unused for five
languages, and #2806 proposed suppressing it for all five. Measurement
kept one of them. This states the contract the census has always had,
declares the one language that cannot be asked it, and renames the
signal to what it measures.

The contract (docs/unreferenced_by_name_contract.md, six corollaries,
one test each in tests/core_engine/test_unreferenced_by_name_contract_2806.py):
one hit is one extracted callable unit whose name occurs nowhere in the
file outside its own definition.

- jcl declares the top-level `invocation_model: "positional"`: a job's
  EXEC steps run in the order they are written and no JCL syntax can
  reference one, so the census is not computed there -- absent, not
  maximal. 334 of 376 crucible steps read unreferenced, and the 42 that
  cleared did so by colliding with a DSN fragment, inline SQL or a
  repeated step name. Both halves were noise. Every other language keeps
  the `by_name` default and is unchanged by construction. The key is a
  top-level language property, beside `lexical_family`, NOT a rule:
  language_lens.py compiles every string value inside `rules` into a
  regex, so a declaration there arrives as re.compile("positional") and
  reads as "not positional" -- with every unit test still green, since
  they build the extractor from LANGUAGE_DEFINITIONS directly.
- abap, m4, makefile and objective-c stay in the census: each has an
  ordinary invoke-by-name form (real-world ABAP reads a 6% unreferenced
  rate through PERFORM), and their 100% reading was keyword-rosetta's
  `main` never dispatching its probes -- a corpus plant gap, fixed in
  the companion PR with no engine change.
- the rename: equations `orphaned_logic` -> `unreferenced_by_name`,
  columns `state_slop_orphans` -> `state_unreferenced` and
  `raw_state_slop_orphans` -> `raw_state_unreferenced`. Both old names
  asserted the count was dead weight, which corollary 3 says the
  measurement cannot claim. record_keeper renames a pre-#2806 database's
  columns in place; the values were never wrong, only the label.

Found on the way and filed rather than fixed here: #2823 (haskell and
scheme read 0.25/file -- a module export list and a type signature are
declarations, not references), #2824 (abap args counts call-site
parameter keywords), and a comment on #2822 (cobol branch counts a bare
out-of-line PERFORM, which is why cobol's plant is deliberately not in
the companion PR).

Closes #2806
Part of #2812

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…absence

Scoped with tests/tools/bless_scope.py plus a full uncapped diff (the
runner's own output is capped at ~28 of 6,200 lines):

- 5,636 key renames, one pair per parsed file: section 7's label
  `Orphaned Logic` -> `Unreferenced By Name`. No attached value moved.
- 564 value differences, every one of them jcl: 150 Tech Debt Exposure
  (`cics-java-jcics-samples` 73.11 -> 0.0), 34 API Exposure and 34
  Documentation Exposure (the orphan -> api conversion has nothing to
  convert), 33 Structural Mass, 267 topological X/Y/Z (the layout
  re-solves corpus-wide when any node's mass changes), and three global
  aggregates that are the same change summed -- avg_tech_debt 31.243 ->
  27.332, avg_documentation 7.942 -> 7.753, jcl impact 916.94 -> 877.94.

No other language moved in either fixture.

Part of #2806

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

🐦‍⬛ Muninn Security Scan

✅ No security issues found.

🐦‍⬛ Powered by Muninn · Skald Lab

squid-protocol and others added 2 commits September 6, 2026 16:35
The #2806 rename pushed `orphan_idx` past the line limit. Format check is
zero-tolerance; only this one file is touched (a bare `ruff format .`
reformats ~19 files main has never formatted).

Part of #2806

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Same rationale as #2765 committing rule_probe.py/bless_scope.py: this
session rebuilt a throwaway script and then lost calls to two traps the
workflow docs did not name.

- `tests/tools/census_probe.py` -- rule_probe.py's sibling for the
  signals that are NOT rules (`unreferenced_by_name`, `duplicate_logic`,
  `api_declared_orphans`, per-function `usage_status`), same
  snapshot/`--compare` shape so an audit table comes out of it directly.
- The pipeline-vs-unit-test trap, in CLAUDE.md's debugging section, the
  ci-push-checklist (before the bless step, where it would have been
  caught) and the probe's own docstring: a real scan goes through
  `language_lens.py`, which compiles every string value inside `rules`
  into a regex; tests and probes do not. A wrong golden-master bless is
  what that costs.
- The measurement ORDER, in rule-contract-audit Phase 0: read the
  crucible rate before the control corpus for every language an issue
  accuses. One number ended #2806's premise for four of its five
  languages. A corpus cell reading 100% defective is not evidence the
  language cannot answer -- check whether the plant asks.
- Two smaller ones: format the whole changed set BEFORE regenerating the
  ruff baseline (a rename touches files you never opened, and CI reads
  the line shift as a new finding -- two wasted pushes here), and re-run
  rosetta_audit.py to 46/46 before opening the PRs, because a signal
  change moves the cells DERIVED from it (jcl's api_orphan_credit).

Part of #2806

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@squid-protocol
squid-protocol merged commit 018edf0 into main Sep 6, 2026
32 checks passed
@squid-protocol
squid-protocol deleted the fix/2806-orphan-invocation-model branch September 6, 2026 20:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

rosetta:rebless-owed Intentionally moves keyword-rosetta counts; audit warns, corpus re-blesses after merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

orphaned_logic reports 100% dead code for five languages whose invocation model never names the callee — the family #2727 closed without covering

1 participant