Repository navigation
feat(report): stop the risk surface claiming more than it measures (#3111, #3112, #3113, #3114) - #3145
Merged
Conversation
…3111, #3112, #3113, #3114) Four issues from one real-repo plausibility audit (IBM's zopeneditor-sample: COBOL/PL-I/JCL/HLASM/REXX, 51 files). None of them changes how anything is computed; all four change what the report CLAIMS. They share one root defect. The brief's Primary Risk Drivers line took the top four vectors by raw value, and two of the thirteen -- spec_match and documentation -- are "fraction NOT covered by a convention" meters that sit at ceiling on any codebase lacking that convention. Both read mode 100 / median 50 on the sample scan, so they occupied 2 of 4 slots on ALL TEN top-10 entries. A line whose entire purpose is to discriminate between files was reprinting two constants, and burying the ones that did differ (Mutation Surface, Guard Balance). #3111 -- spec alignment is now opt-in, `--spec-alignment`, default OFF. Absent from every display surface when off, NOT rendered as 0.0: a 0.0 asserts perfect spec alignment, which is the opposite of "not measured". The stored vector stays dense and numeric (the slot holds 0.0) because ~25 positional consumers require that -- gpu_recorder does int(v * 10), network_risk_sensor multiplies, dev_agent_firewall sums, record_keeper writes REAL columns -- and because it is already how the engine represents the history vectors it ablates. Honesty is enforced at the render surfaces, which is where the issue asked for it. #3114 -- documentation is reclassified, not disabled. It is still computed, still tabled (marked "(coverage)"), and still reported per file as "X% of unit weight undocumented", but it is no longer eligible for the driver narrative. Its own contract (#2908) defines it as a coverage gap over units, so coverage is where it belongs. Both are declared in analysis_lens as OPTIONAL_VECTORS / CONTEXT_VECTORS with a single inactive_vectors() resolver every renderer consults, so a second optional vector needs no renderer edits. A contract test asserts both name real RISK_SCHEMA entries and that the two families stay disjoint -- off and reframed are different asks. #3112 -- the Cumulative Risk composite is removed. It was sum() over the whole 13-entry risk_vector: independently scaled sigmoids added with no weighting and no unit, producing "Cumulative Risk: 556.29". Unrepairable in place, because ~31% of the sum was constant or dead (spec_match/documentation ceiling-defaulted; stability/churn ablated to zero in every scan today), and because summing the vectors re-created exactly the composite risk score that the #2982/#2991 validation record retired. BREAKING: --max-systemic-threat multiplied blast radius by that composite and now multiplies it by structural magnitude. Intent is unchanged and arguably better served ("fail when something structurally heavy is also widely depended upon"), and both factors are now unit-honest. Thresholds DO NOT carry over -- the scales differ -- and are deliberately not silently rescaled, because a fabricated conversion factor fails more quietly than an obvious re-tune. #3113 -- the brief leads with the answer. A new executive summary states the dominant languages, the load-bearing artifact, the top orchestrator and the heaviest artifact before any formula exposition, reusing #2556's zero-guard so a flat import graph is reported as flat rather than ranked by scan order. It independently reproduces the two facts that audit verified against source: INCLUDES/BALSTATS.inc at 3 inbound, JCL/RUN.jcl at 14 outbound. Sections 11 and 12 answered "which files deserve attention first" twice, and the one that ranked did so on the composite #3112 deleted. They are now one ranked list, ordered by structural magnitude, and each entry carries a Blast Radius sentence saying what a change would reach -- assembled from dependency edges the brief already computed but never stated as a consequence. Outbound comes from raw_imports, matching section 7 and the summary, because the resolved out_degree can read 0 directly above a line naming two imports. The 75-line vector lexicon and the non-predictive disclaimer move to Appendix A, verbatim. #3113 was explicit that the honesty is an asset and must be kept; it just should not be the wall a reader climbs to reach the findings. A pointer is left where section 2 was. The AI output directive that told consumers to look for "massive Cumulative Risk" now points at magnitude and blast radius and explicitly forbids summing the vectors. Verified: 9,890 tests pass; both golden-crucible modes pass (fixtures in the following commit); ruff baseline clean; mypy identical to main (5 pre-existing missing-stub errors). docs/vectors.md records all four, including why ~115 historical wiki teardowns keep their retired numbers rather than being edited. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwA8w2WpBqJg3AGf3yvYeF
…luster 2,426 differences per fixture, every one attributable to the preceding commit: 2,230 Specification Exposure per-file spec vector, no longer measured (#3111) 193 spec_match the same vector's ecosystem-summary aggregates 39 (removed paths) 4. High-Value Forensic Report/cumulative_risk (#3112) -- 10 highest + 3 lowest x 3 fields each Zero topological (X/Y/Z) churn, zero added paths, and nothing outside those three groups. Reviewed with bless_scope.py, which is the only way to review these now that #3141 marks them -diff -- and the reason that change was worth making: this is its first real fixture bump, and it renders as "Bin 57926723 -> 57919405 bytes" instead of ~400k lines of unreviewable churn. Each fixture was blessed from an environment matching the CI leg that checks it -- a PyYAML-only venv for the zero-dependency master, a PyYAML+tiktoken+pandas+xgboost venv for full-precision -- after a first attempt recorded "numpy": false into the zero-dep fixture. That would have been my machine leaking into a fixture whose whole purpose is to describe an environment without optional engines, and CI (which installs PyYAML only on that leg) would have disagreed on a leaf that has nothing to do with this change. Missing Dependencies is compared; only Absolute Project Path, the timestamps and the git footprint are sanitized. Both modes verified green locally: pytest -m golden_crucible passes against each fixture from its matching venv. One near-miss worth recording. A clean-corpus run initially disagreed with main on .gitgalaxy/dependency_cache.db's recorded Size (20480 -> 4096), which looked like a volatile scan artifact -- the scan writes that cache INSIDE the tree it is scanning, then records it in the excluded queue. It is not volatile: the file is tracked in language-crucible at v1.3.0 at exactly 20480 bytes, committed by accident in corpus commit 6a2257c. The 4096 came from me deleting a tracked corpus file before re-blessing. main was never broken, the corpus is restored, and the pre-emptive "fix" (redirecting --dependency-cache in both the bless and check paths) was reverted as unnecessary. Filed separately: a dependency-cache DB is checked into the corpus and is now load-bearing for fixture reproducibility, which is fragile even though it currently works. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwA8w2WpBqJg3AGf3yvYeF
Contributor
CodeQL (Note) on #3145: gitgalaxy.recorders.llm_recorder was imported both as `from ... import LLMRecorder` at module scope and as `import ... as module` inside test_the_two_duplicate_hitlists_are_gone. Fixed with inspect.getmodule(LLMRecorder) rather than CodeQL's suggested sys.modules[LLMRecorder.__module__], which would add a `sys` import to do what the already-imported `inspect` does natively. Verified the test is still not vacuous: the resolved module source is the real 93,724-char module, it contains the current "## 11. RANKED ARTIFACTS" heading, and it does not contain either removed heading -- so the assertion is still capable of failing if a duplicate hitlist ever came back. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HwA8w2WpBqJg3AGf3yvYeF
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3111. Closes #3112. Closes #3113. Closes #3114.
Four issues from one real-repo plausibility audit (IBM's zopeneditor-sample — COBOL/PL-I/JCL/HLASM/REXX, 51 files). None changes how anything is computed. All four change what the report claims.
The shared defect
The brief's Primary Risk Drivers line took the top four vectors by raw value. Two of the thirteen —
spec_matchanddocumentation— are "fraction not covered by a convention" meters, so on any codebase lacking that convention they sit at ceiling. Both read mode 100 / median 50 on the sample scan, and they took 2 of 4 slots on all ten top-10 entries.A line whose only job is to discriminate between files was reprinting two constants and burying the vectors that actually differed. Reproduced as a test — with both at 100 and four ordinary vectors below them, the line now reports only the four that differentiate:
#3111 — spec alignment opt-in, default OFF
--spec-alignment. When off it is absent from every display surface, not rendered as0.0— a 0.0 asserts perfect spec alignment, the opposite of "not measured". Section 6 says so explicitly and names the flag.The stored vector stays dense and numeric (the slot holds 0.0). Injecting
Nonewould have meant auditing ~25 positional consumers, several of which crash on it —gpu_recorderdoesint(v * 10),network_risk_sensormultiplies,dev_agent_firewallsums,record_keeperwrites REAL columns — and dense-0.0 is already how the engine represents the history vectors it ablates (galaxyscope.py:1174). Honesty is enforced at the render surfaces, which is where the issue asked for it.#3114 — documentation reframed, not disabled
Still computed, still tabled (marked
_(coverage)_), still reported per file as "X% of unit weight undocumented" — but no longer eligible for the driver narrative. Its own contract (#2908) defines it as a coverage gap over units, so coverage is where it belongs.Both are declared once, in
analysis_lens:with a single
inactive_vectors()resolver every renderer consults, so a second optional vector needs no renderer edits. A contract test asserts both name realRISK_SCHEMAentries and that the families stay disjoint — off and reframed are different asks.#3112 — the Cumulative Risk composite is removed
sum()over the whole 13-entryrisk_vector: independently scaled sigmoids added with no weighting and no unit, yieldingCumulative Risk: 556.29. Unrepairable in place — ~31% of the sum was constant or dead, and summing the vectors re-created exactly the composite risk score that the #2982/#2991 validation record retired.--max-systemic-threat. Wascumulative_risk × blast_radius, nowstructural_magnitude × blast_radius. Intent unchanged and arguably better served ("fail when something structurally heavy is also widely depended upon"), and both factors are now unit-honest. Thresholds do not carry over — the scales differ — and are deliberately not silently rescaled: a fabricated conversion factor fails more quietly than an obvious re-tune. Flag help anddocs/vectors.mdboth say so.#3113 — the brief leads with the answer
A new executive summary, before any formula exposition. It reuses #2556's zero-guard, so a flat import graph is reported as flat rather than ranked by scan order — and it independently reproduced the two facts that audit verified against source:
Sections 11 and 12 answered "which files deserve attention first" twice, and the one that ranked did so on the composite #3112 deleted. Now one list, ordered by magnitude, each entry carrying a consequence assembled from edges the brief already computed but never stated:
Outbound comes from
raw_imports(matching section 7 and the summary) because the resolvedout_degreecan read 0 directly above a line naming two imports — a regression test pins that.The 75-line lexicon and the non-predictive disclaimer move to Appendix A, verbatim. #3113 was explicit that the honesty is an asset and must be kept; it just shouldn't be the wall a reader climbs to reach the findings. A pointer is left where section 2 was. The AI output directive that told consumers to hunt for "massive Cumulative Risk" now points at magnitude and blast radius and explicitly forbids summing the vectors.
Section order:
0, 0.5, 1 (exec), 1.5, 3 … 10.7, 11 (ranked), 12, 12.5, 12.8, Appendix A. Brief went 883 → 845 lines on the sample repo, with the verified content first.Fixtures
Second commit, 2,426 differences per master, all attributable:
Specification Exposure— per-file vector no longer measured (#3111)spec_match— the same vector's ecosystem aggregates4. High-Value Forensic Report/cumulative_riskpaths (#3112)Zero topological churn, zero added paths, nothing outside those groups. This is the first real fixture bump since #3141 marked these
-diff— it renders asBin 57926723 -> 57919405 bytesrather than ~400k lines, and was reviewed withbless_scope.py, exactly as that PR argued.Each fixture was blessed from an environment matching the CI leg that checks it (PyYAML-only for zero-dep; +tiktoken/pandas/xgboost for full-precision), after a first attempt recorded
"numpy": falseinto the zero-dep fixture — my machine leaking into a fixture whose purpose is to describe an environment without optional engines.Missing Dependenciesis compared; onlyAbsolute Project Path, timestamps and the git footprint are sanitized. Both modes verified green locally.Verification
--ignore=tests/security_auditing, and deselecting the known-stale localfidelity_tablecheck)out_degreeregressiongolden_cruciblemodes pass from their matching venvsruff_audit.py --ci: no new findings beyond the 8-finding baselinemypy gitgalaxy/: identical to main — 5 pre-existing missing-stub errors, none introducedNotes
docs/vectors.mdrecords all four, including why ~115 historicaldocs/wiki/teardowns keep their retiredcumulative_risknumbers instead of being edited — they're dated analyses of what the engine really did emit, and rewriting them would falsify the record. (#3112 estimated ~8; it's 115.)docs/gitgalaxy_architecture_brief.mdis regenerated by the scheduled self-scan and drops the metric on its next run.One near-miss recorded in the fixture commit: a clean-corpus run disagreed with main on
.gitgalaxy/dependency_cache.db's recordedSize, which looked like a volatile artifact of the scan recording its own cache. It isn't — that file is tracked in language-crucible at v1.3.0 at exactly 20480 bytes, committed by accident in corpus commit6a2257c. The mismatch came from me deleting a tracked corpus file. main was never broken, the corpus is restored, and the pre-emptive fix was reverted. Worth its own issue: a dependency-cache DB checked into the corpus is now load-bearing for fixture reproducibility.🤖 Generated with Claude Code
https://claude.ai/code/session_01HwA8w2WpBqJg3AGf3yvYeF