Repository navigation
feat(mainframe): named call graph and dataset boundary (#3200, #3201) - #3238
Merged
Merged
Conversation
The counted rules say THAT a COBOL program calls something and touches files (ipc_rpc_bridges -> arch_ipc, io -> arch_io); they never say WHAT. This adds the named channel for both relations, so the master DB answers "program P opens DD X for INPUT; job J step S binds DD X to dataset D" -- two stated absences in docs/refraction_engine_differential.md. Both issues land together because they are one mechanism and one consumer: each gates #3120's lineage switch, so either alone unblocks nothing, and the JCL half of each is the same statement pass. New: - core/mainframe_boundary.py -- extraction off the prism CODE STREAM: COBOL CALL, CICS LINK/XCTL PROGRAM(...), JCL EXEC PGM=; SELECT ... ASSIGN TO <dd> with the OPEN modes actually used; JCL DD -> DSN. - core/invocation_resolver.py -- repo-wide name -> file resolution and the call/exec edge aggregation. - call_site_data -- one row per invocation site, resolved or not. - dataset_data -- the dataset boundary, both the COBOL and the JCL half. - edge_data gains edge_kind 'call' and 'exec'. - galaxy_ir.py: dataset_lineage() and unresolved_calls(). A language opts in with a TOP-LEVEL `boundary_extraction` declaration: language_lens.py re.compile()s every string value inside `rules`, so a helper key put there arrives as a Pattern and extracts nothing while every unit test stays green (#2806). Scores (answer key): every engine column that read "not carried" is now exact -- zopeneditor DD names 12/12, inputs 6/6, outputs 6/6, dynamic CALLs 3/3; CBSA DD names 1/1, outputs 1/1; call targets 3/3 and 45/45. The extractor was scored before being wired in: 147/147 call sites match on verb, form, operand, target AND line; 13/13 dataset records match. JCL has no answer key, so EXEC PGM= was checked against an independent raw-file scan: 24/77/150 steps, exact. Call edges are a separate kind and never enter the DiGraph, so pagerank, popularity, internal_dependency_links, betweenness, archetypes and every risk score are unchanged. Whether they SHOULD count as coupling is #3237. #2992's per-file reconciliation is now scoped to edge_kind='import', and so is galaxy_ir's copy_deps. Verification: golden master PASS both modes, zero diff (no bless owed); refraction snapshots unchanged on all three corpora; ruff/mypy/dead-key clean; every new pattern scales linearly under the repo's ReDoS rule; full-then-incremental scans produce byte-identical boundary rows. Also fixes pre-existing test pollution: two tests in test_galaxyscope.py injected a MagicMock over sys.modules["gitgalaxy.core.state_rehydrator"] and never restored it, so every later test importing StateRehydrator got a mock returning an empty ram_cache. Restored with addCleanup. Out of scope, stated: FD/01 record layouts (data-division extraction, a separate absence) and reachability (the engine has none -- #3198). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
|
||
| def test_a_baseline_written_before_3200_still_rehydrates(recorded): | ||
| """A DB with no boundary tables is a normal baseline, not a failure.""" | ||
| import sqlite3 |
Contributor
The engine stores OS-native paths, so `c.count("/")` was 0 for every
candidate on Windows and the shallower-path tiebreak silently degraded
to alphabetical -- Windows would resolve a shared PROGRAM-ID to a
DIFFERENT file than Linux for the same repository. `_shared()` already
normalised; the tiebreak did not.
The regression test uses `zz/P.cbl` vs `aa/bb/P.cbl`, the case where
depth and alphabetical disagree: it fails without the fix and passes
with it. The earlier fixture never exercised the tiebreak, because the
shared-prefix test already discriminated there.
Same class as #3223.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A parametrized literal becomes part of pytest's test id, and pytest exports that id as PYTEST_CURRENT_TEST. A 40,000-character id exceeds Windows' 32,767-character environment variable limit, so all 16 of these errored at SETUP -- before the test body ran -- on windows-latest 3.9 and 3.12 (28 errors), while every other platform passed. Parametrize over a short shape name and build the payload inside the test instead. Longest id is now 119 characters; the assertions are unchanged. Found by the dispatched OS x Python matrix, which the PR checks do not cover (it triggers on `labeled`, so it never saw the later commit). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`sqlite3` is already imported at module scope, so the function-level import shadowed it for no reason. Flagged by CodeQL on #3238. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Joe Esquibel <joe.m.esquibel@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3200. Closes #3201.
The counted rules say that a COBOL program calls something and touches files (
ipc_rpc_bridges→arch_ipc,io→arch_io). They never say what. This adds the named channel, so the master DB can answer "program P opens DD X for INPUT; job J step S binds DD X to dataset D" — whichdocs/refraction_engine_differential.mdrecorded as two stated absences.Both issues land together because they are one mechanism and one consumer: the skill's ownership table has both gating #3120's lineage switch, so either alone unblocks nothing, and the JCL half of each is the same statement pass (
//STEP EXEC PGM=and//DD DD DSN=).What's new
core/mainframe_boundary.pyCALL, CICSLINK/XCTL PROGRAM(...), JCLEXEC PGM=;SELECT ... ASSIGN TO <dd>with theOPENmodes actually used; JCLDD→ DSNcore/invocation_resolver.pycall/execedge aggregationcall_site_datadataset_dataedge_dataedge_kind'call'/'exec'galaxy_ir.pydataset_lineage()andunresolved_calls()for the refraction toolsA language opts in with a top-level
boundary_extractiondeclaration (cobol, jcl).Scores (answer key)
Every engine column that read
not carriedis now carried, and exact. The forge column is untouched.call targetsis a new scored field. It has no forge column: the DAG architect records only non-literalCALLoperands and never seesEXEC CICS LINK/XCTL— 140 of CBSA's 144 call sites.The extractor was scored directly against the key before being wired in: 147/147 call sites match on verb, form, operand, target and line, and 13/13 dataset records on internal name, DD and modes, with no false positives. JCL has no answer key, so
EXEC PGM=was checked against an independent raw-file scan: 24 / 77 / 150 steps on the three corpora, exact.Verification
45/192/259 files, 0 differ), which is what an engine-side PR requires.ruff formatback at the pre-existing 29.call_site_dataanddataset_data.The one mypy baseline line moved 386 → 388 (pure line shift from two new imports); the message is unchanged.
Three decisions worth reviewing
edge_kindis'call'/'exec'and these never enter the DiGraph, sopagerank_score,popularity,internal_dependency_links, betweenness, archetypes and every risk score are byte-for-byte unchanged. Whether a runtime invocation should count as coupling is a scoring question with its own before/after — filed as Decide whether call/exec edges should feed the dependency graph (pagerank, popularity, blast radius) #3237, deliberately not settled here. Consequence: Persist the dependency graph edge list (enables neighborhood/interaction-level risk hypotheses) #2992's per-file reconciliation is now scoped toWHERE edge_kind = 'import', and so isgalaxy_ir'scopy_deps.SAM2programs.targetkeeps the name regardless, so nothing is lost when the file attribution is debatable.CALL 'CEEGMT'is an LE service,EXEC PGM=IEFBR14a system utility. The table distinguishes "the name itself was unreadable" (target IS NULL) from "named but external" (dst_file_id IS NULL).Out of scope, stated
OPENin an unreachable paragraph is still extracted. The engine has no reachability model, and inventing one here would repeat the mistake cobol usage_status: entry paragraph flagged, case-sensitive matching, and NAME-EXIT counts as a reference to NAME #3198 corrected.Two traps, recorded because each nearly shipped
rulestrap. The declaration was first written as a string insiderules, whichlanguage_lens.pysilentlyre.compile()s. Every unit test stays green while a real scan extracts nothing. It is now top level, with a test asserting it never moves back.sys.modulesentry (fixed here). Two tests intest_galaxyscope.pyinjected aMagicMockovergitgalaxy.core.state_rehydratorand never restored it, so every later test in the session that importedStateRehydratorgot a mock whoseload_statereturns{"commit_hash": "old", "ram_cache": {}}. The new delta tests pass alone and in their own directory, then failed in the full suite with an empty cache and aKeyErrorpointing nowhere near the cause. Restored withaddCleanupat both sites — pre-existing pollution, but this PR is the first thing to trip it.DISPLAY 'GNP CALL FAILED :'read as a call toFAILED— 5 of 102 sites on carddemo, which has no answer key and so was audited separately. Neither scored corpus contains the shape. Guarded per line (quote state never crosses a line, orAUTHOR. James O'Grady.swallows the file — the COBOL modernization answer key: hand-verified per-program truth for the #3120 corpora, scored against forge and engine #3210 trap), and pinned.Filed along the way
function_data.calls_out_toislist(set)[:20], so it keeps a different subset perPYTHONHASHSEED. Pre-existing and untouched here (detector.pyhas no diff on this branch), but it was the only non-additive difference between a pre- and post-change scan and had to be chased down to prove this PR moved nothing.🤖 Generated with Claude Code