Skip to content

feat(mainframe): named call graph and dataset boundary (#3200, #3201) - #3238

Merged
squid-protocol merged 5 commits into
mainfrom
feat/3200-3201-mainframe-edges
Sep 20, 2026
Merged

squid-protocol merged 5 commits into
mainfrom
feat/3200-3201-mainframe-edges

Conversation

@squid-protocol

Copy link
Copy Markdown
Owner

Closes #3200. Closes #3201.

The counted rules say that a COBOL program calls something and touches files (ipc_rpc_bridges → arch_ipc, io → arch_io). They never say what. This adds the named channel, so the master DB can answer "program P opens DD X for INPUT; job J step S binds DD X to dataset D" — which docs/refraction_engine_differential.md recorded as two stated absences.

Both issues land together because they are one mechanism and one consumer: the skill's ownership table has both gating #3120's lineage switch, so either alone unblocks nothing, and the JCL half of each is the same statement pass (//STEP EXEC PGM= and //DD DD DSN=).

What's new

core/mainframe_boundary.py extraction: COBOL CALL, CICS LINK/XCTL PROGRAM(...), JCL EXEC PGM=; SELECT ... ASSIGN TO <dd> with the OPEN modes actually used; JCL DD → DSN
core/invocation_resolver.py repo-wide name → file resolution, and the call/exec edge aggregation
call_site_data one row per invocation site — resolved or not
dataset_data the dataset boundary, both the COBOL and the JCL half
edge_data gains edge_kind 'call' / 'exec'
galaxy_ir.py dataset_lineage() and unresolved_calls() for the refraction tools

A language opts in with a top-level boundary_extraction declaration (cobol, jcl).

Scores (answer key)

Every engine column that read not carried is now carried, and exact. The forge column is untouched.

corpus field before after
zopeneditor-sample DD names not carried P 12/12 · R 12/12
zopeneditor-sample inputs not carried P 6/6 · R 6/6
zopeneditor-sample outputs not carried P 6/6 · R 6/6
zopeneditor-sample dynamic CALLs not carried P 3/3 · R 3/3
zopeneditor-sample call targets not carried P 3/3 · R 3/3
cics-banking-…-cbsa DD names not carried P 1/1 · R 1/1
cics-banking-…-cbsa outputs not carried P 1/1 · R 1/1
cics-banking-…-cbsa call targets not carried P 45/45 · R 45/45

call targets is a new scored field. It has no forge column: the DAG architect records only non-literal CALL operands and never sees EXEC CICS LINK/XCTL — 140 of CBSA's 144 call sites.

The extractor was scored directly against the key before being wired in: 147/147 call sites match on verb, form, operand, target and line, and 13/13 dataset records on internal name, DD and modes, with no false positives. JCL has no answer key, so EXEC PGM= was checked against an independent raw-file scan: 24 / 77 / 150 steps on the three corpora, exact.

Verification

  • Golden master: PASS in both modes, zero diff. No bless owed — no counted signal changed and none of this reaches the audit JSON.
  • Refraction snapshot: unchanged, all three corpora, excerpts and full (45/192/259 files, 0 differ), which is what an engine-side PR requires.
  • ruff / mypy / dead-key audits clean; ruff format back at the pre-existing 29.
  • ReDoS: every new pattern scales linearly (~4× for 4× input, 40k → 160k), no offender under the repo's own rule.
  • Delta parity: a full then an incremental scan of zopeneditor produce byte-identical call_site_data and dataset_data.

The one mypy baseline line moved 386 → 388 (pure line shift from two new imports); the message is unchanged.

Three decisions worth reviewing

  1. A call edge is not a dependency edge. edge_kind is 'call'/'exec' and these never enter the DiGraph, so pagerank_score, popularity, internal_dependency_links, betweenness, archetypes and every risk score are byte-for-byte unchanged. Whether a runtime invocation should count as coupling is a scoring question with its own before/after — filed as Decide whether call/exec edges should feed the dependency graph (pagerank, popularity, blast radius) #3237, deliberately not settled here. Consequence: Persist the dependency graph edge list (enables neighborhood/interaction-level risk hypotheses) #2992's per-file reconciliation is now scoped to WHERE edge_kind = 'import', and so is galaxy_ir's copy_deps.
  2. A CALL resolves by PROGRAM-ID, nearest-wins. That is not the import resolver's rule, which refuses to guess on an ambiguous stem (Dependency resolver drops ambiguous targets: COBOL COPY loses most copybook edges (same stem as the program, or duplicated copybook roots) #3199). An import names a file; a called program is chosen by library concatenation order at link-edit time, which is what the answer key records for zopeneditor's two SAM2 programs. target keeps the name regardless, so nothing is lost when the file attribution is debatable.
  3. Unresolved is data, not a gap. Most real sites resolve to nothing — CALL 'CEEGMT' is an LE service, EXEC PGM=IEFBR14 a system utility. The table distinguishes "the name itself was unreadable" (target IS NULL) from "named but external" (dst_file_id IS NULL).

Out of scope, stated

Two traps, recorded because each nearly shipped

  • The orphaned_logic reports 100% dead code for five languages whose invocation model never names the callee — the family #2727 closed without covering #2806 string-in-rules trap. The declaration was first written as a string inside rules, which language_lens.py silently re.compile()s. Every unit test stays green while a real scan extracts nothing. It is now top level, with a test asserting it never moves back.
  • A leaked sys.modules entry (fixed here). Two tests in test_galaxyscope.py injected a MagicMock over gitgalaxy.core.state_rehydrator and never restored it, so every later test in the session that imported StateRehydrator got a mock whose load_state returns {"commit_hash": "old", "ram_cache": {}}. The new delta tests pass alone and in their own directory, then failed in the full suite with an empty cache and a KeyError pointing nowhere near the cause. Restored with addCleanup at both sites — pre-existing pollution, but this PR is the first thing to trip it.
  • A verb inside a string literal. DISPLAY 'GNP CALL FAILED :' read as a call to FAILED — 5 of 102 sites on carddemo, which has no answer key and so was audited separately. Neither scored corpus contains the shape. Guarded per line (quote state never crosses a line, or AUTHOR. James O'Grady. swallows the file — the COBOL modernization answer key: hand-verified per-program truth for the #3120 corpora, scored against forge and engine #3210 trap), and pinned.

Filed along the way

🤖 Generated with Claude Code

The counted rules say THAT a COBOL program calls something and touches
files (ipc_rpc_bridges -> arch_ipc, io -> arch_io); they never say WHAT.
This adds the named channel for both relations, so the master DB answers
"program P opens DD X for INPUT; job J step S binds DD X to dataset D" --
two stated absences in docs/refraction_engine_differential.md.

Both issues land together because they are one mechanism and one
consumer: each gates #3120's lineage switch, so either alone unblocks
nothing, and the JCL half of each is the same statement pass.

New:
- core/mainframe_boundary.py -- extraction off the prism CODE STREAM:
  COBOL CALL, CICS LINK/XCTL PROGRAM(...), JCL EXEC PGM=; SELECT ...
  ASSIGN TO <dd> with the OPEN modes actually used; JCL DD -> DSN.
- core/invocation_resolver.py -- repo-wide name -> file resolution and
  the call/exec edge aggregation.
- call_site_data -- one row per invocation site, resolved or not.
- dataset_data -- the dataset boundary, both the COBOL and the JCL half.
- edge_data gains edge_kind 'call' and 'exec'.
- galaxy_ir.py: dataset_lineage() and unresolved_calls().

A language opts in with a TOP-LEVEL `boundary_extraction` declaration:
language_lens.py re.compile()s every string value inside `rules`, so a
helper key put there arrives as a Pattern and extracts nothing while
every unit test stays green (#2806).

Scores (answer key): every engine column that read "not carried" is now
exact -- zopeneditor DD names 12/12, inputs 6/6, outputs 6/6, dynamic
CALLs 3/3; CBSA DD names 1/1, outputs 1/1; call targets 3/3 and 45/45.
The extractor was scored before being wired in: 147/147 call sites match
on verb, form, operand, target AND line; 13/13 dataset records match.
JCL has no answer key, so EXEC PGM= was checked against an independent
raw-file scan: 24/77/150 steps, exact.

Call edges are a separate kind and never enter the DiGraph, so pagerank,
popularity, internal_dependency_links, betweenness, archetypes and every
risk score are unchanged. Whether they SHOULD count as coupling is #3237.
#2992's per-file reconciliation is now scoped to edge_kind='import', and
so is galaxy_ir's copy_deps.

Verification: golden master PASS both modes, zero diff (no bless owed);
refraction snapshots unchanged on all three corpora; ruff/mypy/dead-key
clean; every new pattern scales linearly under the repo's ReDoS rule;
full-then-incremental scans produce byte-identical boundary rows.

Also fixes pre-existing test pollution: two tests in test_galaxyscope.py
injected a MagicMock over sys.modules["gitgalaxy.core.state_rehydrator"]
and never restored it, so every later test importing StateRehydrator got
a mock returning an empty ram_cache. Restored with addCleanup.

Out of scope, stated: FD/01 record layouts (data-division extraction, a
separate absence) and reachability (the engine has none -- #3198).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@squid-protocol squid-protocol added enhancement New feature, sensor, or structural signature core-engine Modifications to the central physics and parsing engine legacy-modernization COBOL refractor, dead-code extraction, and JCL forging labels Sep 20, 2026

def test_a_baseline_written_before_3200_still_rehydrates(recorded):
"""A DB with no boundary tables is a normal baseline, not a failure."""
import sqlite3
@github-actions

Copy link
Copy Markdown
Contributor

🐦‍⬛ Muninn Security Scan

✅ No security issues found.

🐦‍⬛ Powered by Muninn · Skald Lab

squid-protocol and others added 4 commits September 20, 2026 09:23
The engine stores OS-native paths, so `c.count("/")` was 0 for every
candidate on Windows and the shallower-path tiebreak silently degraded
to alphabetical -- Windows would resolve a shared PROGRAM-ID to a
DIFFERENT file than Linux for the same repository. `_shared()` already
normalised; the tiebreak did not.

The regression test uses `zz/P.cbl` vs `aa/bb/P.cbl`, the case where
depth and alphabetical disagree: it fails without the fix and passes
with it. The earlier fixture never exercised the tiebreak, because the
shared-prefix test already discriminated there.

Same class as #3223.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A parametrized literal becomes part of pytest's test id, and pytest
exports that id as PYTEST_CURRENT_TEST. A 40,000-character id exceeds
Windows' 32,767-character environment variable limit, so all 16 of these
errored at SETUP -- before the test body ran -- on windows-latest 3.9
and 3.12 (28 errors), while every other platform passed.

Parametrize over a short shape name and build the payload inside the
test instead. Longest id is now 119 characters; the assertions are
unchanged.

Found by the dispatched OS x Python matrix, which the PR checks do not
cover (it triggers on `labeled`, so it never saw the later commit).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`sqlite3` is already imported at module scope, so the function-level
import shadowed it for no reason. Flagged by CodeQL on #3238.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Joe Esquibel <joe.m.esquibel@gmail.com>
@squid-protocol
squid-protocol merged commit fb6b409 into main Sep 20, 2026
31 checks passed
@squid-protocol
squid-protocol deleted the feat/3200-3201-mainframe-edges branch September 20, 2026 13:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core-engine Modifications to the central physics and parsing engine enhancement New feature, sensor, or structural signature legacy-modernization COBOL refractor, dead-code extraction, and JCL forging

Projects

None yet

2 participants