Skip to content

test(refraction): pinned mainframe corpora, generated-output snapshot, tool-regex ReDoS sweep (#3213, #3212, #3214) - #3223

Merged
squid-protocol merged 2 commits into
mainfrom
feat/3213-mainframe-corpus-tooling
Sep 19, 2026
Merged

squid-protocol merged 2 commits into
mainfrom
feat/3213-mainframe-corpus-tooling

Conversation

@squid-protocol

Copy link
Copy Markdown
Owner

Closes #3213, closes #3212 (its second half; the sorted IR sets landed in #3217), closes #3214. Part of #3122.

Tests and tools only. No golden-master bless. The engine is untouched.

#3213: pinned corpora as commands

  • Manifest. tests/cobol_mainframe/corpora.json records name, url, branch, ref, license, languages, answer key and CI excerpt for each corpus.
  • Tool. tests/tools/mainframe_corpus.py has these commands:
    • fetch: a shallow clone at the pinned SHA. It refuses to discard local edits without --force.
    • scan: builds the master DB through galaxy_ir.scan_to_db. The DB is cached by ref, engine HEAD and a hash of any dirty gitgalaxy/ diff, so an engine edit never reuses a stale DB. The scan runs with this checkout first on PYTHONPATH, which avoids the stale-worktree trap.
    • score: the answer-key scorer on the cached clone and DB.
    • path and excerpt.
  • Cache. The gitignored .mainframe_corpora/, or $GITGALAXY_MAINFRAME_CORPORA. Galaxyscope's aperture skips dot-dirs, the same precedent as .ast_accuracy_corpus/.
  • Timings. A cold fetch of all three corpora takes about 5s and scan about 30s. score then reproduces the post-fix(refraction): graveyard reachability with a SECTION model (#3203) + per-path clean-room keys (#3218) #3219 numbers (CBSA dead P 62/62, R 62/62) in 1.8s from cache.

Third corpus / BMS survey. aws-mainframe-modernization-carddemo @ 59cc6c2f (Apache-2.0) has 21 .bms maps, each with its generated symbolic-map copybook; CBSA has 10. On both, the engine's mapset→class_data and map→function_data names match the source labels on 31/31 files. Two limits remain for #3122's BMS row:

  • Neither corpus has a mapset with more than one DFHMDI.
  • The 733 named DFHMDF fields are not named entities in the DB.

#3212: snapshot of the generated clean room + Spring Boot tree

tests/tools/refraction_snapshot.py check|update [--corpus NAME ...]:

  • What it runs. cobol-refractor and then cobol-to-java, in-process, on a copy of the input.
  • Normalisation. The scratch dir becomes <WORK> with / separators, the timestamp becomes <TS>, and JSON is re-serialised canonically.
  • Storage. One <path>.snap file per generated file. A forge PR's diff therefore shows exactly which JCL, schema, IR or Java lines moved.
input snapshot where it runs
committed excerpts tests/cobol_mainframe/refraction_excerpts/ (35 upstream files, unmodified, license files alongside, listed in the manifest, regenerated by mainframe_corpus.py excerpt) refraction_snapshot/excerpts/ (87 files) CI, via test_refraction_snapshot.py, ~2s
full pinned corpora refraction_snapshot/full/ (440 files, 1.27 MB) locally, once fetched; skipped in CI

The excerpts are chosen to exercise what was just fixed:

.gitattributes keeps the excerpts byte-exact on Windows (-text) and the snapshots LF.

What the snapshot found on its first run. The output was not deterministic. Two runs with different PYTHONHASHSEED differed on CardDemo's COCRDLIC: unresolved_calls came from list(set), which reordered the IR, the audit report, the agent job and the generated Java. Fixes:

  • cobol_dag_architect.py: sorted(dynamic_calls).
  • Both controllers and cobol_jcl_auditor.py: sort their rglob/glob walks, because filesystem order differs per OS and orders the report, the jobs and same-name entity writes.
  • The auditor's list(set)[0] becomes sorted(...)[0].

These change ordering only. Output is now byte-identical across three hash seeds.

It also shows #3221: the five zopeneditor IR dumps yield three Java services. cobol-to-java still names classes from metadata.file_name, so the two SAM2 programs collide. This is the Java side of #3218, filed rather than fixed here.

#3214: ReDoS sweep over the tool regexes

tests/tools/tool_regex_redos.py [--ci | --update-baseline]:

tests/extraction/tools/sweep_redos_scaling.py is not extended. It scrapes assert_redos_immune payloads out of the engine's strict suites, and the tools have no such suites to scrape.

Verification

  • Full suite: pytest tests/ -n 8 gives 10250 passed, 7 skipped, 9 xfailed, 3 xpassed.
  • New tests: 36 pass locally with the corpora. Under Python 3.9 with no corpora (the CI situation), 30 pass and 6 skip; this caught a write_text(newline=) that is 3.10-only.
  • Lint: audit_check.py is all clear (ruff, mypy, dead-key, ast-accuracy). xray-inspector . reports 0 anomalies.

Docs

  • tests/cobol_mainframe/readme.md has a new section 7.
  • The answer-key README and docs/refraction_engine_differential.md now give the commands in place of the manual clone steps.

Unblocks #3215 (the cobol-modernization skill), which can now document commands.

🤖 Generated with Claude Code

…, tool-regex ReDoS sweep

#3213: tests/cobol_mainframe/corpora.json pins zopeneditor-sample, CBSA
(July2024Refresh) and aws-mainframe-modernization-carddemo, which carries 21
real BMS maps. tests/tools/mainframe_corpus.py can fetch, scan, score, print
a path, or excerpt each corpus. Clones and DBs live in the gitignored
.mainframe_corpora/, and scans are cached per ref + engine state.

#3212: tests/tools/refraction_snapshot.py runs cobol-refractor and then
cobol-to-java on a copy of the input. It diffs every generated file against
committed .snap files. CI checks the committed excerpts; the full corpora
are checked locally. The snapshot found hash-order and filesystem-order
nondeterminism, now fixed:
- unsorted rglob/glob in both controllers and the JCL auditor
- unresolved_calls built with list(set) in the DAG architect

#3214: tests/tools/tool_regex_redos.py sweeps every regex in both tool
suites. An offender is slow outright, or grows >= 8x for 4x the input
(confirmed best-of-3). Three existing offenders are baselined (#3222).
Reverting #3205's fix makes the sweep fail.

Closes #3213, closes #3212, closes #3214.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@squid-protocol squid-protocol added the legacy-modernization COBOL refractor, dead-code extraction, and JCL forging label Sep 19, 2026
Comment thread tests/tools/tool_regex_redos.py Fixed
@github-actions

Copy link
Copy Markdown
Contributor

🐦‍⬛ Muninn Security Scan

✅ No security issues found.

🐦‍⬛ Powered by Muninn · Skald Lab

…_ESCAPE alternation (CodeQL)

- WindowsPath sorts case-insensitively. multiroot__* and SAM1LIB* therefore
  swapped order on Windows, and a different same-title schema won
  RptStatsDetail.java (the #3221 collision), failing the zopeneditor
  excerpt snapshot on windows-latest 3.9 and 3.12. Sort by p.parts or
  p.name, which compare the same on every OS. Order on Linux is
  unchanged: 0 snapshot differences on excerpts and full corpora.
- tool_regex_redos._ESCAPE: `\\.` and `[^\]]` both consumed a backslash.
  That made the alternation exponential on '[' + many backslashes
  (CodeQL py/redos). The class now excludes the backslash: 0.5 ms on 5k
  backslashes. The sweep results are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@squid-protocol squid-protocol added legacy-modernization COBOL refractor, dead-code extraction, and JCL forging and removed legacy-modernization COBOL refractor, dead-code extraction, and JCL forging labels Sep 19, 2026
@squid-protocol
squid-protocol merged commit 89a740c into main Sep 19, 2026
38 of 40 checks passed
@squid-protocol
squid-protocol deleted the feat/3213-mainframe-corpus-tooling branch September 19, 2026 20:20
squid-protocol added a commit that referenced this pull request Sep 20, 2026
The engine stores OS-native paths, so `c.count("/")` was 0 for every
candidate on Windows and the shallower-path tiebreak silently degraded
to alphabetical -- Windows would resolve a shared PROGRAM-ID to a
DIFFERENT file than Linux for the same repository. `_shared()` already
normalised; the tiebreak did not.

The regression test uses `zz/P.cbl` vs `aa/bb/P.cbl`, the case where
depth and alphabetical disagree: it fails without the fix and passes
with it. The earlier fixture never exercised the tiebreak, because the
shared-prefix test already discriminated there.

Same class as #3223.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
squid-protocol added a commit that referenced this pull request Sep 20, 2026
…#3238)

* feat(mainframe): named call graph and dataset boundary (#3200, #3201)

The counted rules say THAT a COBOL program calls something and touches
files (ipc_rpc_bridges -> arch_ipc, io -> arch_io); they never say WHAT.
This adds the named channel for both relations, so the master DB answers
"program P opens DD X for INPUT; job J step S binds DD X to dataset D" --
two stated absences in docs/refraction_engine_differential.md.

Both issues land together because they are one mechanism and one
consumer: each gates #3120's lineage switch, so either alone unblocks
nothing, and the JCL half of each is the same statement pass.

New:
- core/mainframe_boundary.py -- extraction off the prism CODE STREAM:
  COBOL CALL, CICS LINK/XCTL PROGRAM(...), JCL EXEC PGM=; SELECT ...
  ASSIGN TO <dd> with the OPEN modes actually used; JCL DD -> DSN.
- core/invocation_resolver.py -- repo-wide name -> file resolution and
  the call/exec edge aggregation.
- call_site_data -- one row per invocation site, resolved or not.
- dataset_data -- the dataset boundary, both the COBOL and the JCL half.
- edge_data gains edge_kind 'call' and 'exec'.
- galaxy_ir.py: dataset_lineage() and unresolved_calls().

A language opts in with a TOP-LEVEL `boundary_extraction` declaration:
language_lens.py re.compile()s every string value inside `rules`, so a
helper key put there arrives as a Pattern and extracts nothing while
every unit test stays green (#2806).

Scores (answer key): every engine column that read "not carried" is now
exact -- zopeneditor DD names 12/12, inputs 6/6, outputs 6/6, dynamic
CALLs 3/3; CBSA DD names 1/1, outputs 1/1; call targets 3/3 and 45/45.
The extractor was scored before being wired in: 147/147 call sites match
on verb, form, operand, target AND line; 13/13 dataset records match.
JCL has no answer key, so EXEC PGM= was checked against an independent
raw-file scan: 24/77/150 steps, exact.

Call edges are a separate kind and never enter the DiGraph, so pagerank,
popularity, internal_dependency_links, betweenness, archetypes and every
risk score are unchanged. Whether they SHOULD count as coupling is #3237.
#2992's per-file reconciliation is now scoped to edge_kind='import', and
so is galaxy_ir's copy_deps.

Verification: golden master PASS both modes, zero diff (no bless owed);
refraction snapshots unchanged on all three corpora; ruff/mypy/dead-key
clean; every new pattern scales linearly under the repo's ReDoS rule;
full-then-incremental scans produce byte-identical boundary rows.

Also fixes pre-existing test pollution: two tests in test_galaxyscope.py
injected a MagicMock over sys.modules["gitgalaxy.core.state_rehydrator"]
and never restored it, so every later test importing StateRehydrator got
a mock returning an empty ram_cache. Restored with addCleanup.

Out of scope, stated: FD/01 record layouts (data-division extraction, a
separate absence) and reachability (the engine has none -- #3198).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(mainframe): normalise separators in the nearest-wins tiebreak

The engine stores OS-native paths, so `c.count("/")` was 0 for every
candidate on Windows and the shallower-path tiebreak silently degraded
to alphabetical -- Windows would resolve a shared PROGRAM-ID to a
DIFFERENT file than Linux for the same repository. `_shared()` already
normalised; the tiebreak did not.

The regression test uses `zz/P.cbl` vs `aa/bb/P.cbl`, the case where
depth and alphabetical disagree: it fails without the fix and passes
with it. The earlier fixture never exercised the tiebreak, because the
shared-prefix test already discriminated there.

Same class as #3223.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(tests): short ids for the boundary ReDoS parametrization

A parametrized literal becomes part of pytest's test id, and pytest
exports that id as PYTEST_CURRENT_TEST. A 40,000-character id exceeds
Windows' 32,767-character environment variable limit, so all 16 of these
errored at SETUP -- before the test body ran -- on windows-latest 3.9
and 3.12 (28 errors), while every other platform passed.

Parametrize over a short shape name and build the payload inside the
test instead. Longest id is now 119 characters; the assertions are
unchanged.

Found by the dispatched OS x Python matrix, which the PR checks do not
cover (it triggers on `labeled`, so it never saw the later commit).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* style(tests): drop a redundant sqlite3 import (CodeQL)

`sqlite3` is already imported at module scope, so the function-level
import shadowed it for no reason. Flagged by CodeQL on #3238.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Signed-off-by: Joe Esquibel <joe.m.esquibel@gmail.com>
Co-authored-by: Joe Esquibel <squid-protocol@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

legacy-modernization COBOL refractor, dead-code extraction, and JCL forging

Projects

None yet

2 participants