Repository navigation
test(refraction): pinned mainframe corpora, generated-output snapshot, tool-regex ReDoS sweep (#3213, #3212, #3214) - #3223
Merged
Conversation
…, tool-regex ReDoS sweep #3213: tests/cobol_mainframe/corpora.json pins zopeneditor-sample, CBSA (July2024Refresh) and aws-mainframe-modernization-carddemo, which carries 21 real BMS maps. tests/tools/mainframe_corpus.py can fetch, scan, score, print a path, or excerpt each corpus. Clones and DBs live in the gitignored .mainframe_corpora/, and scans are cached per ref + engine state. #3212: tests/tools/refraction_snapshot.py runs cobol-refractor and then cobol-to-java on a copy of the input. It diffs every generated file against committed .snap files. CI checks the committed excerpts; the full corpora are checked locally. The snapshot found hash-order and filesystem-order nondeterminism, now fixed: - unsorted rglob/glob in both controllers and the JCL auditor - unresolved_calls built with list(set) in the DAG architect #3214: tests/tools/tool_regex_redos.py sweeps every regex in both tool suites. An offender is slow outright, or grows >= 8x for 4x the input (confirmed best-of-3). Three existing offenders are baselined (#3222). Reverting #3205's fix makes the sweep fail. Closes #3213, closes #3212, closes #3214. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Contributor
…_ESCAPE alternation (CodeQL) - WindowsPath sorts case-insensitively. multiroot__* and SAM1LIB* therefore swapped order on Windows, and a different same-title schema won RptStatsDetail.java (the #3221 collision), failing the zopeneditor excerpt snapshot on windows-latest 3.9 and 3.12. Sort by p.parts or p.name, which compare the same on every OS. Order on Linux is unchanged: 0 snapshot differences on excerpts and full corpora. - tool_regex_redos._ESCAPE: `\\.` and `[^\]]` both consumed a backslash. That made the alternation exponential on '[' + many backslashes (CodeQL py/redos). The class now excludes the backslash: 0.5 ms on 5k backslashes. The sweep results are unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This was referenced Sep 19, 2026
squid-protocol
added a commit
that referenced
this pull request
Sep 20, 2026
The engine stores OS-native paths, so `c.count("/")` was 0 for every
candidate on Windows and the shallower-path tiebreak silently degraded
to alphabetical -- Windows would resolve a shared PROGRAM-ID to a
DIFFERENT file than Linux for the same repository. `_shared()` already
normalised; the tiebreak did not.
The regression test uses `zz/P.cbl` vs `aa/bb/P.cbl`, the case where
depth and alphabetical disagree: it fails without the fix and passes
with it. The earlier fixture never exercised the tiebreak, because the
shared-prefix test already discriminated there.
Same class as #3223.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
squid-protocol
added a commit
that referenced
this pull request
Sep 20, 2026
…#3238) * feat(mainframe): named call graph and dataset boundary (#3200, #3201) The counted rules say THAT a COBOL program calls something and touches files (ipc_rpc_bridges -> arch_ipc, io -> arch_io); they never say WHAT. This adds the named channel for both relations, so the master DB answers "program P opens DD X for INPUT; job J step S binds DD X to dataset D" -- two stated absences in docs/refraction_engine_differential.md. Both issues land together because they are one mechanism and one consumer: each gates #3120's lineage switch, so either alone unblocks nothing, and the JCL half of each is the same statement pass. New: - core/mainframe_boundary.py -- extraction off the prism CODE STREAM: COBOL CALL, CICS LINK/XCTL PROGRAM(...), JCL EXEC PGM=; SELECT ... ASSIGN TO <dd> with the OPEN modes actually used; JCL DD -> DSN. - core/invocation_resolver.py -- repo-wide name -> file resolution and the call/exec edge aggregation. - call_site_data -- one row per invocation site, resolved or not. - dataset_data -- the dataset boundary, both the COBOL and the JCL half. - edge_data gains edge_kind 'call' and 'exec'. - galaxy_ir.py: dataset_lineage() and unresolved_calls(). A language opts in with a TOP-LEVEL `boundary_extraction` declaration: language_lens.py re.compile()s every string value inside `rules`, so a helper key put there arrives as a Pattern and extracts nothing while every unit test stays green (#2806). Scores (answer key): every engine column that read "not carried" is now exact -- zopeneditor DD names 12/12, inputs 6/6, outputs 6/6, dynamic CALLs 3/3; CBSA DD names 1/1, outputs 1/1; call targets 3/3 and 45/45. The extractor was scored before being wired in: 147/147 call sites match on verb, form, operand, target AND line; 13/13 dataset records match. JCL has no answer key, so EXEC PGM= was checked against an independent raw-file scan: 24/77/150 steps, exact. Call edges are a separate kind and never enter the DiGraph, so pagerank, popularity, internal_dependency_links, betweenness, archetypes and every risk score are unchanged. Whether they SHOULD count as coupling is #3237. #2992's per-file reconciliation is now scoped to edge_kind='import', and so is galaxy_ir's copy_deps. Verification: golden master PASS both modes, zero diff (no bless owed); refraction snapshots unchanged on all three corpora; ruff/mypy/dead-key clean; every new pattern scales linearly under the repo's ReDoS rule; full-then-incremental scans produce byte-identical boundary rows. Also fixes pre-existing test pollution: two tests in test_galaxyscope.py injected a MagicMock over sys.modules["gitgalaxy.core.state_rehydrator"] and never restored it, so every later test importing StateRehydrator got a mock returning an empty ram_cache. Restored with addCleanup. Out of scope, stated: FD/01 record layouts (data-division extraction, a separate absence) and reachability (the engine has none -- #3198). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(mainframe): normalise separators in the nearest-wins tiebreak The engine stores OS-native paths, so `c.count("/")` was 0 for every candidate on Windows and the shallower-path tiebreak silently degraded to alphabetical -- Windows would resolve a shared PROGRAM-ID to a DIFFERENT file than Linux for the same repository. `_shared()` already normalised; the tiebreak did not. The regression test uses `zz/P.cbl` vs `aa/bb/P.cbl`, the case where depth and alphabetical disagree: it fails without the fix and passes with it. The earlier fixture never exercised the tiebreak, because the shared-prefix test already discriminated there. Same class as #3223. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tests): short ids for the boundary ReDoS parametrization A parametrized literal becomes part of pytest's test id, and pytest exports that id as PYTEST_CURRENT_TEST. A 40,000-character id exceeds Windows' 32,767-character environment variable limit, so all 16 of these errored at SETUP -- before the test body ran -- on windows-latest 3.9 and 3.12 (28 errors), while every other platform passed. Parametrize over a short shape name and build the payload inside the test instead. Longest id is now 119 characters; the assertions are unchanged. Found by the dispatched OS x Python matrix, which the PR checks do not cover (it triggers on `labeled`, so it never saw the later commit). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * style(tests): drop a redundant sqlite3 import (CodeQL) `sqlite3` is already imported at module scope, so the function-level import shadowed it for no reason. Flagged by CodeQL on #3238. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Signed-off-by: Joe Esquibel <joe.m.esquibel@gmail.com> Co-authored-by: Joe Esquibel <squid-protocol@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3213, closes #3212 (its second half; the sorted IR sets landed in #3217), closes #3214. Part of #3122.
Tests and tools only. No golden-master bless. The engine is untouched.
#3213: pinned corpora as commands
tests/cobol_mainframe/corpora.jsonrecords name, url, branch, ref, license, languages, answer key and CI excerpt for each corpus.tests/tools/mainframe_corpus.pyhas these commands:fetch: a shallow clone at the pinned SHA. It refuses to discard local edits without--force.scan: builds the master DB throughgalaxy_ir.scan_to_db. The DB is cached by ref, engine HEAD and a hash of any dirtygitgalaxy/diff, so an engine edit never reuses a stale DB. The scan runs with this checkout first onPYTHONPATH, which avoids the stale-worktree trap.score: the answer-key scorer on the cached clone and DB.pathandexcerpt..mainframe_corpora/, or$GITGALAXY_MAINFRAME_CORPORA. Galaxyscope's aperture skips dot-dirs, the same precedent as.ast_accuracy_corpus/.fetchof all three corpora takes about 5s andscanabout 30s.scorethen reproduces the post-fix(refraction): graveyard reachability with a SECTION model (#3203) + per-path clean-room keys (#3218) #3219 numbers (CBSA dead P 62/62, R 62/62) in 1.8s from cache.Third corpus / BMS survey. aws-mainframe-modernization-carddemo @
59cc6c2f(Apache-2.0) has 21.bmsmaps, each with its generated symbolic-map copybook; CBSA has 10. On both, the engine's mapset→class_dataand map→function_datanames match the source labels on 31/31 files. Two limits remain for #3122's BMS row:#3212: snapshot of the generated clean room + Spring Boot tree
tests/tools/refraction_snapshot.py check|update [--corpus NAME ...]:cobol-refractorand thencobol-to-java, in-process, on a copy of the input.<WORK>with/separators, the timestamp becomes<TS>, and JSON is re-serialised canonically.<path>.snapfile per generated file. A forge PR's diff therefore shows exactly which JCL, schema, IR or Java lines moved.tests/cobol_mainframe/refraction_excerpts/(35 upstream files, unmodified, license files alongside, listed in the manifest, regenerated bymainframe_corpus.py excerpt)refraction_snapshot/excerpts/(87 files)test_refraction_snapshot.py, ~2srefraction_snapshot/full/(440 files, 1.27 MB)The excerpts are chosen to exercise what was just fixed:
COBOL/SAM2+multiroot/sam/SAM2;CALC-DAY-OF-WEEKandPOPULATE-TIME-DATEsections, and the BANKDATA OPEN OUTPUT;.gitattributeskeeps the excerpts byte-exact on Windows (-text) and the snapshots LF.What the snapshot found on its first run. The output was not deterministic. Two runs with different
PYTHONHASHSEEDdiffered on CardDemo's COCRDLIC:unresolved_callscame fromlist(set), which reordered the IR, the audit report, the agent job and the generated Java. Fixes:cobol_dag_architect.py:sorted(dynamic_calls).cobol_jcl_auditor.py: sort theirrglob/globwalks, because filesystem order differs per OS and orders the report, the jobs and same-name entity writes.list(set)[0]becomessorted(...)[0].These change ordering only. Output is now byte-identical across three hash seeds.
It also shows #3221: the five zopeneditor IR dumps yield three Java services.
cobol-to-javastill names classes frommetadata.file_name, so the two SAM2 programs collide. This is the Java side of #3218, filed rather than fixed here.#3214: ReDoS sweep over the tool regexes
tests/tools/tool_regex_redos.py [--ci | --update-baseline]:re.*(...)call incobol_to_cobol/,cobol_to_java/and both controllers. Patterns and flags are evaluated in the module namespace, so f-strings over module constants resolve. It also takes every compiled global. The total is 54 patterns. The 4 built from function locals are listed as skipped, never dropped..{0,600}?CICS HANDLE scan: 0.13s at 160k, but linear. There is no ratio between small single samples (flaky: ReDoS "pre-fix pattern scales quadratically" timing ratios fail on shared macOS runners #2901).tests/cobol_mainframe/tool_regex_redos_baseline.jsonholds three offenders, each with a note → Refraction tools: three regexes with quadratic scaling (baselined by the #3214 sweep) #3222:dag_architectOPEN[^.]*graveyard_finderREPLACING pairsjcl_forgeSELECT/ASSIGN[^.]*test_tool_regex_redos.pyin every pytest leg, ~2s, because it skips re-measuring the baselined offenders.cobol_jcl_forge.pyfrom before fix(refraction): forge parser fixes (#3203 1-4, #3204, #3205, #3206, #3212 determinism) #3217 reportsEXEC\s+CICS.*?END-EXEC\.and itsEXEC SQLtwin as 2 new offenders, so--cifails.tests/extraction/tools/sweep_redos_scaling.pyis not extended. It scrapesassert_redos_immunepayloads out of the engine's strict suites, and the tools have no such suites to scrape.Verification
pytest tests/ -n 8gives 10250 passed, 7 skipped, 9 xfailed, 3 xpassed.write_text(newline=)that is 3.10-only.audit_check.pyis all clear (ruff, mypy, dead-key, ast-accuracy).xray-inspector .reports 0 anomalies.Docs
tests/cobol_mainframe/readme.mdhas a new section 7.docs/refraction_engine_differential.mdnow give the commands in place of the manual clone steps.Unblocks #3215 (the cobol-modernization skill), which can now document commands.
🤖 Generated with Claude Code