Skip to content

Collation hazards: extract alphanumeric relation conditions, guard the port, and prove it with boundary inputs #4035

Description

@squid-protocol

Follow-up to #3986 (whose PR adds util/CobolCompare and a porting rule in every ticket).

The gap #3986 leaves

The #3986 fix is only a general rule: the porting agent still has to work out for itself which comparisons are alphanumeric, and nothing checks what it returns. The equivalence harness only catches a String.compareTo port if its inputs happen to straddle the EBCDIC/ASCII disagreement (a digit vs a letter, lower vs upper case, trailing spaces), and generated inputs rarely do. So a wrong port can pass the proof.

Proposed: three layers, shipped together (each alone leaves a hole)

  1. Fact: extract the relation conditions (gitgalaxy/core/, bounded regex over the procedure-division tokens data_moves.py already produces, shaped like rounding_facts). Emit one row per relation condition: the line, the context (IF / EVALUATE WHEN / PERFORM UNTIL / SEARCH WHEN), the two operands, the operator, and each operand's class (PIC X / 9 / group / literal / figurative), resolved from the record layouts galaxy_ir already holds. The port ticket lists each alphanumeric comparison with the exact CobolCompare call, the same way it lists rounding statements. Cases that need care:
    • an abbreviated condition (IF A > 'X' AND < 'Y') must carry its subject forward;
    • a level-88 VALUE 'A' THRU 'Z' range depends on collation too;
    • class conditions (IS ALPHABETIC) and sign conditions aren't comparisons, so they should be recognised and skipped;
    • a PROGRAM COLLATING SEQUENCE / ALPHABET clause should be detected per program, instead of relying on the rule's TODO.
  2. Guardrail: check the returned Java. Add a collation violation to cobol_to_java_guardrail.py: .compareTo( / .equals( where either side is a String field known from the guardrail baseline (entity or DTO fields whose type is String).
  3. Proof: boundary inputs. For every field in a fact-table comparison, have tests/tools/equivalence_inputs.py generate values on the disagreement boundary: digit vs letter, lower vs upper case, trailing spaces, HIGH-/LOW-VALUES. Then a compareTo port fails the field-by-field comparison instead of passing by luck.

Later (separate issue): deterministic level-88 predicates

Level-88s are purely declarative, so the generator can write them itself: 88 IS-LETTER VALUE 'A' THRU 'Z' becomes a generated isIsLetter() on the entity or DTO, built on CobolCompare. No full procedure translation is proposed: that would mean a compiler front end, which works against the AST-free design.

Generalising

The same three layers (fact, guardrail, boundary inputs) fit the other Globe hazards too: DECIMAL-POINT IS COMMA (#3984), DBCS SO/SI offsets (#3985) and BiDi screens (#3987). It's worth building the mechanism once, with collation as the first user.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    data-integrityIssues relating to RAM state or SQLite database fidelityenhancementNew feature, sensor, or structural signaturelegacy-modernizationCOBOL refractor, dead-code extraction, and JCL forgingpriority: medium

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions