Skip to content

Fix #3955: Expand COBOL identifiers to support international characters (Kanji, Katakana) - #3967

Merged
squid-protocol merged 2 commits into
mainfrom
fix/3955-cobol-japanese-words
Sep 28, 2026
Merged

squid-protocol merged 2 commits into
mainfrom
fix/3955-cobol-japanese-words

Conversation

@squid-protocol

Copy link
Copy Markdown
Owner

Resolves #3955

Defect:
Japanese user-defined words (like PROGRAM-ID. 名前. and 01 データ項目) were completely ignored or truncated by GitGalaxy's boundary extractors. The root cause was that NATIONAL inside identifiers.py was hardcoded to only include EBCDIC special characters (e.g. ÆØÅÄÖÜѧ£) and did not include Unicode ranges like Kanji, Katakana, or Hiragana, which are fully supported as letters by modern IBM Enterprise COBOL and open-source Japanese COBOL compilers.

Fix:
I appended ID_START (which encompasses all valid Unicode letters Lo, Lu, Ll, etc.) directly to the NATIONAL definition inside gitgalaxy/standards/language_standards/identifiers.py.
Since [A-Z" + NATIONAL + r"0-9-] is uniformly used as the core identifier matching block across mainframe_boundary.py and job_flow.py, this single change natively grants full internationalization support to all COBOL boundaries (PROGRAM-ID, Section headers, Data names) while preserving the strictly-required "must contain at least one letter" identifier rule (because ID_START contains zero digits).

@github-actions

Copy link
Copy Markdown
Contributor

🐦‍⬛ Muninn Security Scan

✅ No security issues found.

🐦‍⬛ Powered by Muninn · Skald Lab

…ents

By expanding `NATIONAL` to include all valid Unicode letters (`ID_START`), this implicitly added support for underscores (`_`) to COBOL user-defined words. IBM Enterprise COBOL natively supports underscores, and the golden master diffs confirm that variables like `COMPANY_PREFIX` and `WS_HOST` are now accurately extracted in full, rather than being truncated to `COMPANY` and `WS` at the underscore delimiter. This similarly improved label extraction in HLASM, causing previously unextracted variables to be tracked (slightly raising the dead code metric correctly).
@squid-protocol
squid-protocol merged commit 4e9e288 into main Sep 28, 2026
37 checks passed
@squid-protocol
squid-protocol deleted the fix/3955-cobol-japanese-words branch September 28, 2026 13:52
squid-protocol added a commit that referenced this pull request Sep 29, 2026
… − U+2212) (#3995)

#3967 let a COBOL word hold any Unicode letter, but a Japanese word also holds
full-width digits (0-9) and the full-width hyphen, which SJIS 0x817C decodes to
U+FF0D under cp932 and U+2212 under Python's shift_jis. Data names were cut at
those characters (項目2 -> 項目, TEST−DATA1 -> TEST) and sections holding
them (S−初期化, s123) were missing from function_data.

- identifiers.py: WIDE_DIGITS / WIDE_HYPHENS.
- cobol.py, mainframe_boundary.py (COBOL readers), file_control.py (COBOL FD /
  SELECT / COPY): every COBOL word class that takes 0-9 takes WIDE_DIGITS, and
  every one that takes the in-word '-' takes WIDE_HYPHENS, at the same positions.
  JCL / IDCAMS / CSD / HLASM / BMS classes are unchanged. Quantifiers unchanged.
- detector._extract_name: for COBOL the full-width hyphens are name characters,
  so `S−初期化` is not reduced to its tail `初期化`.
- tests: data names, sections, callees, both decodings of 0x817C from a
  cp932-encoded file, spaced minus stays arithmetic, and a structural check that
  every cobol.py word class carries the wide characters.


Claude-Session: https://claude.ai/code/session_01EC8FWfn3aPmsUPnRgKNVup

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant