Repository navigation
Fix #3955: Expand COBOL identifiers to support international characters (Kanji, Katakana) - #3967
Merged
Merged
Conversation
Contributor
…ents By expanding `NATIONAL` to include all valid Unicode letters (`ID_START`), this implicitly added support for underscores (`_`) to COBOL user-defined words. IBM Enterprise COBOL natively supports underscores, and the golden master diffs confirm that variables like `COMPANY_PREFIX` and `WS_HOST` are now accurately extracted in full, rather than being truncated to `COMPANY` and `WS` at the underscore delimiter. This similarly improved label extraction in HLASM, causing previously unextracted variables to be tracked (slightly raising the dead code metric correctly).
This was referenced Sep 28, 2026
squid-protocol
added a commit
that referenced
this pull request
Sep 29, 2026
… − U+2212) (#3995) #3967 let a COBOL word hold any Unicode letter, but a Japanese word also holds full-width digits (0-9) and the full-width hyphen, which SJIS 0x817C decodes to U+FF0D under cp932 and U+2212 under Python's shift_jis. Data names were cut at those characters (項目2 -> 項目, TEST−DATA1 -> TEST) and sections holding them (S−初期化, s123) were missing from function_data. - identifiers.py: WIDE_DIGITS / WIDE_HYPHENS. - cobol.py, mainframe_boundary.py (COBOL readers), file_control.py (COBOL FD / SELECT / COPY): every COBOL word class that takes 0-9 takes WIDE_DIGITS, and every one that takes the in-word '-' takes WIDE_HYPHENS, at the same positions. JCL / IDCAMS / CSD / HLASM / BMS classes are unchanged. Quantifiers unchanged. - detector._extract_name: for COBOL the full-width hyphens are name characters, so `S−初期化` is not reduced to its tail `初期化`. - tests: data names, sections, callees, both decodings of 0x817C from a cp932-encoded file, spaced minus stays arithmetic, and a structural check that every cobol.py word class carries the wide characters. Claude-Session: https://claude.ai/code/session_01EC8FWfn3aPmsUPnRgKNVup Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Resolves #3955
Defect:
Japanese user-defined words (like
PROGRAM-ID. 名前.and01 データ項目) were completely ignored or truncated by GitGalaxy's boundary extractors. The root cause was thatNATIONALinsideidentifiers.pywas hardcoded to only include EBCDIC special characters (e.g.ÆØÅÄÖÜѧ£) and did not include Unicode ranges like Kanji, Katakana, or Hiragana, which are fully supported as letters by modern IBM Enterprise COBOL and open-source Japanese COBOL compilers.Fix:
I appended
ID_START(which encompasses all valid Unicode lettersLo,Lu,Ll, etc.) directly to theNATIONALdefinition insidegitgalaxy/standards/language_standards/identifiers.py.Since
[A-Z" + NATIONAL + r"0-9-]is uniformly used as the core identifier matching block acrossmainframe_boundary.pyandjob_flow.py, this single change natively grants full internationalization support to all COBOL boundaries (PROGRAM-ID, Section headers, Data names) while preserving the strictly-required "must contain at least one letter" identifier rule (becauseID_STARTcontains zero digits).