Follow-up to #3955 (fixed by #3967). Found while rescanning the opensourcecobol4j fresh estate (#3806, epic #3803) on origin/main 6991970d. This was a pre-trial scan with no porting.
Defect
#3967 lets COBOL words contain Unicode letters. Japanese COBOL words also contain two other kinds of character that the name classes still reject:
What happens:
- Data names are truncated at the first such character. Different items can then collide:
TEST−DATA1 and TEST−RECORD1 both become TEST.
- Sections are dropped entirely.
主処理 SECTION is found, but S−初期化 SECTION, S-終了 SECTION and s123 SECTION are not in function_data at all.
On the estate (independent Unicode-aware reading vs the master DB)
| suite |
Japanese data names truncated (before #3967 -> now) |
Japanese sections/paragraphs missing (before -> now) |
| i18n_sjis |
22 -> 14 |
11 -> 9 |
| cobol_utf8 |
22 -> 7 |
16 -> 12 |
| i18n_utf8 |
15 -> 4 |
9 -> 7 |
Every remaining truncation is at a full-width hyphen or digit: O−文字列 -> O, 項目2 -> 項目, TEST−DATA1 -> TEST, 横浜−1 -> 横浜, wk-03 -> WK, 項目ABCDEFGH012345 -> 項目ABCDEFGH.
(The missing-data-name count went to 0 and the Japanese PROGRAM-IDs are all refracted now. #3967 did fix the first half.)
Minimal reproduction (origin/main 6991970)
mkdir -p estate
cat > estate/JPWORDS.cbl <<'COB'
IDENTIFICATION DIVISION.
PROGRAM-ID. JPWORDS.
DATA DIVISION.
WORKING-STORAGE SECTION.
01 項目2 PIC X(4).
01 O−文字列 PIC X(4).
01 WK-03 PIC X(4).
01 横浜-1 PIC X(4).
PROCEDURE DIVISION.
主処理 SECTION.
MOVE 'A' TO 項目2.
ASCII-SEC SECTION.
DISPLAY 'X'.
S−初期化 SECTION.
DISPLAY 'Y'.
s123 SECTION.
DISPLAY 'Z'.
S-終了 SECTION.
STOP RUN.
COB
GITGALAXY_DISABLE_GIT_HISTORY=1 python -m gitgalaxy.galaxyscope estate --db-only --output out
sqlite3 out/estate_galaxy_master.db "select item_name from record_data; select func_name from function_data"
Actual:
項目
O
WK
横浜-1
ASCII-SEC
主処理
Expected: 項目2, O−文字列, WK-03, 横浜-1, plus all five sections.
Notes
- Treat
- / − as the COBOL hyphen inside a word (the same rules as -: not first, not last). Accept full-width digits wherever ASCII digits are accepted. ID_CONTINUE already covers digits (Nd), so the COBOL class probably just needs to use it rather than ID_START alone after the first character.
- Case folding looks right:
wk is stored as WK.
- A test should cover both SJIS decodings of 0x817C (U+2212 via
shift_jis, U+FF0D via cp932).
🤖 Generated with Claude Code
Follow-up to #3955 (fixed by #3967). Found while rescanning the opensourcecobol4j fresh estate (#3806, epic #3803) on origin/main
6991970d. This was a pre-trial scan with no porting.Defect
#3967 lets COBOL words contain Unicode letters. Japanese COBOL words also contain two other kinds of character that the name classes still reject:
-U+FF0D, which cp932 decodes byte 0x817C to, and−U+2212, which Python'sshift_jisdecodes the same byte to. COBOL: Japanese user-defined words are invisible -- a Japanese PROGRAM-ID drops the program from refraction; data names truncated or missing; sections missing #3955 named both.0-9(U+FF10-FF19).What happens:
TEST−DATA1andTEST−RECORD1both becomeTEST.主処理 SECTIONis found, butS−初期化 SECTION,S-終了 SECTIONands123 SECTIONare not in function_data at all.On the estate (independent Unicode-aware reading vs the master DB)
Every remaining truncation is at a full-width hyphen or digit:
O−文字列 -> O,項目2 -> 項目,TEST−DATA1 -> TEST,横浜−1 -> 横浜,wk-03 -> WK,項目ABCDEFGH012345 -> 項目ABCDEFGH.(The missing-data-name count went to 0 and the Japanese PROGRAM-IDs are all refracted now. #3967 did fix the first half.)
Minimal reproduction (origin/main 6991970)
Actual:
Expected:
項目2,O−文字列,WK-03,横浜-1, plus all five sections.Notes
-/−as the COBOL hyphen inside a word (the same rules as-: not first, not last). Accept full-width digits wherever ASCII digits are accepted. ID_CONTINUE already covers digits (Nd), so the COBOL class probably just needs to use it rather than ID_START alone after the first character.wkis stored asWK.shift_jis, U+FF0D viacp932).🤖 Generated with Claude Code