What's wrong
ReplaceNonAlphaNumericWithSpace keeps a code point only if char.IsLetter is true or it is '0'..'9'. Characters in category LetterNumber (Nl) fail both tests, so they are replaced with a space and dropped. Nl covers the Roman numeral characters Ⅰ–ↈ and ⅰ–ⅿ (U+2160–U+2188), Gothic and Old Persian number letters, and similar characters. These characters are cased (Ⅻ ↔ ⅻ), and ToLowercaseFirstChar already maps between the two forms. ToTitleCase does not go through this path and keeps them, so the converters disagree on the same input.
Repro (main @ 4a74d14)
| Input |
Call |
Actual |
Expected |
"chapter Ⅻ" |
ToTitleCase() |
Chapter Ⅻ |
(correct) |
"chapter Ⅻ" |
ToPascalCase() |
Chapter |
ChapterⅫ |
"chapter Ⅻ" |
ToSnakeCase() / ToMacroCase() |
chapter / CHAPTER |
chapter_ⅻ / CHAPTER_Ⅻ |
"Ⅻ" |
ToPascalCase() / ToSnakeCase() |
"" |
Ⅻ / ⅻ |
Why it matters
Text is deleted with no error, and an input made only of these characters becomes the empty string. This is the same class of bug as #70 (astral letters deleted) and #80 (combining marks and non-ASCII digits deleted). #80 covers marks and char.IsDigit but not LetterNumber, so its fix would not cover this case.
Suggested fix
In ReplaceNonAlphaNumericWithSpace, also keep code points whose CharUnicodeInfo.GetUnicodeCategory is LetterNumber. Decide whether IsWordBoundary treats them as letters (no break in "chapterⅫ") or as digits (a break), and apply the choice the same way in every converter. Add rows for "chapter Ⅻ" and "Ⅻ" across every converter.
What's wrong
ReplaceNonAlphaNumericWithSpacekeeps a code point only ifchar.IsLetteris true or it is'0'..'9'. Characters in categoryLetterNumber(Nl) fail both tests, so they are replaced with a space and dropped. Nl covers the Roman numeral characters Ⅰ–ↈ and ⅰ–ⅿ (U+2160–U+2188), Gothic and Old Persian number letters, and similar characters. These characters are cased (Ⅻ ↔ ⅻ), andToLowercaseFirstCharalready maps between the two forms.ToTitleCasedoes not go through this path and keeps them, so the converters disagree on the same input.Repro (main @ 4a74d14)
"chapter Ⅻ"ToTitleCase()Chapter Ⅻ"chapter Ⅻ"ToPascalCase()ChapterChapterⅫ"chapter Ⅻ"ToSnakeCase()/ToMacroCase()chapter/CHAPTERchapter_ⅻ/CHAPTER_Ⅻ"Ⅻ"ToPascalCase()/ToSnakeCase()""Ⅻ/ⅻWhy it matters
Text is deleted with no error, and an input made only of these characters becomes the empty string. This is the same class of bug as #70 (astral letters deleted) and #80 (combining marks and non-ASCII digits deleted). #80 covers marks and
char.IsDigitbut notLetterNumber, so its fix would not cover this case.Suggested fix
In
ReplaceNonAlphaNumericWithSpace, also keep code points whoseCharUnicodeInfo.GetUnicodeCategoryisLetterNumber. Decide whetherIsWordBoundarytreats them as letters (no break in "chapterⅫ") or as digits (a break), and apply the choice the same way in every converter. Add rows for"chapter Ⅻ"and"Ⅻ"across every converter.