Skip to content

Ignore case for letters outside the Basic Multilingual Plane [patch] - #107

Merged
matt-edmondson merged 1 commit into
mainfrom
fix/supplementary-case-90
Oct 6, 2026
Merged

matt-edmondson merged 1 commit into
mainfrom
fix/supplementary-case-90

Conversation

@matt-edmondson

Copy link
Copy Markdown
Contributor

Fixes #90

What was wrong

Fuzzy.CodepointsEqual compared a surrogate pair one code unit at a time. Letters outside the Basic Multilingual Plane, such as Deseret, Osage and Adlam, therefore matched only in the exact case, although the README says case is always ignored.

Change

  • Comparison: a surrogate pair is now compared as one codepoint in the new SurrogatePairsEqualIgnoringCase. Like CharsEqualIgnoringCase (Case-insensitive matching fails for Greek final sigma and micro sign: Contains("ΟΔΟΣ", "οδος") is false #88), it matches on either the lowercase or the uppercase form.
    • On .NET Core 3.0 and later it uses Rune.ToLowerInvariant / Rune.ToUpperInvariant.
    • On netstandard2.0 and netstandard2.1, where Rune doesn't exist, it falls back to invariant string casing of the two-character pair.
  • Remark: the XML remark on CodepointsEqual now gives the real reason a pair has to be compared whole. It no longer says supplementary codepoints are compared exactly.

Not in this PR

The camelCase bonus still reads case from the lead surrogate, so a supplementary-plane uppercase letter does not earn it. The issue lists this as optional. Fixing it changes the internal ApplyBonuses(…, char strChar, char strLower, char strUpper, …) signature that the existing tests call directly, so it is better as its own change. It does not affect the acceptance criteria: the other-case match now scores the same as the exact-case one.

Tests

New tests in the Surrogate Pair region:

  • Deseret and Adlam in both directions, plus a two-letter Osage case
  • Each other-case score equals the exact-case score
  • Two different Deseret letters still don't match

Results:

  • Before the fix, 8 of the new cases fail.
  • After the fix, the full suite passes: 74 of 74 on net10.0, the test project's only target.
  • The library builds clean, with 0 warnings, for net10.0, net9.0, net8.0, netstandard2.0 and netstandard2.1.
  • The netstandard fallback is compiled but not run by the tests, because there is no netstandard test leg.

🤖 Generated with Claude Code

https://claude.ai/code/session_01QnUaFtfBEaU4WcCPkkdpvM


Generated by Claude Code

CodepointsEqual compared a surrogate pair code unit by code unit, so
supplementary-plane letters such as Deseret, Osage and Adlam only matched in
the exact case, although case is always meant to be ignored. A surrogate
pair is now compared as one codepoint, on its lowercase or uppercase form
like CharsEqualIgnoringCase: through Rune where it exists, and through the
invariant string casing on netstandard.

Fixes #90

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QnUaFtfBEaU4WcCPkkdpvM
@sonarqubecloud

sonarqubecloud Bot commented Oct 6, 2026

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Case-insensitive matching fails for supplementary-plane letters: Contains("𐐀", "𐐨") is false (Deseret, Osage, Adlam, …)

1 participant