Use safe scalar code for UTF-8 Equals(OrdinalIgnoreCase) - #134991
Merged
Merged
Conversation
Ordinal.EqualsIgnoreCaseUtf8 is only used by UTF-8 floating-point parsing to match short NaN/Infinity symbols, so replace the vectorized/unsafe implementation with the same safe scalar loop used by StartsWithIgnoreCaseUtf8. The ASCII loop has no calls so it stays cheap when inlined into callers; runes are decoded in a separate non-ASCII fallback. Remove the now unused Utf8Utility ignore-case helpers. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c5a38d5-eacd-41f0-b96d-4bcc0eecb0f4
|
Azure Pipelines: Successfully started running 3 pipeline(s). 13 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
Contributor
|
Tagging subscribers to this area: @dotnet/area-system-globalization |
- Compare 8 ASCII bytes at a time in EqualsIgnoreCaseUtf8 using safe BinaryPrimitives reads and the UInt64OrdinalIgnoreCaseAscii SWAR helper. - Reject ASCII vs non-ASCII bytes inline instead of calling the rune loop. - Keep StartsWithIgnoreCaseUtf8 on the single rune loop, as on main. - Mark EqualsIgnoreCaseUtf8 AggressiveInlining; its ASCII path has no calls and the rune loop stays out of line. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c5a38d5-eacd-41f0-b96d-4bcc0eecb0f4
Member
Author
|
@EgorBot -osx_arm64 -windows_intel -linux_amd using System.Globalization;
using System.Text;
using BenchmarkDotNet.Attributes;
using BenchmarkDotNet.Running;
BenchmarkSwitcher.FromAssembly(typeof(Bench).Assembly).Run(args);
public class Bench {
static readonly NumberFormatInfo Inv = NumberFormatInfo.InvariantInfo;
// ru-RU symbols: NaN is 16 UTF-8 bytes (non-ASCII)
static readonly NumberFormatInfo Ru = new() {
NaNSymbol = "\u043D\u0435\u00A0\u0447\u0438\u0441\u043B\u043E",
PositiveInfinitySymbol = "\u221E",
NegativeInfinitySymbol = "-\u221E",
};
// sv-SE style minus sign (U+2212)
static readonly NumberFormatInfo U2212 = new() { NegativeSign = "\u2212" };
// LRM-wrapped signs (U+200E), the longest signs across ICU cultures
static readonly NumberFormatInfo Lrm = new() { PositiveSign = "\u200E+\u200E", NegativeSign = "\u200E-\u200E" };
// Artificial long ASCII symbol (>= 16 bytes)
static readonly NumberFormatInfo Long = new() { NaNSymbol = "Not-A-Number-Value" };
static readonly Dictionary<string, (string Input, NumberFormatInfo Nfi, double Expected)> Data = new() {
["Inv_Infinity"] = ("Infinity", Inv, double.PositiveInfinity),
["Inv_INFINITY"] = ("INFINITY", Inv, double.PositiveInfinity),
["Inv_NegInfinity"] = ("-Infinity", Inv, double.NegativeInfinity),
["Inv_PlusInfinity"] = ("+infinity", Inv, double.PositiveInfinity),
["Inv_NaN"] = ("NaN", Inv, double.NaN),
["Inv_nan"] = ("nan", Inv, double.NaN),
["Inv_MinusNaN"] = ("-nan", Inv, double.NaN),
["Inv_NaN_WS"] = (" NaN ", Inv, double.NaN),
["Inv_Invalid"] = ("abc", Inv, 0),
["Inv_Number"] = ("123.5", Inv, 123.5),
["Ru_NaN"] = ("\u043D\u0435\u00A0\u0447\u0438\u0441\u043B\u043E", Ru, double.NaN),
["Ru_NaN_Upper"] = ("\u041D\u0415\u00A0\u0427\u0418\u0421\u041B\u041E", Ru, double.NaN),
["Ru_Infinity"] = ("\u221E", Ru, double.PositiveInfinity),
["Ru_NegInfinity"] = ("-\u221E", Ru, double.NegativeInfinity),
["Ru_Invalid"] = ("abc", Ru, 0),
["U2212_MinusNaN"] = ("\u2212nan", U2212, double.NaN),
["Lrm_PlusInfinity"] = ("\u200E+\u200EInfinity", Lrm, double.PositiveInfinity),
["Long_NaN"] = ("not-a-number-value", Long, double.NaN),
};
public static IEnumerable<string> Cases => Data.Keys;
[ParamsSource(nameof(Cases))]
public string Case { get; set; } = "";
byte[] _input = default!;
NumberFormatInfo _nfi = default!;
[GlobalSetup]
public void Setup() {
var (input, nfi, expected) = Data[Case];
_input = Encoding.UTF8.GetBytes(input);
_nfi = nfi;
if (!Parse().Equals(expected)) {
throw new Exception($"{Case}: unexpected result");
}
}
[Benchmark]
public double Parse() {
double.TryParse(_input, NumberStyles.Float, _nfi, out double result);
return result;
}
}Note This benchmark comment was generated with GitHub Copilot. |
This was referenced Oct 1, 2026
Member
Author
|
PTAL @MihaZupan follow up to #134857 |
NLS doesn't case-fold supplementary characters, so U+10400 and U+10428 don't match under OrdinalIgnoreCase for both UTF-16 and UTF-8 parsing. Fixes dotnet#135014 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c5a38d5-eacd-41f0-b96d-4bcc0eecb0f4
Member
Author
|
Had to change the new test I added in #134857 Becuase it was not NLS-friendly Mininal repro Console.WriteLine("𐐀Infinity".StartsWith("𐐨", StringComparison.OrdinalIgnoreCase));prints |
EgorBo
enabled auto-merge (squash)
October 1, 2026 14:49
Member
Author
|
/ba-g failures are #135031 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow up to #134857
Ordinal.EqualsIgnoreCaseUtf8is only used by UTF-8 floating-point parsing to match the NaN/Infinity symbols, so the vectorizedref/Unsafe-based implementation is replaced with safe span-based code:BinaryPrimitives.ReadUInt64LittleEndianand theUInt64OrdinalIgnoreCaseAsciiSWAR helper, then byte by byte. It has no calls, so it's inlined intoTryMatchSpecialValueSymbol.StartsWithIgnoreCaseUtf8, which is unchanged.The now unused
Utf8Utility.UInt32OrdinalIgnoreCaseAsciiandVector128OrdinalIgnoreCaseAsciiare removed.Benchmark
double.TryParseon UTF-8 input; EgorBot results, benchmark source in this comment. Cells aremain → PR (PR/main); below 1.00 is faster.Inv_*: invariant culture.Inv_Number(123.5) doesn't reach the changed code.Ru_*: ru-RU symbols;NaNSymbolisне число(16 bytes) and the infinity symbols are∞/-∞.U2212_MinusNaN:NegativeSignis U+2212 (e.g. sv-SE).Lrm_PlusInfinity: signs wrapped in U+200E.Long_NaN: artificial 18-byte ASCIINaNSymbol.