Skip to content

Speed up the lexer: own float parser, string scan, and \u table - #5738

Open
nlohmann wants to merge 1 commit into
developfrom
json-view/02b-float-parser
Open

nlohmann wants to merge 1 commit into
developfrom
json-view/02b-float-parser

Conversation

@nlohmann

@nlohmann nlohmann commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Part of the stack for the zero-copy view (#5295). This change speeds up json::parse on its own; the view uses the same number converter and string kernels later in the stack. This PR combines #5738, #5618, and #5619, which were reviewed separately before.

Summary

  • Own float converter for binary32/binary64. float, double, and long double where it is IEEE 754 binary64 (MSVC, Apple arm64) are now converted by the library itself: correctly rounded (to nearest, ties to even), independent of the locale, and without calling the C or C++ library.
    1. Split the token into sign, significand w (at most 19 digits), and decimal exponent q. The scanners already record the positions of the decimal point and the exponent, so no character is classified again; digits are read eight at a time.
    2. Clinger's fast path where w and 10^|q| are exact (skipped where FLT_EVAL_METHOD != 0).
    3. Eisel-Lemire otherwise, templated for binary32 and binary64.
    4. For tokens with more than 19 significant digits whose w and w + 1 round differently: an exact big-integer comparison with the midpoint between the two candidates (fast_float's digit comparison, simplified because Eisel-Lemire already yields both candidates).
    • Parsed values are bit-identical wherever the previous chain (Eisel-Lemire in front of std::from_chars/strtod) was correctly rounded. Out-of-range values still throw out_of_range.406, and underflow still gives a zero with the sign of the token.
    • eisel_lemire() and decimal_to_float() are always inlined, so callers' hot loops keep the whole conversion inline; this makes the view's traversal of canada.json 8% faster later in the stack.
    • The strtold fallback left for long double formats other than binary64 (x87, binary128) now copies a decimal point longer than one byte into a copy of the token instead of substituting its first byte, fixing Floats are truncated at the decimal point under locales with a multi-byte decimal point (fa_IR.UTF-8) #5660 completely for a locale such as fa_IR.UTF-8.
  • Faster string scanning. The SWAR kernels in string_scan.hpp find the first stop byte with the trailing-zero count of the comparison mask instead of a byte loop once a word contains a hit (the lowest flagged byte is always a real hit, since borrows can only flag bytes above one). Words are assembled in little-endian order on every platform. scalar_string_bulk_run() validates a run of multi-byte UTF-8 sequences one after another instead of re-searching for the next special byte after each sequence, which helps text in non-Latin scripts. These kernels serve the lexer's contiguous fast path, the serializer, and the binary formats.
  • Faster \u escape decoding. get_codepoint() decoded four hex digits with four calls to get(), each going through a chain of range comparisons. For contiguous input it now uses one table lookup per byte (after yyjson's read_hex_u16); an invalid digit shows in the OR of the four values, and the lexer then skips the four bytes and updates position counters as before. The streaming path and all error positions are unchanged.
  • Docs (number_handling.md, template_parameters.md, number_float_t.md) no longer say parsing uses strtod/strtof/strtold, except for non-binary64 long double.
  • Credit: the fast_float entry in README.md/license.md now also names the digit comparison.

Performance

json_view

The view uses these converters and kernels from group 3 on; it is not affected directly by this PR.

Core library (json::parse / dump)

Measured on the regrouped stack

Measured on the regrouped stack (µs, best of 3 interleaved rounds of 15 runs; files from nativejson-benchmark). Apple M1 Max with Apple clang at -O2; x86-64 on a KVM Haswell VPS, pinned to one core, GCC 13 / Clang 18 at -O2 (the VPS is noisy, about ±5–10%).

json::parse and json::dump, develop → this PR:

Apple M1 x86-64 GCC x86-64 Clang
parse canada 8958 → 7529 (−16%) 16494 → 15553 (−6%) 15709 → 14950 (−5%)
parse twitter 1528 → 1402 (−8%) 4445 → 4034 (−9%) 4386 → 3990 (−9%)
parse citm_catalog 2994 → 2974 (−1%) 7898 → 7843 (−1%) 6511 → 6546 (+1%)
dump twitter 490 → 411 (−16%) 1058 → 1007 (−5%) 1075 → 875 (−19%)
dump citm_catalog 534 → 521 (−2%) 1382 → 1352 (−2%) 1019 → 1144 (+12%, noise)

Earlier measurements (on the old stack)

Apple clang, -O2, end-to-end json::parse from a string, ns per number (structure included), vs. develop:

input vs. develop
canada.json 0.88
mesh.json 0.95
numbers.json 0.90
100k random doubles 0.87
canada.json, float json 0.79

json::parse in a separate process, best of 5, Apple M1 Max:

file change
twitterescaped.json (\u table) −14%
poet.json (string scan) −27%
random.json (string scan) −8%
twitter.json (string scan) −7%

tests/benchmarks, median of interleaved repetitions (string scan also speeds up dump()):

benchmark change
Dump/twitter −16%
Dump/jeopardy −7%
Dump/citm_catalog −5%
other ParseString/Dump rows within ±2%

Tests

  • 508 generated hard float-parsing cases with expected binary32 and binary64 bits (exact midpoints and their neighbours, boundary values, huge exponents, long tokens), checked through the converter and through json::parse with both scanners.
  • Exact-bit tests for double and float (ties to even, subnormal/overflow boundaries), a 200,000-value round trip, and a new 100,000-value float round trip.
  • The string-scan kernels are compared against byte-by-byte reference scans on 100,000 generated buffers at three alignments; the \u table is checked for valid escapes, surrogate pairs, truncation at every distance from the end, an invalid digit at each of the four positions, and 3,000 seeded random escapes.
  • unit-locale-cpp.cpp checks a multi-byte decimal point for double exactly and for long double against the "C" locale.
  • 0 mismatches against strtod_l/strtof_l over several million random tokens and the full 240,187 hard cases; clean under ASan/UBSan and the CI's compiler warning flags, including GCC 16, clang-tidy, and x86_64 under Rosetta (x87 long double).
Generator of float_hard_cases.hpp

The 508 embedded cases are produced by compact_hard_cases.py 5 (importing hard_cases.py, the fuller generator used for the 240,187-case offline checks), then formatted as one C++ initializer per output line. For every format boundary and a set of random values, it computes the exact midpoint to the next representable value and emits that midpoint, one unit above/below it, the midpoint with digits appended just past the rounding boundary, and the midpoint truncated at several digit counts around 17–30, in all three JSON number notations, with 30% negative. Both scripts round with exact rational arithmetic and cross-check against Python's float().

Public API

No breaking changes. Everything added or removed is in nlohmann::detail. Parsed values are unchanged wherever the previous conversion was correctly rounded; they change only where it was not (a locale with a multi-byte decimal point, or a C library whose strtod is not correctly rounded). No change to dump() or any other output.

Fixes #5660.


Written by Claude Code.

🤖 Generated with Claude Code

@nlohmann
nlohmann added this pull request to stack #5739 September 30, 2026 13:21
@nlohmann
nlohmann force-pushed the json-view/02b-float-parser branch from d92b762 to 39092df Compare September 30, 2026 18:06
Base automatically changed from json-view/02-eisel-lemire to develop September 30, 2026 18:06
@nlohmann
nlohmann force-pushed the json-view/02b-float-parser branch from 39092df to 192b99b Compare September 30, 2026 18:06
@nlohmann
nlohmann marked this pull request as ready for review September 30, 2026 18:15
@nlohmann
nlohmann force-pushed the json-view/02b-float-parser branch from 192b99b to 9c71689 Compare September 30, 2026 18:19
@nlohmann nlohmann added the review needed It would be great if someone could review the proposed changes. label Sep 30, 2026
- The library converts integers and floating-point numbers itself, independent of the locale. Floating-point
numbers are correctly rounded (to nearest, ties to even). Only a `#!c long double` that is not IEEE 754 binary64
(e.g., the 80-bit x87 format) is converted with `#!cpp std::from_chars` where available, or with
[`std::strtold`](https://en.cppreference.com/w/cpp/string/byte/strtof), which gets the decimal point of the

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't we convert the decimal point back to period?

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes. The token always holds .. Only the strtold fallback (for a long double that is not binary64) swaps it for the locale's decimal point during the call, and puts . back afterwards. So the result does not depend on the locale. The note read as if the input had to use the locale's decimal point; 9d88ead rewrites it.


This comment was written by Claude Code on behalf of @nlohmann.

Comment thread tests/src/unit-class_lexer.cpp Dismissed
Give the library its own correctly rounded float converter for
binary32 and binary64 (IEEE 754), and speed up the lexer's string
and escape scanning.

The converter splits a number token into sign, significand, and
decimal exponent, then tries Clinger's fast path, then a templated
Eisel-Lemire step, and falls back to an exact big-integer digit
comparison for tokens with more than 19 significant digits whose two
candidate values round differently. This replaces std::from_chars
and strtod/strtof for both formats, so parsed values no longer
depend on the C/C++ library or the current locale. The strtold
fallback kept for other long double formats (x87, binary128) now
also copies a multi-byte decimal point correctly, fixing #5660.
eisel_lemire() and decimal_to_float() are always inlined so callers
keep the whole conversion in their hot loop.

The string-scanning kernels in string_scan.hpp find a stop byte with
the trailing-zero count of the SWAR mask instead of a byte loop, and
scalar_string_bulk_run() validates a run of multi-byte UTF-8
sequences one after another instead of re-searching after each one.

get_codepoint() decodes a contiguous \uXXXX escape with one table
lookup per byte instead of four range-checked get() calls; the
streaming path and all error positions are unchanged.

Adds 508 generated hard float-parsing cases with expected binary32
and binary64 bits, and kernel-comparison tests for the string scans
and the escape table against byte-by-byte references.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
@nlohmann
nlohmann removed this pull request from stack #5739 October 6, 2026 09:26
@nlohmann
nlohmann force-pushed the json-view/02b-float-parser branch from 15b0cc0 to 953d74d Compare October 6, 2026 09:28
@nlohmann nlohmann changed the title Convert float and double with the library's own correctly rounded parser Speed up the lexer: own float parser, string scan, and \u table Oct 6, 2026
@nlohmann
nlohmann added this pull request to stack #5768 October 6, 2026 09:29

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation L review needed It would be great if someone could review the proposed changes. tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Floats are truncated at the decimal point under locales with a multi-byte decimal point (fa_IR.UTF-8)

3 participants