Skip to content

Skip integer conversion in accept()/SAX validation when the value is unused - #5484

Merged
nlohmann merged 3 commits into
developfrom
issue-5411-lexer-skip-conversion
Sep 9, 2026
Merged

nlohmann merged 3 commits into
developfrom
issue-5411-lexer-skip-conversion

Conversation

@nlohmann

@nlohmann nlohmann commented Sep 5, 2026

Copy link
Copy Markdown
Owner

Summary

Fixes #5411.

lexer::scan_number() eagerly converts every numeric token with strtoull()/strtoll()/strtof() at the end of scanning, even though for accept() (and any SAX consumer that discards the numeric value, i.e. json_sax_acceptor) that converted value is never used. The only thing the conversion still affects for accept()/reject is the parser's non-finite check on value_float tokens.

This PR threads a discard_number_values flag from basic_json::accept() down through parser/lexer. When set, lexer::scan_number() can skip strtoull()/strtoll() entirely for value_unsigned/value_integer tokens once the digit count alone guarantees the value fits into 64 bits (a decimal number with up to 18 digits always fits into both std::int64_t and std::uint64_t, so errno could never have been set to ERANGE). Such tokens are, by construction, always finite and unconditionally accepted regardless of their actual value, so only the token classification is needed — not the converted value.

Numbers with 19+ digits (rare in practice) fall through to the exact, completely unmodified conversion code, so their handling — including reclassification to value_float when the value overflows 64 bits, and the existing 406 "number overflow" rejection when that reclassified value is not even finite as a double — is bit-for-bit identical to before this change.

Scope and what was deliberately left out

The originating issue's "suggested direction" also proposed (2) a cheap digit-count-based reclassification check to fully replace strtoull/strtoll even near the 64-bit boundary, and (3) a cheap decimal-magnitude finiteness check to replace strtof for value_float tokens. Both are out of scope for this PR:

  • value_float handling is completely unchanged — the existing strtof/strtod/strtold conversion is still always called, and the parser's std::isfinite check on it is untouched.
  • For value_unsigned/value_integer, the fast path is only taken when the digit count makes 64-bit overflow provably impossible (≤18 digits); everything else (≥19 digits) uses the exact original code path, unchanged. I judged a full digit-count-based reclassification of the ≥19-digit boundary cases to be higher risk for a security-sensitive parser without more extensive validation than fits this PR, so I scoped this down to the always-safe subset.

parse() (and the public sax_parse() API for user-supplied SAX consumers) are completely unaffected — they never set discard_number_values, so they keep calling the full conversion exactly as before.

Behavior preservation / testing

  • Preserved invariants verified: accept("1e999") == false, accept("9999999999999999999999999999") == true (28-digit integer, overflows uint64_t but is finite as double).
  • Added a differential regression test in tests/src/unit-class_parser.cpp (SECTION("issue #5411 - skip conversion when accept() does not need the numeric value")) comparing json::accept() against json::parse() across: small/large integers (both signs), the 18/19/20-digit boundary (both signs), the 64-bit boundaries (INT64_MAX/INT64_MIN/UINT64_MAX/UINT64_MAX+1), the 28-digit example from the issue, huge digit-only integers that overflow even a double (must reject), 1e999/1e400-style exponent overflow (must reject), values straddling DBL_MAX, and a mix of valid/invalid numeric syntax.
  • Ran a standalone differential harness comparing accept() output byte-for-byte between the pre-change and post-change headers (both the split include/ headers and the amalgamated single_include/nlohmann/json.hpp) over ~2000 cases (hand-picked edge cases + randomized fuzz corpus of digit/exponent combinations): zero mismatches.
  • Ran the full required test files locally (offline, without the json_test_data submodule available in this environment): unit-class_lexer.cpp, unit-class_parser.cpp, unit-deserialization.cpp, unit-regression1.cpp, unit-regression2.cpp, unit-udt.cpp, unit-class_parser_diagnostic_positions.cpp — all pass, with the exception of 7 pre-existing assertions in unit-regression1.cpp that fail identically on develop in this offline environment because they read fixture files from the (here, unavailable) json_test_data test corpus; unrelated to this change (verified by running the identical test against unmodified develop and getting the same 7 failures).
  • make amalgamate was run and single_include/nlohmann/json.hpp/json_fwd.hpp are included in this PR.

Breaking change?

No breaking changes. This only adds a new defaulted trailing parameter to the internal (nlohmann::detail::parser/lexer, and the JSON_PRIVATE_UNLESS_TESTED basic_json::parser(...) factory) constructors; the public basic_json::accept()/parse()/sax_parse() API signatures are unchanged.

— opened by Claude Code on behalf of @nlohmann

@nlohmann nlohmann added the review needed It would be great if someone could review the proposed changes. label Sep 5, 2026
@nlohmann

nlohmann commented Sep 5, 2026

Copy link
Copy Markdown
Owner Author

Verified with a local benchmark (json::accept() over a 100k-element JSON array, clang++ -O2 -DNDEBUG, Apple Clang 21 / arm64 macOS, median of 80 timed iterations, comparing develop @ 09b6b6b vs this branch's tip):

  • Integer array (100k plain integers, ~1.04 MB): 4.88 ms before → 3.41 ms after (about 30% faster)
  • Float array (100k floats, ~1.14 MB): 5.71 ms before → 5.86 ms after (no measurable change, within noise — expected, since the fast path only applies to integer/unsigned tokens)

The integer-array speedup matches the PR's intent: skipping strtoull()/strtoll() for accept()-only parses whose digit count already guarantees no 64-bit overflow, while leaving float parsing and the exact/value-returning path untouched.

— posted by Claude Code on behalf of @nlohmann

Comment thread include/nlohmann/detail/input/lexer.hpp
…unused

lexer::scan_number() always converted every numeric token with
strtoull()/strtoll() before returning, even though accept() (and any
consumer using json_sax_acceptor) immediately discards the converted
value. For value_unsigned/value_integer tokens whose digit count
already guarantees the value fits into 64 bits, the conversion cannot
change the accept/reject decision (such tokens are always finite and
unconditionally accepted), so scan_number() can skip strtoull()/
strtoll() entirely in that case when the caller signals it does not
need the value. Numbers with more digits keep using the exact,
unmodified conversion path, so overflow reclassification to
value_float (and the finiteness check on it) is unaffected.

parse() and value_float handling are completely unchanged.

Fixes #5411

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
…nsigned_t/number_integer_t width

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
…ferential test

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
@nlohmann
nlohmann force-pushed the issue-5411-lexer-skip-conversion branch from 68c6981 to 03270c6 Compare September 8, 2026 11:12
@nlohmann nlohmann added 🚀 ready to merge Ready to merge - just waiting for CI to complete. and removed review needed It would be great if someone could review the proposed changes. labels Sep 8, 2026
@nlohmann
nlohmann merged commit faa35cc into develop Sep 9, 2026
157 of 158 checks passed
@nlohmann
nlohmann deleted the issue-5411-lexer-skip-conversion branch September 9, 2026 07:46
@nlohmann nlohmann added this to the Release 3.13.0 milestone Sep 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

L 🚀 ready to merge Ready to merge - just waiting for CI to complete. tests

Projects

None yet

2 participants