Description
lexer::scan_number() eagerly converts every numeric token with strtoull/strtoll/strtof at the end of scanning (lexer.hpp:1282-1327). The syntactic validity of the number is already decided by the scanner's state machine before that conversion; the conversion only produces the numeric value.
For accept() and any SAX validation consumer, that value is immediately discarded. The conversion's only remaining effect on the accept/reject decision is the parser's non-finite check (parser.hpp:266-275):
const auto res = m_lexer.get_number_float();
if (JSON_HEDLEY_UNLIKELY(!std::isfinite(res)))
return sax->parse_error(..., out_of_range::create(406, "number overflow ..."));
So the full conversion is performed on every number in the document purely to answer "does this float overflow to ±inf?" — and for integer-classified tokens, not even that:
- A
value_unsigned / value_integer token is, by construction, one whose strtoull/strtoll succeeded without ERANGE — i.e. it fits 64 bits and is therefore always finite and always accepted. Its conversion result is never used by accept().
- Only a
value_float token can be rejected, and only when its magnitude exceeds the double range (≈1.8e308). That is a decimal-magnitude question that does not require a full strtof.
(Confirmed behavior: accept("1e999") == false, accept("9999999999999999999999999999") == true — the 28-digit integer overflows uint64 but is finite as a double, so it is accepted as a float.)
Measurement
g++ 13.3.0 -O2, accept() over a 100k-element array, comparing the current code against a build that short-circuits scan_number's conversion (upper bound):
| input |
baseline accept |
conversion skipped |
speedup |
| integers (100k) |
7.09 ms |
4.12 ms |
1.7× |
| floats (100k) |
17.74 ms |
6.35 ms |
2.8× |
On float-heavy input the strtof conversion is ~64% of accept()'s total time and is entirely discarded.
Relationship to #5283
This is distinct from the open parser-performance PR #5283. That PR speeds up number conversion (Clinger fast path, contiguous fast path) but still performs the full conversion for the validation path — its accept() gains come from bulk string/UTF-8 scanning, not from numbers. Eliminating the conversion for validation is an orthogonal, additive win and is most naturally implemented on top of that PR's number_parse.hpp rework.
Suggested direction
Make the value computation lazy / skippable on the validation path, preserving accept/reject exactly:
- Thread a compile-time "values not required" property from the SAX consumer (the acceptor /
json_sax_acceptor) so the parser can skip get_number_*() + sax->number_*() for tokens that cannot be rejected.
- For
value_unsigned / value_integer tokens: skip conversion entirely (always finite, always accepted). The scanner still needs to detect the overflow→float reclassification, which can be done by digit count (≤18 digits always fit; ≥21 always overflow uint64; only 19–20 need a precise compare) instead of a full strtoull.
- For
value_float tokens: replace strtof with a cheap decimal-magnitude finiteness check (reject when the effective base-10 exponent pushes the value outside the double range), keeping the exact out_of_range.406 boundary.
Full parse() is unaffected — it still calls get_number_*() once per number and converts exactly as today.
Behavior preservation
The optimization must keep accept()'s current results bit-for-bit, including: 1e999/1e400 rejected; huge-but-finite integers (e.g. 28 digits) accepted as finite doubles; the exact 406 overflow boundary. A differential test (accept old vs new over a large mixed + edge corpus, including magnitudes straddling DBL_MAX) should gate it.
Compiler and operating system
g++ 13.3.0 (Ubuntu 24.04, x86-64), -O2
Library version
develop @ 01853ed6bcf9ebe88ec2e248ea757b86417b1487
Description
lexer::scan_number()eagerly converts every numeric token withstrtoull/strtoll/strtofat the end of scanning (lexer.hpp:1282-1327). The syntactic validity of the number is already decided by the scanner's state machine before that conversion; the conversion only produces the numeric value.For
accept()and any SAX validation consumer, that value is immediately discarded. The conversion's only remaining effect on the accept/reject decision is the parser's non-finite check (parser.hpp:266-275):So the full conversion is performed on every number in the document purely to answer "does this float overflow to ±inf?" — and for integer-classified tokens, not even that:
value_unsigned/value_integertoken is, by construction, one whosestrtoull/strtollsucceeded withoutERANGE— i.e. it fits 64 bits and is therefore always finite and always accepted. Its conversion result is never used byaccept().value_floattoken can be rejected, and only when its magnitude exceeds thedoublerange (≈1.8e308). That is a decimal-magnitude question that does not require a fullstrtof.(Confirmed behavior:
accept("1e999") == false,accept("9999999999999999999999999999") == true— the 28-digit integer overflowsuint64but is finite as adouble, so it is accepted as a float.)Measurement
g++ 13.3.0
-O2,accept()over a 100k-element array, comparing the current code against a build that short-circuitsscan_number's conversion (upper bound):acceptOn float-heavy input the
strtofconversion is ~64% ofaccept()'s total time and is entirely discarded.Relationship to #5283
This is distinct from the open parser-performance PR #5283. That PR speeds up number conversion (Clinger fast path, contiguous fast path) but still performs the full conversion for the validation path — its
accept()gains come from bulk string/UTF-8 scanning, not from numbers. Eliminating the conversion for validation is an orthogonal, additive win and is most naturally implemented on top of that PR'snumber_parse.hpprework.Suggested direction
Make the value computation lazy / skippable on the validation path, preserving
accept/reject exactly:json_sax_acceptor) so the parser can skipget_number_*()+sax->number_*()for tokens that cannot be rejected.value_unsigned/value_integertokens: skip conversion entirely (always finite, always accepted). The scanner still needs to detect the overflow→float reclassification, which can be done by digit count (≤18 digits always fit; ≥21 always overflowuint64; only 19–20 need a precise compare) instead of a fullstrtoull.value_floattokens: replacestrtofwith a cheap decimal-magnitude finiteness check (reject when the effective base-10 exponent pushes the value outside thedoublerange), keeping the exactout_of_range.406boundary.Full
parse()is unaffected — it still callsget_number_*()once per number and converts exactly as today.Behavior preservation
The optimization must keep
accept()'s current results bit-for-bit, including:1e999/1e400rejected; huge-but-finite integers (e.g. 28 digits) accepted as finite doubles; the exact406overflow boundary. A differential test (acceptold vs new over a large mixed + edge corpus, including magnitudes straddlingDBL_MAX) should gate it.Compiler and operating system
g++ 13.3.0 (Ubuntu 24.04, x86-64),
-O2Library version
develop@01853ed6bcf9ebe88ec2e248ea757b86417b1487