Repository navigation
Move the float conversion chain out of the lexer - #5616
Merged
Merged
Conversation
nlohmann
added this pull request to stack #5636
September 29, 2026 14:20
gregmarr
approved these changes
Sep 29, 2026
nlohmann
marked this pull request as ready for review
September 29, 2026 14:35
7 tasks done
nlohmann
removed this pull request from stack #5636
September 30, 2026 13:19
nlohmann
added this pull request to stack #5739
September 30, 2026 13:21
Base automatically changed from
json-view/00-amalgamate-external
to
develop
September 30, 2026 18:05
lexer::convert_number() converted float tokens with std::from_chars (when available), Clinger's fast path, and the locale-aware strtod fallback, all as lexer members. They are now free functions in number_parse.hpp: - convert_float_fast(): std::from_chars, then Clinger's fast path, skipped when the mantissa has too many significant digits - convert_float_locale_aware(): strtof/strtod/strtold with the decimal point of the current locale, retried when the locale changed (#5198) so that other code converting JSON number tokens gets the same values. No change in behavior; the lexer no longer includes <clocale> and <cstdlib>. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
nlohmann
force-pushed
the
json-view/01-float-chain
branch
from
September 30, 2026 18:06
f6c115a to
5432675
Compare
nlohmann
added a commit
that referenced
this pull request
Oct 1, 2026
…5731) * Share the diagnostic-position setter of the DOM SAX parsers json_sax_dom_parser and json_sax_dom_callback_parser each had a private copy of handle_diagnostic_positions_for_json_value(), identical except for comments. Move the body into one static member function, detail::diagnostic_positions::set_from_lexer(value, lexer), which both classes call with their lexer pointer. basic_json befriends the new struct (only when JSON_DIAGNOSTIC_POSITIONS is enabled), as the position members are private. The discarded case is reached through the callback parser, so the LCOV_EXCL markers that only the dom parser's copy had are gone. The NOLINT on the unreachable default case loses the stray "-warnings-as-errors", which is not a check name. The start-position setup in start_object()/start_array() is left alone, as #5706 is editing the callback parser's versions. Behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Correct the parser comments on recursion and skip_to_state_evaluation The class documentation called the parser a recursive descent parser, but sax_parse_internal() is a loop that keeps the open containers on an explicit stack. The comment at the end of an array and of an object said the flag is set to false while the code below it sets it to true. Describe what the code does instead. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Update the discard_number_values comments to the current number path The comments explaining the accept() shortcut in convert_number() and the member documentation still argued in terms of strtoull()/strtoll() and errno, which #5283 replaced with convert_integer(), and pointed at scan_number() instead of convert_number(). They also did not say that scan_number_bulk_contiguous() converts integers itself, so the shortcut is only reached for input without bulk access, with JSON_DIAGNOSTIC_POSITIONS, or when the bulk scanner falls back. Rewrite both comments to describe the digit-count check in front of convert_integer(), keeping the 18-digit bound and the json_sax_acceptor argument. The stale <cstdlib> comment is left for after #5616, which edits that include block. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * List the UTF-8 validators instead of calling the DFA the only one The documentation of decode() called the Hoehrmann DFA the single source of truth for UTF-8 validation. It is used only by the serializer and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text strings). The lexer's scan_string() switch, validate_one_utf8() / valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer) and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges on their own. Replace the sentence with a list of the four validators, what each is used for, and a note that they must accept the same sequences. Sharing code between them was considered and dropped: it would save a few lines in a validator that is entangled with BON8 pushback, and #5677 is editing the BON8 byte path. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix stale doc comments and include lists in the input headers input_adapters.hpp included <memory> and <numeric> for the removed shared_ptr-based adapter design but used neither; it called (std::min) without including <algorithm>. json_sax.hpp used std::numeric_limits without including <limits>. Also corrected comments that no longer matched the code: input_stream_adapter does not skip the input's BOM (the lexer's skip_bom() does), the span_input_adapter comment named the no-longer-existing input_buffer_adapter type, lexer::get_string() does not reset the token, binary_reader's get_number() doc opened with /* instead of /*! (so Doxygen skipped it) and omitted BON8 from its endianness note, and the UBJSON-binary-types note did not mention that BJData 'B' arrays are read as binary. Left out: the lgtm suppression on lexer.hpp's scan_number() (in #5616's hunk) and the "-1 if unknown" wording in json_sax.hpp's start_object/start_array docs (in draft #5267's hunk), per the verdict's conflict list. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the strict-EOF/release_lookahead/error block in parser::parse() json_sax_dom_callback_parser and json_sax_dom_parser branches of parser::parse() ran the same ~25 lines after sax_parse_internal(): the strict-mode EOF check (raising parse_error.101 through the SAX parser), release_lookahead() in non-strict mode, and mapping an errored SAX parser to a discarded result. The two copies had already drifted apart in formatting and in the second copy's "see above" comment. Add a private parse_dom(DomSax&, strict) member that runs this shared sequence once and returns whether the SAX parser did not error; both branches of parse() now only construct their DOM SAX parser, call parse_dom(), and (for the callback parser) map a discarded top-level value to null. sax_parse() is left untouched, since it only runs the EOF check and release_lookahead() when sax_parse_internal() succeeded, unlike parse(), which runs them unconditionally. Behavior-preserving: same operations in the same order for both SAX parser kinds. Verified with unit-class_parser (strict/non-strict, callback and non-callback), unit-deserialization and unit-disabled_exceptions (JSON_NOEXCEPTION), plus a clean make amalgamate / make check-amalgamation diff. Overlaps #5601, which touches the same lines. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5712 item 2 * Share the code point to UTF-8 encoding between the wide-string helpers and the lexer The 1/2/3/4-byte UTF-8 encoding ladder was written out by hand three times: in wide_string_input_helper<..., 4>::fill_buffer() for a UTF-32 code point, in the UTF-16 helper for both a BMP code unit and a valid surrogate pair, and in the lexer's \uXXXX/\uXXXX\uYYYY handling. The copies had drifted: the UTF-32 helper masked the leading bits of each byte (& 0x1Fu, & 0x0Fu, & 0x07u) where the others relied on the shift alone, even though both give the same result for a code point that is already known to be in range. Add detail::encode_utf8(cp, out) in string_utils.hpp, a single encoder that invokes a callable once per output byte, most significant byte first. Use it in the three valid-code-point branches (UTF-32 code points up to U+10FFFF, UTF-16 code units outside the surrogate range, and valid UTF-16 surrogate pairs) and in the lexer's \u handling, where out forwards to add(). The UTF-16 helper's deliberate pass-through of malformed surrogate units and the UTF-32 helper's 0xFF sentinel for code points above U+10FFFF are untouched, since neither reaches the new helper. Behavior-preserving: same bytes in the same order for every valid code point, verified with unit-class_lexer, unit-class_parser, unit-deserialization, unit-wstring and the non-test-data parts of unit-unicode1..5 (ASan/UBSan, C++11/17/20), and an escape-heavy parse microbenchmark that shows no change (about 73 ms either way, median of 3, 1M escape sequences). single_include/ regenerated with make amalgamate; make check-amalgamation leaves a clean tree. Overlaps #5704, which rewrites the wide_string_input_helper specializations touched here. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5712 item 6 --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of the stack for the zero-copy view (#5295).
Summary
lexer::convert_number()converts a float token withstd::from_chars(where available), then Clinger's fast path, then the locale-awarestrtodfallback (#5198, #5597). These steps are now free functions indetail/input/number_parse.hpp. The zero-copy view converts its float tokens with the same code, and therefore gets the same values asjson::parse. The next PR adds a step to this chain.No change in behavior.
Changes
convert_float_fast<FloatType>():std::from_chars, then Clinger's fast path. The fast path is skipped when the mantissa has too many significant digits, or whenFLT_EVAL_METHODindicates excess precision (as before).convert_float_locale_aware<StringType, FloatType>():strtof/strtod/strtoldwith the decimal point of the current locale, retried when the locale changed after the lexer was created (TOCTOU race between lexer construction and locale changes causes float truncation #5198)lexer::convert_number()calls both;lexer.hppno longer includes<clocale>and<cstdlib>.Tests
The existing tests (among them
unit-class_lexer,unit-locale-cpp,unit-testsuites,unit-regression*) pass unchanged under AddressSanitizer and UndefinedBehaviorSanitizer, in C++11 and C++17.Public API
No change: the moved functions are in
detail.Written by Claude Code.
🤖 Generated with Claude Code