Repository navigation
Speed up parsing of contiguous input (numbers, strings, UTF-8) - #5283
Conversation
c724f41 to
84939c1
Compare
4183304 to
c4a3b4d
Compare
Quick review notes – PR #5283
What's missing (should be addressed before merge):
|
c4a3b4d to
509494e
Compare
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
622e91e to
d8cfdc9
Compare
b3c6db6 to
c65313e
Compare
…simdjson world) The number scanner converted its already-validated digit buffer with std::strtoull/std::strtoll/std::strtod. Those pull in locale and errno machinery and dominate number-heavy parsing (strtod runs at ~6 M/s). Replace them with dedicated parsers over the validated buffer: - parse_integer_unsigned / parse_integer_signed: accumulate digits with overflow detection, falling back to the float path on overflow exactly as the strtoull/strtoll round-trip check did. Overflow behavior is unchanged for narrower or wider custom number types. - parse_float_fast: Clinger's exact fast path for `double` (<=19 significant digits, |exp10| <= 22, significand < 2^53), where significand * 10^exp is exact under IEEE round-to-nearest. This is the same fast path used by fast_float/simdjson. It is bit-identical to strtod on this subset and declines (falling back to strtod) otherwise. Only `double` uses it; float and long double keep std::strtof/std::strtold via a templated overload. Measured on representative data (g++ 13, -O3): - integers: DOM parse +11%, SAX +25-34% - floats: DOM parse +37%, SAX +70% (clang: float DOM ~1.9x) No dependencies added; header-only and C++11-clean. Existing parser, lexer, conversion and deserialization unit tests pass unchanged; a 3M-value random-double fuzz matches strtod bit-for-bit. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA Signed-off-by: Niels Lohmann <mail@nlohmann.me>
scan_string() read the input one character at a time through the input adapter and classified every byte with a large switch. For contiguous byte buffers we can instead scan 8 bytes at a time with a SWAR word test that finds the first byte needing individual handling (the closing quote, an escape, a control character, or a non-ASCII UTF-8 byte) and bulk-append the ordinary run in one go. - input adapters expose supports_bulk_scan / bulk_data / bulk_remaining / bulk_skip for provably-contiguous, same-type, 1-byte iterator ranges (raw pointers in every standard; std::string/std::vector/std::array and friends additionally in C++20 via std::contiguous_iterator). - the lexer gains a bulk_scan capability (gated on lazy_token_string so bypassing the per-character capture cannot lose error diagnostics) and a scan_string_bulk() fast path; streaming/wide/user adapters are unchanged and keep the byte-at-a-time scanner. The run contains no newline (all bytes < 0x20 are treated as special), so position bookkeeping stays exact, and error tokens are still reconstructed lazily from the consumed byte range. The SWAR special-byte test is pure uint64_t arithmetic - no intrinsics, no runtime dispatch, C++11-clean. Measured on representative data, pointer input, g++ 13 -O3 (string values discarded by accept() see the largest gains): long ASCII strings: DOM +4.5x, SAX +14x, accept +17x (to ~2 GB/s) short strings: DOM +15%, SAX +62%, accept +85% escape-heavy: DOM +31%, SAX +26%, accept +28% Same-input parity verified: 200k randomized documents (escapes, multibyte UTF-8, surrogate pairs) accept/parse identically via the contiguous SWAR path and the streaming byte path; unit lexer/parser/diagnostic-position/ deserialization/conversions suites pass unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA Signed-off-by: Niels Lohmann <mail@nlohmann.me>
…I text) The SWAR bulk string path stopped at the first non-ASCII byte and handed every multibyte character to the byte-at-a-time scanner, whose per-byte get()/next_byte_in_range()/add() machinery runs at roughly half the speed of validating straight from the buffer. As a result, dense non-ASCII text (CJK, emoji, accented Latin) parsed ~10-15x slower than ASCII. Fold well-formed UTF-8 into the bulk run: scan_string_bulk() now, on a non-ASCII lead byte, validates one sequence with validate_one_utf8() - which mirrors scan_string()'s per-byte switch ranges exactly (rejecting overlong forms, surrogates, and out-of-range code points) - and appends it in place, continuing until the closing quote, an escape, a control byte, or an ill-formed sequence. All error handling still defers to the byte path, so error messages and positions are byte-for-byte unchanged. Because only well-formed content is fast-pathed and every rejection falls through to the existing scanner, behavior is identical; the win is purely throughput. Measured on pointer input (accept, string values discarded): content g++ 13 clang 18 dense CJK 277 -> 648 ~605 MB/s (~2.3x) dense emoji 299 -> 857 ~702 MB/s (~2.6-2.9x) mixed 90% ASCII 246 -> 331 ~334 MB/s (~1.35x) pure ASCII unchanged (~3.2 / 4.1 GB/s) Verified: 2,000,000 randomized documents built from arbitrary bytes (overlong, surrogate, truncated, out-of-range sequences) accept/reject and parse identically via the contiguous path and the streaming byte path; lexer/parser/diagnostic-position/deserialization/conversions suites pass unchanged. Pure C++11, no intrinsics. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA Signed-off-by: Niels Lohmann <mail@nlohmann.me>
…ths in C++11) json::parse(std::string) - the most common entry point - did not benefit from the contiguous fast paths (bulk string scanning, UTF-8 bulk validation, memcpy for binary formats) in C++11..17: std::string::iterator is a library wrapper, not a raw pointer, and pre-C++20 there is no portable way to prove it contiguous, so supports_bulk_scan was false. Only raw pointers, string literals, and C-arrays (and, in C++20, anything modelling std::contiguous_iterator) took the fast path. Detect contiguous single-byte containers (std::string, std::vector<char>, std::vector<std::uint8_t>, std::string_view, ...) via is_contiguous_byte_ container and route them through an iterator_input_adapter built from data()/data()+size(). The generic iterator-based container overload is constrained to exclude these, so the two overloads are disjoint and there is no ambiguity (a plain competing overload loses to the greedy forwarding-reference container overload on reference binding, and a factory partial-specialization is ambiguous - both were tried and rejected). The pointer keeps the container's own element type, so char_type - and therefore all parsing behavior - is byte-for-byte identical to the iterator path (const char* for std::string, const std::uint8_t* for std::vector<std::uint8_t>); only the raw pointer additionally turns on the fast paths. Lifetimes are unchanged: the container outlives the adapter for the full parse expression, exactly as the iterators it replaces did. Measured, C++11, json::parse/accept(std::string), g++ 13: long ASCII strings: accept 201 -> 3200 MB/s (~16x), parse 174 -> 1444 dense CJK: accept 263 -> 697 MB/s (~2.6x) short strings: accept 163 -> 243 MB/s (~1.5x) Verified: char_type preserved for std::string (char) and std::vector<std::uint8_t> (uint8_t); CBOR/MsgPack round-trips from std::vector<std::uint8_t> unchanged; 1,000,000 randomized documents accept and parse identically via std::string and via std::istream; deserialization/user-defined-input/parser/lexer/conversions/diagnostic- position suites pass (20,480 assertions); warning-clean on g++ and clang in C++11/17/20. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA Signed-off-by: Niels Lohmann <mail@nlohmann.me>
is_contiguous_byte_container accepted any type with a data() returning a
pointer to a single-byte integral plus a size(). That is duck typing: the two
members say nothing about size() counting the units data() points at.
A type where it does not - fixed-size records, say - was routed to the
pointer-based adapter and parsed as [data(), data() + size()) bytes, silently
truncating input the iterator-based adapter had read in full:
struct record_buffer {
using value_type = std::array<char, 4>;
std::string bytes;
const char* data() const; // raw bytes
std::size_t size() const; // in records
const char* begin() const; const char* end() const;
};
json::parse(record_buffer{"[1,2,3,4,5]"}); // parse error at column 3
Requiring the container's own value_type to be that same element type ties the
two together. Every contiguous standard container satisfies it, so std::string,
std::vector<char>, std::array<char, N> and std::string_view keep the fast path;
anything else falls back to the iterator-based adapter, which is always correct.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
simdutf.h rejects anything below C++17 with an #error, so defining JSON_USE_SIMDUTF in a C++11 or C++14 translation unit did not fail with a message about simdutf being unavailable - it failed to compile at all, taking the library's C++11 support with it. Nothing caught this because no build ever compiled that path. Gate the include and both uses on JSON_HAS_CPP_17, the same way number_parse.hpp gates std::from_chars. Below C++17 the macro now has no effect and the scalar validator runs; it accepts and rejects exactly the same input, so the macro is safe to set project-wide even when some translation units use an older standard. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T Signed-off-by: Niels Lohmann <mail@nlohmann.me>
JSON_USE_SIMDUTF was documented and shipped but never compiled by anything in the repository, so nothing held the backend to the behavior the docs promise. Add JSON_TestSimdutf (OFF by default), which fetches simdutf and defines JSON_USE_SIMDUTF for every test target, and a ci_test_simdutf target that runs the whole suite in that configuration. Because simdutf needs C++17, the suite is built at C++11 as well, so one job covers both the scalar fallback with the macro defined and simdutf itself. The dependency hangs off test_main, whose usage requirements every test target inherits. The library target and the installed CMake package are deliberately untouched: making nlohmann_json link simdutf would put a find_dependency() in the exported package, which is a separate decision. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T Signed-off-by: Niels Lohmann <mail@nlohmann.me>
simdutf needs C++17: without it the dependency does not even compile, and with a C++17 compiler but no C++17-or-later standard under test it builds and then goes unused. Either way the option silently did nothing useful, or broke the configure step outright. Resolve the tested standards first, then check them: when none of them can reach simdutf, skip the dependency and say so, naming which of the two reasons applies and how to fix it. The tests then run against the scalar validator, which is what would have happened anyway. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The record_buffer test type declares data() and size() so the is_contiguous_byte_container trait can see both and still reject the type on its value_type. data() was never called, so clang's -Wunneeded-member-function (under -Weverything -Werror) failed the C++20 build. Assert that data() points at the underlying bytes: it ODR-uses the member and documents the property the type is meant to demonstrate. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG Signed-off-by: Niels Lohmann <mail@nlohmann.me>
parse_float_fast() (Clinger) needs a significand below 2^53, so it always declines once the mantissa has 17 or more significant digits. convert_number() called it unconditionally, so those numbers were walked an extra time before strtod had to run anyway. On streaming input, where scanning is byte-at-a-time and there is no compensating win, that made canada.json about 6% slower than develop. Derive the significant-digit count from token_buffer indices - the digits are not scanned again - and skip the call when it is guaranteed to decline. Both scanners pass the offset where the mantissa ends; the count only has to be corrected for a leading "0", which the JSON grammar admits nowhere else. The integer path returns before the check, so integer-heavy input is unaffected. Values are unchanged: this only avoids an attempt that would have failed. Verified bit-exact against develop over every number in canada.json, floats.json, signed_ints.json, unsigned_ints.json, small_signed_ints.json, citm_catalog.json and twitter.json, for both the contiguous and the streaming scanner. parse, streaming develop before after canada.json 19.4ms 20.5ms 19.3ms floats.json 135.9ms 131.8ms 128.0ms parse, contiguous develop before after canada.json 15.5ms 12.9ms 11.7ms floats.json 98.6ms 69.8ms 66.7ms Signed-off-by: Niels Lohmann <mail@nlohmann.me>
c65313e to
621bb04
Compare
lexer::skip_whitespace() called get() for every whitespace byte, and get() checks the (almost always false, once past the first character) next_unget flag on every call. skip_whitespace() now reads its first character with get() (needed to honor a pending unget() left over from finishing the previous token, e.g. scan_number() always ungets the character that terminated the number) and every further whitespace character with a new get_ignoring_pending_unget() variant that skips that branch, since nothing in the loop calls unget(). This is a narrower fix than the full contiguous-buffer bulk-skip suggested in the issue (scan a run of whitespace directly in the adapter's buffer and update position counters once per run). That approach depends on bulk-scan adapter infrastructure (supports_bulk_scan/bulk_data()/bulk_skip()) introduced by the open, unmerged parser-performance PR #5283, which this change intentionally does not depend on or replicate. Building new bulk-scan adapter infrastructure from scratch was judged out of scope/riskier than warranted here, so this change is limited to the safe, always-correct improvement of removing redundant per-character bookkeeping from the existing byte-at-a-time loop; full bulk-skipping is left as future work once #5283 (or equivalent adapter support) lands. Line/column/byte-offset bookkeeping is untouched and verified bit-for-bit identical before and after this change, including for pretty-printed (dump(4)) input with embedded newlines. Fixes #5412 Stacked on top of the PR for #5411 (branch issue-5411-lexer-skip-conversion). Signed-off-by: Niels Lohmann <mail@nlohmann.me>
lexer::skip_whitespace() called get() for every whitespace byte, and get() checks the (almost always false, once past the first character) next_unget flag on every call. skip_whitespace() now reads its first character with get() (needed to honor a pending unget() left over from finishing the previous token, e.g. scan_number() always ungets the character that terminated the number) and every further whitespace character with a new get_ignoring_pending_unget() variant that skips that branch, since nothing in the loop calls unget(). This is a narrower fix than the full contiguous-buffer bulk-skip suggested in the issue (scan a run of whitespace directly in the adapter's buffer and update position counters once per run). That approach depends on bulk-scan adapter infrastructure (supports_bulk_scan/bulk_data()/bulk_skip()) introduced by the open, unmerged parser-performance PR #5283, which this change intentionally does not depend on or replicate. Building new bulk-scan adapter infrastructure from scratch was judged out of scope/riskier than warranted here, so this change is limited to the safe, always-correct improvement of removing redundant per-character bookkeeping from the existing byte-at-a-time loop; full bulk-skipping is left as future work once #5283 (or equivalent adapter support) lands. Line/column/byte-offset bookkeeping is untouched and verified bit-for-bit identical before and after this change, including for pretty-printed (dump(4)) input with embedded newlines. Fixes #5412 Stacked on top of the PR for #5411 (branch issue-5411-lexer-skip-conversion). Signed-off-by: Niels Lohmann <mail@nlohmann.me>
lexer::skip_whitespace() called get() for every whitespace byte, and get() checks the (almost always false, once past the first character) next_unget flag on every call. skip_whitespace() now reads its first character with get() (needed to honor a pending unget() left over from finishing the previous token, e.g. scan_number() always ungets the character that terminated the number) and every further whitespace character with a new get_ignoring_pending_unget() variant that skips that branch, since nothing in the loop calls unget(). This is a narrower fix than the full contiguous-buffer bulk-skip suggested in the issue (scan a run of whitespace directly in the adapter's buffer and update position counters once per run). That approach depends on bulk-scan adapter infrastructure (supports_bulk_scan/bulk_data()/bulk_skip()) introduced by the open, unmerged parser-performance PR #5283, which this change intentionally does not depend on or replicate. Building new bulk-scan adapter infrastructure from scratch was judged out of scope/riskier than warranted here, so this change is limited to the safe, always-correct improvement of removing redundant per-character bookkeeping from the existing byte-at-a-time loop; full bulk-skipping is left as future work once #5283 (or equivalent adapter support) lands. Line/column/byte-offset bookkeeping is untouched and verified bit-for-bit identical before and after this change, including for pretty-printed (dump(4)) input with embedded newlines. Fixes #5412 Stacked on top of the PR for #5411 (branch issue-5411-lexer-skip-conversion). Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Speed up whitespace skipping in the lexer lexer::skip_whitespace() called get() for every whitespace byte, and get() checks the (almost always false, once past the first character) next_unget flag on every call. skip_whitespace() now reads its first character with get() (needed to honor a pending unget() left over from finishing the previous token, e.g. scan_number() always ungets the character that terminated the number) and every further whitespace character with a new get_ignoring_pending_unget() variant that skips that branch, since nothing in the loop calls unget(). This is a narrower fix than the full contiguous-buffer bulk-skip suggested in the issue (scan a run of whitespace directly in the adapter's buffer and update position counters once per run). That approach depends on bulk-scan adapter infrastructure (supports_bulk_scan/bulk_data()/bulk_skip()) introduced by the open, unmerged parser-performance PR #5283, which this change intentionally does not depend on or replicate. Building new bulk-scan adapter infrastructure from scratch was judged out of scope/riskier than warranted here, so this change is limited to the safe, always-correct improvement of removing redundant per-character bookkeeping from the existing byte-at-a-time loop; full bulk-skipping is left as future work once #5283 (or equivalent adapter support) lands. Line/column/byte-offset bookkeeping is untouched and verified bit-for-bit identical before and after this change, including for pretty-printed (dump(4)) input with embedded newlines. Fixes #5412 Stacked on top of the PR for #5411 (branch issue-5411-lexer-skip-conversion). Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix codegen regression in skip_whitespace() from #5490 Benchmarking found the get()/get_ignoring_pending_unget() split in skip_whitespace() made long whitespace runs (e.g. indentation in pretty-printed JSON) 1.75x-3.2x SLOWER instead of faster, reproducible with both Apple Clang and GCC. Root cause: rewriting the loop from a plain do-while into an initial get() followed by a while-loop defeated the compiler's ability to keep the input adapter's read/end pointers in registers across iterations; both compilers instead reloaded them from memory on every character. The function split itself was not the problem (it still fully inlines); the loop's control-flow shape was. The fix keeps the same two-function structure but restores a do-while shape (guarded by an if for the "first char not whitespace" case), which lets both compilers hoist the pointers back into registers, matching or beating pre-#5490 performance. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Share the position-counter bump between get() and get_ignoring_pending_unget() Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Extract current_is_whitespace() to deduplicate skip_whitespace()'s two whitespace checks Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Use a raw string literal for the multi-line error-position test input Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix clang-tidy raw-string-literal finding and guard a new test against JSON_NOEXCEPTION The issue #5412 whitespace-skipping test added a check_error() helper that relies on catching json::parse_error to verify the exception message; under JSON_NOEXCEPTION, JSON_THROW aborts instead of throwing, which crashed ci_test_noexceptions (and cascaded into the other ci_cmake_options jobs). Guard the whole section with #if !defined(JSON_NOEXCEPTION), matching the existing pattern used by sibling tests in this file. Also switch one escaped string literal to a raw string literal to satisfy clang-tidy's modernize-raw-string-literal check. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
…ormance-research-4hkdf8 Resolves the number-parsing conflict between the discard_number_values fast path (an early return for accept()-only calls that skips value conversion for short digit runs) and the contiguous number scanner's convert_number()/convert_integer() split: the discard check now runs as the first step of convert_number(), so both the byte-at-a-time scanner and the bulk contiguous path benefit from it. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add missing headers to BUILD.bazel and make its generator reproduce it The "json" cc_library did not list three headers that the library includes: - detail/meta/logic.hpp (added in #5016, included by from_json.hpp) - detail/input/number_parse.hpp (added in #5283, included by lexer.hpp) - detail/input/string_scan.hpp (added in #5283, included by lexer.hpp and serializer.hpp) Bazel's sandbox only exposes declared headers, so any target depending on @nlohmann_json//:json and including <nlohmann/json.hpp> failed with "'nlohmann/detail/meta/logic.hpp' file not found". The file could not simply be regenerated, because the generator behind "make BUILD.bazel" was stale: it wrote only the "json" cc_library and dropped the load() statements, the license block, and the "singleheader-json" target that were added by hand in #4584. The generator now emits the complete file, so its output differs from the previous BUILD.bazel only by the three headers. It also resolves the glob against the project root instead of the working directory and sorts the list explicitly. "make BUILD.bazel" is now phony: in a fresh checkout, BUILD.bazel is not older than the headers, so make considered it up to date, and a removed header would never trigger a rebuild. "make check-amalgamation" also checks that BUILD.bazel is up to date. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Check in CI that BUILD.bazel is up to date The "Check amalgamation" workflow now also regenerates BUILD.bazel, so a pull request that adds, renames, or removes a header without updating the Bazel header list fails, and the attached amalgamation.patch contains the fix. The failure comment and the contribution guidelines mention the new check, and the comment now links to the existing "Amalgamate the source code" section instead of the "Files to change" anchor that was removed in #4560. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
…5731) * Share the diagnostic-position setter of the DOM SAX parsers json_sax_dom_parser and json_sax_dom_callback_parser each had a private copy of handle_diagnostic_positions_for_json_value(), identical except for comments. Move the body into one static member function, detail::diagnostic_positions::set_from_lexer(value, lexer), which both classes call with their lexer pointer. basic_json befriends the new struct (only when JSON_DIAGNOSTIC_POSITIONS is enabled), as the position members are private. The discarded case is reached through the callback parser, so the LCOV_EXCL markers that only the dom parser's copy had are gone. The NOLINT on the unreachable default case loses the stray "-warnings-as-errors", which is not a check name. The start-position setup in start_object()/start_array() is left alone, as #5706 is editing the callback parser's versions. Behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Correct the parser comments on recursion and skip_to_state_evaluation The class documentation called the parser a recursive descent parser, but sax_parse_internal() is a loop that keeps the open containers on an explicit stack. The comment at the end of an array and of an object said the flag is set to false while the code below it sets it to true. Describe what the code does instead. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Update the discard_number_values comments to the current number path The comments explaining the accept() shortcut in convert_number() and the member documentation still argued in terms of strtoull()/strtoll() and errno, which #5283 replaced with convert_integer(), and pointed at scan_number() instead of convert_number(). They also did not say that scan_number_bulk_contiguous() converts integers itself, so the shortcut is only reached for input without bulk access, with JSON_DIAGNOSTIC_POSITIONS, or when the bulk scanner falls back. Rewrite both comments to describe the digit-count check in front of convert_integer(), keeping the 18-digit bound and the json_sax_acceptor argument. The stale <cstdlib> comment is left for after #5616, which edits that include block. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * List the UTF-8 validators instead of calling the DFA the only one The documentation of decode() called the Hoehrmann DFA the single source of truth for UTF-8 validation. It is used only by the serializer and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text strings). The lexer's scan_string() switch, validate_one_utf8() / valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer) and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges on their own. Replace the sentence with a list of the four validators, what each is used for, and a note that they must accept the same sequences. Sharing code between them was considered and dropped: it would save a few lines in a validator that is entangled with BON8 pushback, and #5677 is editing the BON8 byte path. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix stale doc comments and include lists in the input headers input_adapters.hpp included <memory> and <numeric> for the removed shared_ptr-based adapter design but used neither; it called (std::min) without including <algorithm>. json_sax.hpp used std::numeric_limits without including <limits>. Also corrected comments that no longer matched the code: input_stream_adapter does not skip the input's BOM (the lexer's skip_bom() does), the span_input_adapter comment named the no-longer-existing input_buffer_adapter type, lexer::get_string() does not reset the token, binary_reader's get_number() doc opened with /* instead of /*! (so Doxygen skipped it) and omitted BON8 from its endianness note, and the UBJSON-binary-types note did not mention that BJData 'B' arrays are read as binary. Left out: the lgtm suppression on lexer.hpp's scan_number() (in #5616's hunk) and the "-1 if unknown" wording in json_sax.hpp's start_object/start_array docs (in draft #5267's hunk), per the verdict's conflict list. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the strict-EOF/release_lookahead/error block in parser::parse() json_sax_dom_callback_parser and json_sax_dom_parser branches of parser::parse() ran the same ~25 lines after sax_parse_internal(): the strict-mode EOF check (raising parse_error.101 through the SAX parser), release_lookahead() in non-strict mode, and mapping an errored SAX parser to a discarded result. The two copies had already drifted apart in formatting and in the second copy's "see above" comment. Add a private parse_dom(DomSax&, strict) member that runs this shared sequence once and returns whether the SAX parser did not error; both branches of parse() now only construct their DOM SAX parser, call parse_dom(), and (for the callback parser) map a discarded top-level value to null. sax_parse() is left untouched, since it only runs the EOF check and release_lookahead() when sax_parse_internal() succeeded, unlike parse(), which runs them unconditionally. Behavior-preserving: same operations in the same order for both SAX parser kinds. Verified with unit-class_parser (strict/non-strict, callback and non-callback), unit-deserialization and unit-disabled_exceptions (JSON_NOEXCEPTION), plus a clean make amalgamate / make check-amalgamation diff. Overlaps #5601, which touches the same lines. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5712 item 2 * Share the code point to UTF-8 encoding between the wide-string helpers and the lexer The 1/2/3/4-byte UTF-8 encoding ladder was written out by hand three times: in wide_string_input_helper<..., 4>::fill_buffer() for a UTF-32 code point, in the UTF-16 helper for both a BMP code unit and a valid surrogate pair, and in the lexer's \uXXXX/\uXXXX\uYYYY handling. The copies had drifted: the UTF-32 helper masked the leading bits of each byte (& 0x1Fu, & 0x0Fu, & 0x07u) where the others relied on the shift alone, even though both give the same result for a code point that is already known to be in range. Add detail::encode_utf8(cp, out) in string_utils.hpp, a single encoder that invokes a callable once per output byte, most significant byte first. Use it in the three valid-code-point branches (UTF-32 code points up to U+10FFFF, UTF-16 code units outside the surrogate range, and valid UTF-16 surrogate pairs) and in the lexer's \u handling, where out forwards to add(). The UTF-16 helper's deliberate pass-through of malformed surrogate units and the UTF-32 helper's 0xFF sentinel for code points above U+10FFFF are untouched, since neither reaches the new helper. Behavior-preserving: same bytes in the same order for every valid code point, verified with unit-class_lexer, unit-class_parser, unit-deserialization, unit-wstring and the non-test-data parts of unit-unicode1..5 (ASan/UBSan, C++11/17/20), and an escape-heavy parse microbenchmark that shows no change (about 73 ms either way, median of 3, 1M escape sequences). single_include/ regenerated with make amalgamate; make check-amalgamation leaves a clean tree. Overlaps #5704, which rewrites the wide_string_input_helper specializations touched here. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5712 item 6 --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Review and extend the documentation, and check it in CI A review of all documentation pages found factual errors, dead links, missing cross-references, and gaps in examples. This fixes them and adds checks so the same problems are caught automatically. Fixes: - wrong signatures and version histories (operator!= C++20 member, binary() subtype type, get<PointerType>(), JSON_NO_THREAD_LOCAL, ...) - stale descriptions (number parsing since #5283, UBJSON table, SAX example that no longer compiled, tsl::ordered_map advice) - dead internal and external links; repology.org badges (the domain is suspended) replaced by badges that query the registries directly - deprecation notes link the migration guide; the guide itself fixed Additions: - "See also" sections, cross-references, 25 runnable examples, 12 Mermaid diagrams, new API pages for json_pointer::operator<=> and byte_container_with_subtype::operator==/!= - landing page, guides for untrusted input and performance - "unreleased" badge after versions newer than the latest release Checks: - strict documentation build (broken links/anchors fail it); CI and the publish workflow fetch the full history the build needs - weekly external link check, Mermaid syntax check in CI - check_structure.py: example titles, heading levels, alt texts, header links, docset index coverage; its unused-example check works again - all examples produce the same output on every platform Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Keep the customer links that could not be fixed A dead link on the customers page is still the evidence of where the use of the library was documented. Keep the original URLs of the entries without a working replacement (Marne, Cisco Webex Desk Camera, Philips Hue, CyberArk) and exclude exactly these URLs from the link check. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Correct the duplicate-key recipe's claim about SAX positions The SAX interface's key() receives no position either; only parse_error() does. Also note that the recipe does not report the path to the repeated key (see discussion #5085). Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Say the library is available as a single header and mention json_fwd.hpp Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Correct documentation errors found while hunting for bugs - patch/patch_inplace: list the JSON pointer errors parse_error.106-109 and out_of_range.402/404, and quote the actual parse_error.105 message. - unflatten: list parse_error.106/107/108 and out_of_range.404. - to_bson: list out_of_range.415 (binary subtype above 255) and note that 412 and 415 are new in 3.13.0. - to_string: state that string_t must be convertible to std::string, also in the StringType requirements table. - JSON Lines: a `while (input >> j)` loop also throws after the last value for concatenated JSON values; show a loop that works for both. - BON8: a string gets 0xFF only if nothing follows it in the message; a string at the end of an array or object is ended by 0xFE. - custom_string_type.hpp: add operator+=(char), which the "Always required" list asks for (json_pointer::to_string, flatten, unflatten, and diff did not compile), and an ADL int_to_string for diff and items. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Cache the release headers with functools.lru_cache Codacy (Pylint) flagged the mutable default argument that header() used as its cache. functools.lru_cache keeps the same memoization without it. The script's output is unchanged. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
What & why
The parser reads input one character at a time through the input adapter and converts numbers with
strtod/strtoull. For the common case of contiguous byte input (std::string,std::vector<char>, string literals,const char*ranges), the per-characterget()/add()overhead and the locale/errno-heavy number conversion dominate. This PR adds fast paths for that case, borrowing ideas from the simdjson / fast_float ecosystem, while keeping the library header-only with a C++11 fallback on every path. Streaming inputs (files,std::istream, wide strings, user adapters) are untouched and keep the byte-at-a-time scanner. Every fast path defers to the byte-path scanner on anything malformed, so parse errors keep the same message and the samebyteoffset; the one deliberate exception is the reported column after a newline, described under Public API impact below.Stack
The changes (each independent and reversible)
strtoull/strtoll/strtodover the validated digit buffer with dedicated parsers: hand-rolled integer parsing with overflow→float fallback (matching the old semantics exactly), and Clinger's exact fast path fordouble, bit-identical tostrtodon its subset and declining otherwise.float/long doublekeepstrtof/strtold.uint64_tword test to bulk-append ordinary characters. Pure C++11, no intrinsics.std::string& friends through a pointer-based adapter so 1–3 apply tojson::parse(std::string)in C++11–17 (previously onlyconst char*/literals, and C++20 contiguous iterators, qualified). The pointer keeps the container's element type, sochar_type— and all behavior, including binary formats fromstd::vector— is unchanged. Detection requires the container's ownvalue_typeto be the single-byte elementdata()points at, so a type whosesize()counts something else keeps the iterator-based adapter instead of being read as[data(), data() + size())bytes.get()/add()), materializingtoken_bufferin one copy with the locale decimal point substituted, then reusing the sharedconvert_number()tail.JSON_USE_SIMDUTF, off by default) — since simdutf ships a.cppand uses runtime dispatch, it is not vendored; when defined, bulk UTF-8 validation routes throughsimdutf::validate_utf8(the user supplies/links simdutf), otherwise the C++11 scalar validator is used. The library stays header-only either way. simdutf requires C++17 — its header rejects older standards with an#error— so the backend is gated onJSON_HAS_CPP_17and the macro is inert below that, which keeps it safe to set project-wide in a mixed-standard build.detail/input/number_parse.hpp(302 lines) anddetail/input/string_scan.hpp(241 lines), as free functions operating on raw bytes, solexer.hppkeeps only the state machine and the adapter-facing glue.lexer.hppitself grows (1760 → 2025 lines) with the bulk scanners and the number fast path; the only code moved out of it is the ~30-linestrtoull/strtoll/strtodblock replaced in item 1.std::counted_iterator+std::default_sentinel_treaches them too (per @gregmarr's review, matching Extend memcpy fast path to sized sentinels (e.g. std::counted_iterator) #5268).Measured gains
Cumulative,
json::parse/accept(std::string), C++11, g++ 13-O3, vs the v3.12.0 baseline (representative synthetic data):parseacceptaccept()/SAX/validation see the largest gains (not allocation-bound); full-DOMparse()gains on object-heavy input are capped by DOM allocation, which is a separate concern. WithJSON_USE_SIMDUTF, dense-UTF-8 validation has further headroom (~4.7 GB/s ceiling measured with an SSE validator; ~7× the scalar path).Public API impact
No breaking changes to the documented public API.
parse(),accept(),sax_parse(), and thefrom_*()signatures are unchanged.Two internal
nlohmann::detailchanges are worth recording:detail::input_adapter(container)now returns a pointer-based adapter for contiguous byte containers, sodetail::string_input_adapter_typechanges fromiterator_input_adapter<std::string::iterator>toiterator_input_adapter<const char*>. Source-compatible (these types are normally deduced), but the mangled names change, which matters to anyone shipping a precompiled wrapper that names them.iterator_input_adaptergainssupports_bulk_scan/bulk_data()/bulk_remaining()/bulk_skip(), andsupports_seekis relaxed from "same-type iterator/sentinel" to "distance computable in O(1)".One deliberate behavior change: error column after a newline
scan_number()reads the character that terminates a number and then ungets it, so the reported column is the one reached after the number's last character. When that terminator is a newline,get()has already clearedchars_read_current_lineandunget()could only restorelines_read, leaving the column at 0 — a number followed by a newline reported a different position than the same number followed by a space.unget()now remembers the column the newline was read at, so both terminators report the same position:Only the column changes — message text,
byteoffset, and accept/reject are identical, and the contiguous and streaming routes agree. 22 of 481k differential documents are affected, all of this shape.Verification
Differential testing against the merge base across all routes (
std::string/std::istream/std::vector/accept), C++11 and C++17, clang/libc++ and g++/libstdc++, in theCandde_DE/fr_FRlocales:byteoffsets, and message text; zero disagreement between the contiguous and streaming routes across 298,333 error documents.strtod, both withstd::from_charsactive under libstdc++ and under-ffast-math.JSON_DIAGNOSTIC_POSITIONSstart/end offsets: identical over 60k documents.%.17gdoubles; escapes) parse and accept/reject identically; theJSON_USE_SIMDUTFbuild matches too.make check-amalgamationclean.JSON_TestSimdutf=ON, the whole suite (162 tests, built at C++11 and C++17) passes against the simdutf backend, and its results are byte-identical to the scalar validator over the differential corpus.Notes for reviewers / remaining work
tests/cases cover the number fast path (including an exhaustive grammar-parity sweep over the number alphabet), the string and UTF-8 bulk scanners (exhaustive contiguous-vs-streaming parity, special bytes at every offset of the SWAR stride, and the boundaries of every accepted UTF-8 range), contiguous-container detection, the counted-iterator bulk paths (including counts that end before the underlying buffer), and contiguous-vs-streaming error-position parity.JSON_TestSimdutf(off by default) builds the unit tests against simdutf, fetched withFetchContent;ci_test_simdutfruns the suite in that configuration and is part of theci_cmake_optionsmatrix. It warns and falls back to the scalar validator when no tested standard can reach simdutf. The library target and the installed CMake package are deliberately untouched: makingnlohmann_jsonlinksimdutf::simdutfwould add afind_dependency(simdutf)to the exported package, which is a separate decision.JSON_USE_SIMDUTFdoc page says "Added in version 3.13.0", matching the other pages documenting unreleased features — adjust if this lands in a different release.JSON_USE_SIMDUTFmacro page, nav entry, and supported-macros overview)make amalgamate.Generated by Claude Code