Skip to content

Scan json_view strings with SIMD and index large objects - #5628

Open
nlohmann wants to merge 22 commits into
json-view/13-view-dumpfrom
json-view/16-view-simd
Open

nlohmann wants to merge 22 commits into
json-view/13-view-dumpfrom
json-view/16-view-simd

Conversation

@nlohmann

@nlohmann nlohmann commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Part of the stack for the zero-copy view (#5295). This PR combines #5628 and #5629, which were reviewed separately before, with the parser-side part of #5765's x86-64 Linux work folded in: SIMD scanning of strings and an object index, plus the run-time SSSE3 dispatch, a vector-first string scan on x86-64, and a store-forwarding fix that closed the remaining gap between GCC and Clang.

Summary

  • Long runs of string bytes are scanned 16 bytes at a time with NEON (AArch64, GCC and Clang) and SSE2 (x86-64). Both belong to the baseline instruction sets.
  • Keys and values: keys keep 16 table checks before the vector loop, because their lengths repeat from record to record. String values get 8, because their lengths vary more.
  • Non-ASCII text is validated 16 bytes at a time with simdjson's "lookup4" check (Keiser and Lemire, 2021): with NEON on AArch64, and on x86-64 with SSSE3.
    • SSSE3 on every x86-64 CPU that has it. SSSE3 is not part of baseline x86-64, so instead of requiring it at compile time, the check is compiled for SSSE3 with a function attribute (GCC 4.9 and later, Clang; MSVC compiles the intrinsics anyway) and used where CPUID reports SSSE3, which all x86-64 CPUs since about 2011 do. The answer is cached in a std::atomic<int> initialized at compile time, so there is no guard of a local static and no global constructor (-Wglobal-constructors); the definitions do not depend on compiler flags, so different translation units cannot violate the ODR. JSON_VIEW_USE_SSSE3 now only skips the CPU check; the same input is accepted either way.
  • Opt-out: JSON_VIEW_NO_SIMD selects the portable code.
  • Strings vector-first on x86-64: with SSE2, string runs are checked 16 bytes at a time from their first byte instead of one byte at a time for the first 8-16 bytes; one compare finds the end of most keys and short values, which is faster than a branch per byte on x86-64. AArch64 keeps the byte-wise steps: there a NEON mask costs more and the branches predict well, and vector-first was 20% slower on Apple M1.
  • No stall when entering or leaving an object or array: open() now stores the parent's frame field by field instead of building it on the stack and copying it; the copy was read back by loads wider than its stores, which waited until the stores retired (store forwarding fails). This was 18% of the parse time of citm_catalog on x86-64; Apple M1 forwards such stores and does not change.
  • Credit: simdjson is credited in simd.hpp's SPDX block, the README, and license.md.
  • Object index: lookups in objects are linear, as for ordered_json. Objects with 128 members or more now get a hash table after parsing:
    • open addressing;
    • of duplicate keys, the first is kept, as for the linear search;
    • operator[], at(), find(), contains(), count(), value(), and JSON pointers take constant time on average in such objects.
    • The idea of switching to a hash table for large objects comes from Boost.JSON.
    • The parser notes a large object when it closes it, out of line, so the parse loop only has a call for it; done inline, it made canada.json 37% slower. The object node keeps the number of its table.
    • The node index section of the architecture page describes the tables and what extra holds for objects.
  • Docs: JSON_VIEW_USE_SSSE3 describes the run-time check. read() notes that reusing a document avoids the page faults of a fresh node index, about 40% of the parse time of a 55 MB document on x86-64 Linux.

Performance

json_view

Measured on the regrouped stack

Measured on the regrouped stack (µs, best of 3 interleaved rounds of 15 runs; files from nativejson-benchmark). Apple M1 Max with Apple clang at -O2; x86-64 on a KVM Haswell VPS, pinned to one core, GCC 13 / Clang 18 at -O2 (the VPS is noisy, about ±5–10%).

json_document::parse, previous PR → this PR:

Apple M1 x86-64 GCC x86-64 Clang
twitter 245 → 187 (−24%) 560 → 333 (−41%) 529 → 357 (−33%)
citm_catalog 389 → 374 (−4%) 967 → 786 (−19%) 916 → 816 (−11%)
canada 1019 → 1027 (+1%) 1972 → 1827 (−7%) 2282 → 2028 (−11%)

Earlier measurements (on the old stack)

json_document::parse, best of 7 runs in separate processes (Apple M1 Max), before -> after:

file change
poet.json (CJK text) -72%
random.json -25%
twitter.json -22%
gsoc-2018.json -20%
semanticscholar -19%
github_events -11%
apache_builds -9.5%
update-center -7%
citm_catalog / canada -6% / -5%
lottie / tree-pretty / twitterescaped / instruments +4% / +2.5% / +1.5% / +1.4%

Object index: looking up every key of an object with 10,000 members, 59.8 ms -> 0.16 ms. Parsing (json_document::parse, best of 7, separate processes) stays within 1% for most files; canada +5%, mesh.pretty +3%, citm +3%.

x86-64 (KVM Haswell, from #5765's SSSE3 dispatch, vector-first scan, and the open() stall fix): twitter 706 -> 400 us (GCC) and 631 -> 367 us (Clang); citm_catalog about -25%.

Core library (json::parse / dump)

Unaffected. All changes are inside include/nlohmann/detail/view/ (scan.hpp, simd.hpp, new object_index.hpp, builder.hpp, document_data.hpp, lookup.hpp) and the view's own macro_unscope.hpp; no file the core lexer or serializer uses is touched.

Tests

  • Every two-byte sequence, and three- and four-byte sequences with continuation bytes at the edges of their ranges, each placed at every offset around the vector blocks of keys and values, and also cut short.
  • Long runs of text with a damaged byte, all compared with json::accept and json::parse.
  • Checked locally with NEON, the portable code, SSE2, and SSSE3 (the x86-64 ones under Rosetta); CMake builds the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with JSON_VIEW_USE_SSSE3 -mssse3.
  • Old compilers in Docker (linux/amd64): GCC 4.9, 5, 7, 9, 12 and Clang 5, 10 build it, detect SSSE3 at run time, and agree with json::accept on 200,000 random UTF-8 strings each, including ill-formed ones.
  • Objects with 127, 128, 129, and 10,000 members; escaped, empty, duplicate, and missing keys, and comparisons; nested large objects, and documents reused with read().
  • Fuzzing: 846,567 libFuzzer runs of fuzzer-parse_json_view with ASan and UBSan, with the run-time SSSE3 dispatch active.

Public API

No breaking change:

  • Two new configuration macros, documented under Macros: JSON_VIEW_NO_SIMD (portable scanning) and JSON_VIEW_USE_SSSE3 (skip the CPU check; without it, x86-64 builds now use SSSE3 when the CPU has it, and the same inputs are accepted or rejected either way).
  • Faster, otherwise unchanged lookups in large objects; no new members.

Written by Claude Code.

🤖 Generated with Claude Code

@nlohmann
nlohmann added this pull request to stack #5636 September 29, 2026 14:20
@nlohmann nlohmann changed the title json view/16 view simd Scan the strings of json_view with NEON and SSE2 Sep 29, 2026
@github-actions

Copy link
Copy Markdown

🔴 Amalgamation check failed! 🔴

The source code has not been amalgamated and/or formatted correctly, or BUILD.bazel is out of date.

@nlohmann
nlohmann removed this pull request from stack #5636 September 30, 2026 13:19
@nlohmann
nlohmann force-pushed the json-view/16-view-simd branch from 50f0a1b to d33586c Compare September 30, 2026 13:20
@nlohmann
nlohmann added this pull request to stack #5739 September 30, 2026 13:21
@nlohmann
nlohmann force-pushed the json-view/16-view-simd branch 3 times, most recently from da1c822 to d4ac1b7 Compare September 30, 2026 18:06
@nlohmann nlohmann added the review needed It would be great if someone could review the proposed changes. label Sep 30, 2026
@nlohmann
nlohmann marked this pull request as ready for review September 30, 2026 18:18
@nlohmann
nlohmann force-pushed the json-view/16-view-simd branch from d4ac1b7 to 1f0c3be Compare September 30, 2026 19:19
@nlohmann
nlohmann removed this pull request from stack #5739 October 6, 2026 09:26
@nlohmann
nlohmann changed the base branch from json-view/15-view-bench to json-view/13-view-dump October 6, 2026 09:26
@nlohmann
nlohmann force-pushed the json-view/16-view-simd branch from e4f4c76 to 6ee5e89 Compare October 6, 2026 09:28
@nlohmann nlohmann changed the title Scan the strings of json_view with NEON and SSE2 Scan json_view strings with SIMD and index large objects Oct 6, 2026
@nlohmann
nlohmann added this pull request to stack #5768 October 6, 2026 09:29
Speed up json_view's parser with SIMD scanning and a hash table
for large objects.

Long runs of string bytes are scanned 16 bytes at a time with NEON
(AArch64, GCC and Clang) and SSE2 (x86-64), both baseline
instruction sets. Keys keep 16 table checks before the vector
loop, because their lengths repeat from record to record; string
values get 8, because their lengths vary more. Non-ASCII text is
validated 16 bytes at a time with simdjson's "lookup4" check
(Keiser and Lemire, 2021), with NEON on AArch64 and, on x86-64,
with SSSE3. SSSE3 is not part of baseline x86-64, so the check is
compiled for SSSE3 with a function attribute and used only where
CPUID reports it, which all x86-64 CPUs since about 2011 do; the
answer is cached in a statically initialized atomic, so there is
no guard of a local static and no global constructor. The same
input is accepted either way. JSON_VIEW_NO_SIMD selects the
portable code.

On x86-64, string runs are now checked vector-first: one SSE2
compare from the first byte finds the end of most keys and short
values, instead of a branch per byte for the first 8-16 bytes.
AArch64 keeps the byte-wise steps, where a NEON mask costs more and
the branches predict well. Entering an object or array no longer
stalls: open() stores the parent's frame field by field instead of
building it on the stack and reading it back with wider loads,
which waited for the narrower stores to retire.

Objects with 128 members or more get an open-addressing hash table
built when the object closes, so operator[], at(), find(),
contains(), count(), value(), and JSON pointers take constant time
on average in such objects; of duplicate keys, the first is kept,
as for the linear search. The idea comes from Boost.JSON.

simdjson is credited in simd.hpp's SPDX block, the README, and
license.md.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
@nlohmann
nlohmann force-pushed the json-view/16-view-simd branch from 6ee5e89 to 8daec2b Compare October 7, 2026 14:44
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC ignores the target attribute in modules, so the SSSE3 dispatch is
disabled for the module interface (the check stays portable, SSE2 is kept).
Make the 8-vs-16 byte unrolling condition in scan_string_run a
preprocessor/template split to avoid a constant condition (C4127).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/view/document_data.hpp
#	single_include/nlohmann/json_view.hpp
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The key hash is not seeded, so keys chosen to collide made building the table quadratic (20,000 colliding keys took 470 ms to parse). A key may now sit at most 64 slots from its home slot; if a key would sit further away, the table is dropped and the object is searched linearly. Lookups stop after the same distance.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
shrink_to_fit() now trims the tables of large objects like the node array and the decoded strings, and the list of large objects is released as soon as the tables are built.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
WebAssembly with -msse2 defines __SSE2__, but has no <cpuid.h> or x86 intrinsics headers.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects with 128 members or more now return the last member of
a repeated key, like the linear search of smaller objects and like
materialize() and basic_json::parse(). build_object_index let the first
occurrence win, so the same text gave different results depending on the
size of the object.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects with 128 members or more return the first member of a
repeated key again, like the linear search of smaller objects. This undoes
the code change of a695420; its test now expects the first member from
lookups (with and without a table) and the last value from materialize()
and basic_json::parse().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Silence bugprone-casting-through-void for the SSE loads (a reinterpret_cast
would trip -Wcast-align=strict), use auto for a cast initialiser, add
parentheses to a mixed expression, and name bugprone-std-namespace-modification
in the NOLINTs of the tuple_size/tuple_element specialisations.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
…uilder.cpp

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
….cpp

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CMake documentation L review needed It would be great if someone could review the proposed changes. tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant