Repository navigation
Conversation
nlohmann
added this pull request to stack #5636
September 29, 2026 14:20
nlohmann
removed this pull request from stack #5636
September 30, 2026 13:19
nlohmann
force-pushed
the
json-view/22-view-dump-fast
branch
from
September 30, 2026 13:20
2202ac0 to
b6f0daf
Compare
nlohmann
added this pull request to stack #5739
September 30, 2026 13:21
nlohmann
force-pushed
the
json-view/22-view-dump-fast
branch
3 times, most recently
from
September 30, 2026 18:06
d50ae18 to
61afa69
Compare
nlohmann
marked this pull request as ready for review
September 30, 2026 18:16
nlohmann
force-pushed
the
json-view/22-view-dump-fast
branch
from
September 30, 2026 19:19
61afa69 to
1df24a0
Compare
nlohmann
removed this pull request from stack #5739
October 6, 2026 09:26
This was referenced Oct 6, 2026
nlohmann
force-pushed
the
json-view/22-view-dump-fast
branch
from
October 6, 2026 09:28
2db0e30 to
64c3dbd
Compare
nlohmann
added this pull request to stack #5768
October 6, 2026 09:29
| { | ||
| if (n <= 32 && static_cast<std::size_t>(src_end - from) >= 32) | ||
| { | ||
| std::memcpy(w, from, 32); |
| { | ||
| if (n <= 32 && static_cast<std::size_t>(src_end - from) >= 32) | ||
| { | ||
| std::memcpy(w, from, 32); |
The default dump() (no indentation, no ensure_ascii) gets its own writer that makes the same walk and produces the same output: - the write position stays in a local variable instead of a member, so the compiler keeps it in a register across stores through aliasing char pointers; - strings and number tokens are copied with fixed-size 32-byte moves wherever enough source bytes remain, instead of one memcpy call per token; - the innermost open container lives in local variables; a stack that starts as a local array of 32 entries holds the rest; - unedited documents are walked through the node array in order, and integer tokens are read from the source directly. On top of that, float tokens of at most 15 significant digits are written straight from their digits via zmij::to_shortest() and write_shortest(), without converting to a double and back: such decimals are farther apart than a double's rounding interval, so the token's digits are the double's shortest digits. Tokens of 16+ digits, or edited values, still go through decimal_to_float(). The view's own NEON write_decimal() is removed in favor of the shared writer, and the dump output now grows in 64 KiB steps instead of being resized to its estimate at once. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
nlohmann
force-pushed
the
json-view/22-view-dump-fast
branch
from
October 7, 2026 14:44
64c3dbd to
accba4d
Compare
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of the stack for the zero-copy view (#5295). This PR combines #5633 and #5635, plus the view's serializer changes from #5765 (the rest of #5765 is in earlier PRs of this stack).
Summary
A dedicated writer for the default dump()
The default
dump()(no indentation, noensure_ascii) gets its own writer. It makes the same walk and produces the same output, with these changes:charpointers may alias anything, so with the position as a member of the output buffer, the compiler reloads it after every store.memcpycall per token. Up to 256 bytes, several 32-byte moves are used. The buffer keeps 64 bytes of slack for the overshoot, and its first size includes that slack, so it does not grow just before the end.Indented output and
ensure_asciistill use the general writer. Editable documents take the new path too (through the navigation policy); strings and tokens written by edits are copied with a plainmemcpy.Floats written from their digits
json_view::dump()writes a float token of at most 15 significant digits from its digits, without converting it to a double and back. The output is unchanged: it is whatjson::dump()writes for the same value.DBL_DIG. So the token's digits, without trailing zeros, are the shortest digits of its double. The library now writes exactly those (Żmij, viazmij::to_shortest()anddtoa_impl::write_shortest()); with Grisu2 this would not hold.decimal_to_float(), so the token is not read twice.write_short_decimal()handles tokens of up to 15 digits once their digit count is known from the token, without the checks of the generalto_chars()path.write_decimal()is removed; both NEON and non-NEON targets now go through the shared writer.Output growth
dump()'s output now grows in steps of 64 KiB within its reserved size estimate (the source extent), instead of being resized to the estimate at once.resize()zero-fills, and for a pretty-printed source the estimate is far larger than the compact output.The techniques come from the prototype: when the view was split into PRs, the writer was simplified, and
dump()became 2–3x slower. Keeping the innermost container in registers is new.Performance
json_view
Measured on the regrouped stack
Measured on the regrouped stack (µs, best of 3 interleaved rounds of 15 runs; files from nativejson-benchmark). Apple M1 Max with Apple clang at
-O2; x86-64 on a KVM Haswell VPS, pinned to one core, GCC 13 / Clang 18 at-O2(the VPS is noisy, about ±5–10%).dump()of the whole document, before this PR (measured at the SIMD PR; the edit and image PRs keep the general writer) → this PR:Earlier measurements (on the old stack)
dump()of a whole document, before vs. after this change:Writing floats from their digits (numbers 0.31, marine_ik 0.38, canada 0.86), plus the serializer part of #5765 (writing doubles with
write_shortest()in registers):json_viewdump of canada 7.84 → 7.22 ms on x86-64 and 4.49 → 3.76 ms on Apple M1.In
bench_edit(parse, edit, serialize): twitter goes from 463 to 311 µs, about 1.5–1.6x faster than yyjson's mutable documents.Core library (json::parse / dump)
Not affected: this PR only changes
include/nlohmann/detail/view/serializer.hppandjson_view.hpp.Tests
No new test files: the existing tests compare
dump()withordered_json::dump()on thousands of generated documents, on edited documents (the differential test), and on loaded images. Added coverage: 20,000 float tokens with 1 to 17 significant digits, in every spelling (with/without a point,e/Eexponents, signs, leading and trailing zeros, subnormals through 1e300), checked againstjson::dump()of the same text; on AArch64 against the NEON path too.Public API
No change.
Written by Claude Code.
🤖 Generated with Claude Code