Skip to content

Write json_view's dump() without a call or conversion per token - #5633

Open
nlohmann wants to merge 2 commits into
json-view/21-imagesfrom
json-view/22-view-dump-fast
Open

nlohmann wants to merge 2 commits into
json-view/21-imagesfrom
json-view/22-view-dump-fast

Conversation

@nlohmann

@nlohmann nlohmann commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Part of the stack for the zero-copy view (#5295). This PR combines #5633 and #5635, plus the view's serializer changes from #5765 (the rest of #5765 is in earlier PRs of this stack).

Summary

A dedicated writer for the default dump()

The default dump() (no indentation, no ensure_ascii) gets its own writer. It makes the same walk and produces the same output, with these changes:

  • The write position stays in a local variable. Stores through char pointers may alias anything, so with the position as a member of the output buffer, the compiler reloads it after every store.
  • Strings and number tokens of the source are copied with fixed-size moves of 32 bytes wherever the source has that many bytes left, instead of one memcpy call per token. Up to 256 bytes, several 32-byte moves are used. The buffer keeps 64 bytes of slack for the overshoot, and its first size includes that slack, so it does not grow just before the end.
  • The innermost open container is kept in local variables; the stack holds only the ones around it. That stack starts in a local array of 32 entries and moves to the heap only for deeper nesting, and its address does not escape, so its pointers stay in registers. Shallow documents need no allocation besides the output.
  • Documents that are not edited are walked through the node array in order, so a container only needs its end, and integer tokens are read from the source directly.

Indented output and ensure_ascii still use the general writer. Editable documents take the new path too (through the navigation policy); strings and tokens written by edits are copied with a plain memcpy.

Floats written from their digits

json_view::dump() writes a float token of at most 15 significant digits from its digits, without converting it to a double and back. The output is unchanged: it is what json::dump() writes for the same value.

  • Why this is exact. Two decimals of at most 15 significant digits are farther apart than the rounding interval of a normal double; this is the argument behind DBL_DIG. So the token's digits, without trailing zeros, are the shortest digits of its double. The library now writes exactly those (Żmij, via zmij::to_shortest() and dtoa_impl::write_shortest()); with Grisu2 this would not hold.
  • When it applies. The exponent must keep the value away from subnormals and overflow (leading digit between 10^−290 and 10^304).
  • Other tokens (16 or more digits, or edited values) are converted from the digits already read, with decimal_to_float(), so the token is not read twice.
  • write_short_decimal() handles tokens of up to 15 digits once their digit count is known from the token, without the checks of the general to_chars() path.
  • The view's own NEON write_decimal() is removed; both NEON and non-NEON targets now go through the shared writer.
  • Doubles are written into the output directly, not through a local buffer.

Output growth

dump()'s output now grows in steps of 64 KiB within its reserved size estimate (the source extent), instead of being resized to the estimate at once. resize() zero-fills, and for a pretty-printed source the estimate is far larger than the compact output.

The techniques come from the prototype: when the view was split into PRs, the writer was simplified, and dump() became 2–3x slower. Keeping the innermost container in registers is new.

Performance

json_view

Measured on the regrouped stack

Measured on the regrouped stack (µs, best of 3 interleaved rounds of 15 runs; files from nativejson-benchmark). Apple M1 Max with Apple clang at -O2; x86-64 on a KVM Haswell VPS, pinned to one core, GCC 13 / Clang 18 at -O2 (the VPS is noisy, about ±5–10%).

dump() of the whole document, before this PR (measured at the SIMD PR; the edit and image PRs keep the general writer) → this PR:

Apple M1 x86-64 GCC x86-64 Clang
canada 5287 → 3680 (−30%) 8179 → 6949 (−15%) 7986 → 7178 (−10%)
citm_catalog 378 → 132 (−65%) 710 → 266 (−63%) 631 → 279 (−56%)
twitter 238 → 74 (−69%) 379 → 166 (−56%) 400 → 187 (−53%)

Earlier measurements (on the old stack)

dump() of a whole document, before vs. after this change:

file before after ratio
twitter 225 µs 70 µs 0.31
twitterescaped 265 µs 79 µs 0.30
update-center 298 µs 78 µs 0.26
poet 601 µs 244 µs 0.41
citm_catalog 365 µs 136 µs 0.37
gsoc-2018 767 µs 506 µs 0.66
marine_ik 8.55 ms 7.35 ms 0.86
mesh.pretty 2.70 ms 2.39 ms 0.89
canada 9.61 ms 8.98 ms 0.93

Writing floats from their digits (numbers 0.31, marine_ik 0.38, canada 0.86), plus the serializer part of #5765 (writing doubles with write_shortest() in registers): json_view dump of canada 7.84 → 7.22 ms on x86-64 and 4.49 → 3.76 ms on Apple M1.

In bench_edit (parse, edit, serialize): twitter goes from 463 to 311 µs, about 1.5–1.6x faster than yyjson's mutable documents.

Core library (json::parse / dump)

Not affected: this PR only changes include/nlohmann/detail/view/serializer.hpp and json_view.hpp.

Tests

No new test files: the existing tests compare dump() with ordered_json::dump() on thousands of generated documents, on edited documents (the differential test), and on loaded images. Added coverage: 20,000 float tokens with 1 to 17 significant digits, in every spelling (with/without a point, e/E exponents, signs, leading and trailing zeros, subnormals through 1e300), checked against json::dump() of the same text; on AArch64 against the NEON path too.

Public API

No change.


Written by Claude Code.

🤖 Generated with Claude Code

@nlohmann
nlohmann added this pull request to stack #5636 September 29, 2026 14:20
@nlohmann
nlohmann removed this pull request from stack #5636 September 30, 2026 13:19
@nlohmann
nlohmann force-pushed the json-view/22-view-dump-fast branch from 2202ac0 to b6f0daf Compare September 30, 2026 13:20
@nlohmann
nlohmann added this pull request to stack #5739 September 30, 2026 13:21
@nlohmann
nlohmann force-pushed the json-view/22-view-dump-fast branch 3 times, most recently from d50ae18 to 61afa69 Compare September 30, 2026 18:06
@nlohmann nlohmann added the review needed It would be great if someone could review the proposed changes. label Sep 30, 2026
@nlohmann
nlohmann marked this pull request as ready for review September 30, 2026 18:16
@nlohmann
nlohmann force-pushed the json-view/22-view-dump-fast branch from 61afa69 to 1df24a0 Compare September 30, 2026 19:19
Comment thread include/nlohmann/detail/view/serializer.hpp Dismissed
Comment thread include/nlohmann/detail/view/serializer.hpp Fixed
Comment thread include/nlohmann/detail/view/serializer.hpp Dismissed
Comment thread include/nlohmann/detail/view/serializer.hpp Dismissed
Comment thread single_include/nlohmann/json_view.hpp Dismissed
Comment thread single_include/nlohmann/json_view.hpp Fixed
Comment thread single_include/nlohmann/json_view.hpp Dismissed
Comment thread single_include/nlohmann/json_view.hpp Dismissed
@nlohmann
nlohmann removed this pull request from stack #5739 October 6, 2026 09:26
@nlohmann
nlohmann force-pushed the json-view/22-view-dump-fast branch from 2db0e30 to 64c3dbd Compare October 6, 2026 09:28
@nlohmann nlohmann changed the title Write compact dumps of json_view without a library call per token Write json_view's dump() without a call or conversion per token Oct 6, 2026
@nlohmann
nlohmann added this pull request to stack #5768 October 6, 2026 09:29
{
if (n <= 32 && static_cast<std::size_t>(src_end - from) >= 32)
{
std::memcpy(w, from, 32);
{
if (n <= 32 && static_cast<std::size_t>(src_end - from) >= 32)
{
std::memcpy(w, from, 32);
The default dump() (no indentation, no ensure_ascii) gets its own
writer that makes the same walk and produces the same output:

- the write position stays in a local variable instead of a
  member, so the compiler keeps it in a register across stores
  through aliasing char pointers;
- strings and number tokens are copied with fixed-size 32-byte
  moves wherever enough source bytes remain, instead of one
  memcpy call per token;
- the innermost open container lives in local variables; a stack
  that starts as a local array of 32 entries holds the rest;
- unedited documents are walked through the node array in order,
  and integer tokens are read from the source directly.

On top of that, float tokens of at most 15 significant digits are
written straight from their digits via zmij::to_shortest() and
write_shortest(), without converting to a double and back: such
decimals are farther apart than a double's rounding interval, so
the token's digits are the double's shortest digits. Tokens of
16+ digits, or edited values, still go through decimal_to_float().
The view's own NEON write_decimal() is removed in favor of the
shared writer, and the dump output now grows in 64 KiB steps
instead of being resized to its estimate at once.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
@nlohmann
nlohmann force-pushed the json-view/22-view-dump-fast branch from 64c3dbd to accba4d Compare October 7, 2026 14:44
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

L review needed It would be great if someone could review the proposed changes. tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants