Skip to content

Add BON8 support - #2998

Merged
nlohmann merged 17 commits into
developfrom
bon8
Sep 27, 2026
Merged

nlohmann merged 17 commits into
developfrom
bon8

Conversation

@nlohmann

@nlohmann nlohmann commented Sep 4, 2021 •

Copy link
Copy Markdown
Owner

This PR adds support for BON8 (Binary Object Notation 8), see #2980. BON8 uses the byte values that cannot begin a UTF-8 character as type markers, so strings are stored as plain UTF-8 without a length prefix. Small integers, true/false/null, and -1.0/0.0/1.0 take one byte, and containers with up to four elements need no terminator.

The branch was rebuilt from scratch on top of develop; the previous commits (2021/2022) could not read back their own output.

Size

BON8 is the most compact of the supported formats on the benchmark files (sizes relative to minified JSON):

Format canada.json twitter.json citm_catalog.json jeopardy.json
CBOR 50.5 % 86.3 % 68.4 % 88.0 %
MessagePack 50.5 % 86.0 % 68.5 % 87.9 %
BON8 50.5 % 83.8 % 63.5 % 87.5 %

Changes

  • to_bon8 (vector, uint8_t and char output adapters) and from_bon8 (input, iterator/sentinel), plus input_format_t::bon8 for sax_parse.
  • Reader: non-recursive, built on container_stack like the other binary readers. A string ends at the first byte that cannot continue it, which is the first byte (or, for integers that start with a UTF-8 lead byte, the first two bytes) of the next value; the reader hands these bytes back to the next value through a two-byte pushback buffer. Strings are fully validated as UTF-8. Non-canonical input (longer encodings than necessary, 0x85/0x8B containers with few elements, unsorted keys, unneeded 0xFF) is accepted.
  • Writer: produces the specification's canonical representation (shortest encodings, binary32 whenever lossless, -0.0/infinity/NaN as binary32, NaN as 0x7F800001, sorted keys for json), except for NFC normalization. Unsigned integers above int64 throw out_of_range.407, strings that are not valid UTF-8 throw type_error.316, and binary values are written as arrays of integers (like UBJSON/BJData).
  • Tests: unit-bon8.cpp with the integer test vectors of the reference implementation (HikoGUI, Boost Software License), the examples of the specification and of BON8 support #2980, every integer from -300000 to 600000, string-termination cases, error messages, SAX aborts, deep nesting, and byte-for-byte round trips over 140 files of the test data. BON8 was also added to the other cross-format tests (sizes, alternative string/object/binary types, ordered_json, output sinks, deque arrays), the benchmarks, and a new fuzzer fuzzer-parse_bon8.cpp (picked up by OSS-Fuzz's build.sh automatically).
  • Test data: the .bon8 files were generated with HikoGUI's reference encoder and published in json_test_data v3.2.0 (Add BON8 test files json_test_data#6). Before that, the reference encoder and to_bon8 were checked to produce identical bytes on every file, and each decoder to read the other's output back. JSON_TEST_DATA_VERSION is bumped to 3.2.0.
  • Docs: a feature page with the type mappings, API pages for to_bon8/from_bon8, examples, and BON8 in the size table, README, and format lists.

Note that the specification's example {"a":["b","c"],"d":1} omits the 0xFF bytes after "b" and "c" that its own rules require; the tests use the corrected encoding 88 61 82 62 FF 63 FF 64 91.

Public API

No breaking changes. The change only adds API: basic_json::to_bon8, basic_json::from_bon8, and the enumerator input_format_t::bon8, which is appended at the end so the values of the existing enumerators are unchanged.


This description and the implementation were written by Claude Code.

🤖 Generated with Claude Code

@nlohmann nlohmann added the aspect: binary formats BSON, CBOR, MessagePack, UBJSON label Sep 4, 2021
@nlohmann nlohmann added this to the Release 3.10.3 milestone Sep 4, 2021
@nlohmann nlohmann self-assigned this Sep 4, 2021
@nlohmann
nlohmann marked this pull request as draft September 4, 2021 12:56
@coveralls

coveralls commented Sep 4, 2021 •

Copy link
Copy Markdown

Coverage Status

Coverage decreased (-0.2%) to 99.842% when pulling f9003d1 on bon8 into 8fcdbf2 on develop.

@nlohmann nlohmann removed this from the Release 3.10.3 milestone Oct 8, 2021
@gregmarr

Copy link
Copy Markdown
Contributor

@nlohmann @takev Is this still in progress?

@nlohmann

Copy link
Copy Markdown
Owner Author

It is currently on hold from my side. It's definitely a nice idea for the format, but I currently don't find the time to push this further.

Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.

The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.

The round-trip tests need the .bon8 files of json_test_data 3.2.0.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
@nlohmann

Copy link
Copy Markdown
Owner Author

I revived this. I still thing BON8 is a nice format, and from all implemented, it's the most compact one.

Comment thread docs/mkdocs/docs/api/basic_json/to_bon8.md Outdated
Comment thread include/nlohmann/detail/input/binary_reader.hpp
Comment thread include/nlohmann/detail/output/binary_writer.hpp Outdated
Comment thread include/nlohmann/detail/output/binary_writer.hpp Outdated
Comment thread include/nlohmann/detail/output/binary_writer.hpp
Comment thread include/nlohmann/detail/output/binary_writer.hpp
pull Bot pushed a commit to sysfce2/json_test_data that referenced this pull request Sep 25, 2026
The files were created with the BON8 encoder of HikoGUI and cross-checked
against the BON8 implementation of JSON for Modern C++ (nlohmann/json#2998).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
  now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
  exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Comment thread tests/src/unit-bon8.cpp Fixed
Comment thread tests/src/unit-bon8.cpp Fixed
Comment thread tests/src/unit-bon8.cpp Fixed
Comment thread tests/src/unit-bon8.cpp Fixed
get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- compare the float in write_bon8_float with number_float_t constants,
  so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
  conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
@nlohmann nlohmann added this to the Release 3.13.0 milestone Sep 25, 2026
- copy the valid UTF-8 of a string in one step when the input is
  contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
  jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
  now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
  value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
@nlohmann nlohmann added the 🚀 ready to merge Ready to merge - just waiting for CI to complete. label Sep 25, 2026
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

/// whether the input is a contiguous block of bytes that BON8 strings can
/// be copied from in bulk; see @ref get_bon8_string_bulk
static constexpr bool bon8_bulk_scan =

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is this bon8 specific?

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point: it's a property of the input adapter, not of BON8. Renamed it to bulk_scan, like the lexer's equivalent, in 53ab406. To show it's not BON8-specific, BSON now uses it too (a1381a6): its keys and array indices are null-terminated and were read one byte at a time. They're now read up to the 0x00 in one step, which reads twitter.json 27% faster and citm_catalog.json and jeopardy.json 12% faster. The length-prefixed formats don't need it; they already memcpy contiguous input via get_elements().

(Written by Claude Code.)

// ASCII characters and complete well-formed UTF-8 sequences; n if all of it is
// valid UTF-8. Unlike scalar_string_bulk_run(), quotes, escapes, and control
// characters are ordinary characters here. ASCII is skipped 8 bytes at a time.
inline std::size_t valid_utf8_prefix(const unsigned char* data, std::size_t n) noexcept

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this different than the other utf8 bulk functions?

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, in which bytes stop it. string_bulk_run() / scalar_string_bulk_run() scan the inside of a JSON string, so they also stop at ", \ and control characters. In a BON8 string those are ordinary characters: only an invalid or incomplete UTF-8 sequence ends the run, and that is also where the string ends. That's also why the simdutf path doesn't carry over. It validates the run up to the next delimiter in one call, but a BON8 string has no delimiter to find in advance; its end only shows up as the first byte that isn't valid UTF-8. Apart from the stop set, the loop is the same: skip ASCII 8 bytes at a time, then validate_one_utf8() for anything else.

(Written by Claude Code.)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
  when exceptions are disabled: they catch the parse errors of invalid
  input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).

Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
@nlohmann
nlohmann merged commit 1e101ec into develop Sep 27, 2026
162 of 163 checks passed
@nlohmann
nlohmann deleted the bon8 branch September 27, 2026 14:56
nlohmann added a commit that referenced this pull request Sep 27, 2026
New public members get an entry in docs/docset/docSet.sql (as done for
to_bon8/from_bon8 in #2998). Without it, the Dash/Zeal docset built from
the documentation cannot find basic_json::as_base_class.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
nlohmann added a commit that referenced this pull request Oct 4, 2026
… base classes (#5589)

* Add basic_json::as_base_class and document name conflicts with custom base classes

Members of basic_json hide members of a custom base class with the same
name, and future releases may add members that hide ones accessible
today. Document this in json_base_class_t and add as_base_class() to
reach hidden members without spelling out the cast.

Also make json_base_class_t a public member type. It was documented
since 3.12.0, but declared private, so users could not name it.

Supersedes #3899.

Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add as_base_class to the docset search index

New public members get an entry in docs/docset/docSet.sql (as done for
to_bon8/from_bon8 in #2998). Without it, the Dash/Zeal docset built from
the documentation cannot find basic_json::as_base_class.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Silence clang-tidy for the hidden type_name() in the base class test

ci_clang_tidy failed with readability-convert-member-functions-to-static
on base_class_with_hidden_members::type_name(). It must stay a
non-static member: the test shows that it is hidden by the non-static
basic_json::type_name() and reachable through as_base_class().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

aspect: binary formats BSON, CBOR, MessagePack, UBJSON CI CMake documentation L 🚀 ready to merge Ready to merge - just waiting for CI to complete. tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants