Repository navigation
Bound UBJSON optimized arrays of a valueless type - #5504
Merged
nlohmann merged 2 commits intoSep 11, 2026
Merged
Conversation
nlohmann
marked this pull request as ready for review
September 6, 2026 17:43
nlohmann
force-pushed
the
ubjson-valueless-count-cap
branch
from
September 7, 2026 02:56
d89acce to
5bf814b
Compare
nlohmann
force-pushed
the
ubjson-valueless-count-cap
branch
from
September 7, 2026 05:59
5bf814b to
ee69490
Compare
nlohmann
force-pushed
the
ubjson-valueless-count-cap
branch
from
September 8, 2026 11:13
ee69490 to
42420d4
Compare
gregmarr
reviewed
Sep 8, 2026
| An array whose type marker is `Z` (null), `T` (true) or `F` (false) stores no payload at all, because the marker | ||
| already is the value. Its declared count is therefore the only thing that decides how much memory the receiving side | ||
| allocates, and a handful of bytes can describe billions of elements. `from_ubjson` rejects such an array with | ||
| [`out_of_range.408`](../../home/exceptions.md#jsonexceptionout_of_range408) when the count exceeds 1,048,576, and |
Contributor
There was a problem hiding this comment.
Put (1 << 20) or (0x100000) next to the decimal number so that it's obvious it's a "nice round number" in binary and hex domains? Also applies in other .md file.
Owner
Author
There was a problem hiding this comment.
Good idea — added (1 << 20) next to the decimal number in both spots (here and the matching text in exceptions.md), matching how max_valueless_container_size is actually defined in binary_reader.hpp. Pushed.
nlohmann
added a commit
that referenced
this pull request
Sep 8, 2026
Addresses review feedback from @gregmarr on PR #5504: spell out the binary/hex form next to the decimal count so it reads as the round power-of-two it is, matching how include/nlohmann/detail/input/binary_reader.hpp defines max_valueless_container_size. Applied in both docs/exceptions.md and ubjson.md, as requested. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
gregmarr
approved these changes
Sep 8, 2026
nlohmann
force-pushed
the
ubjson-valueless-count-cap
branch
from
September 9, 2026 08:21
5de50cd to
9edfb53
Compare
nlohmann
force-pushed
the
ubjson-valueless-count-cap
branch
from
September 10, 2026 15:15
9edfb53 to
7c44c0c
Compare
gregmarr
approved these changes
Sep 10, 2026
An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its type marker alone, so an optimized UBJSON array of one of those has no payload: reading an element consumes no input at all. Its declared count is therefore the only thing that decides how much is allocated, and nothing bounded it. "[$Z#l" and a four-byte count is nine bytes of input describing two billion values; #2793 reports 35 GB and 150 seconds from ten bytes, and OSS-Fuzz has an out-of-memory and a timeout report for the same shape. Every other type costs at least one byte per element, so the end of the input bounds it. 'N' (no-op) is already skipped rather than stored. Objects are not affected either: each element is preceded by its key, which costs bytes. And BJData already refuses these markers as an optimized type, so this is a plain UBJSON matter. Reject a count above 1,048,576 elements for those three types with out_of_range.408, the code this reader already uses for a declared size it will not honour. The check runs before the SAX start event, so no container is opened and then abandoned. Rejecting on the read side alone would break the guarantee that anything to_ubjson() writes can be read back, and would trip the round-trip assertion in fuzzer-parse_ubjson.cpp. So the writer falls back to the unoptimized encoding, one byte per element, for arrays of these types above the same limit. Its decision depends only on the array's size, which is identical for a value and for anything parsed back from it, so the round trip is stable. No existing test changes: the largest such count in the test suite is 65,793. The excessive-size test that already used this shape still passes, now rejected a little earlier than by the max_size() check it used to reach. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Addresses review feedback from @gregmarr on PR #5504: spell out the binary/hex form next to the decimal count so it reads as the round power-of-two it is, matching how include/nlohmann/detail/input/binary_reader.hpp defines max_valueless_container_size. Applied in both docs/exceptions.md and ubjson.md, as requested. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
nlohmann
force-pushed
the
ubjson-valueless-count-cap
branch
from
September 11, 2026 06:29
7c44c0c to
f50144c
Compare
This was referenced Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #2793. Also the root cause of the OSS-Fuzz out-of-memory and timeout reports for
parse_ubjson_fuzzer(issues 42506968 and 42506917, both open since January 2022).What
An element of type
'Z'(null),'T'(true) or'F'(false) is encoded by its type marker alone, so an optimized UBJSON array of one of those has no payload: reading an element consumes no input at all. Its declared count is therefore the only thing that decides how much is allocated, and nothing bounded it.[$Z#lplus a four-byte count is nine bytes of input describing two billion values. #2793 reports 35 GB and 150 seconds from ten bytes of input.Everything else is already bounded by the end of the input, because it costs at least one byte per element:
'N'(no-op) is skipped rather than storedHow
Reject a count above 1,048,576 elements for those three types with
out_of_range.408, the code this reader already uses for a declared size it will not honour. The check runs before the SAX start event, so no container is opened and then abandoned.Rejecting on the read side alone would break the guarantee that anything
to_ubjson()writes can be read back, and would trip the round-trip assertion infuzzer-parse_ubjson.cpp. So the writer falls back to the unoptimized encoding, one byte per element, for arrays of these types above the same limit. Its decision depends only on the array's size, which is identical for a value and for anything parsed back from it, so the round trip is stable:Verification
The #2793 payload is now rejected in 0 ms instead of allocating tens of gigabytes. At the limit the optimized form is still used (9 bytes for a million nulls); one past it the writer emits the unoptimized form and it still round-trips. No existing test changes: the largest such count in the test suite is 65,793.
API impact
No breaking changes to the public API. One accepted-input change: a plain UBJSON array of
$Z/$T/$Fdeclaring more than 1,048,576 elements is now rejected without_of_range.408instead of being materialised. Documented inexceptions.mdandubjson.md.Checklist
make amalgamate.🤖 Generated with Claude Code