Repository navigation
Report BON8 input that ends after a UTF-8 lead byte as truncated - #5677
Merged
Merged
Conversation
A lead byte (0xC2..0xF7) inside a string begins either another character
(if a continuation byte follows) or an integer (otherwise). When the input
ended right after the lead byte, the reader took the missing byte as "not
a continuation byte", ended the string before the lead byte, and treated
the lead byte as the start of the next value. With strict=false, a message
cut off there was therefore read as a shorter value: the 11 bytes of
"😀😀é" cut after 9 bytes gave "😀😀", and ["aé"] cut after 3 of its 5
bytes gave ["a"]. With strict=true, the input was rejected with a
misleading message ("expected end of input"), or, for a key, with
parse_error.112 instead of 110.
Either reading of the lead byte leaves the message incomplete: a string at
the end of a message must be terminated by 0xFF, so the lead byte cannot
belong to a following message. Report parse_error.110 (unexpected end of
input) for strings and keys, as the comment on get_bon8_string() already
requires and as the reference decoder (HikoGUI) does.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
7 tasks done
nlohmann
added a commit
that referenced
this pull request
Oct 1, 2026
…5731) * Share the diagnostic-position setter of the DOM SAX parsers json_sax_dom_parser and json_sax_dom_callback_parser each had a private copy of handle_diagnostic_positions_for_json_value(), identical except for comments. Move the body into one static member function, detail::diagnostic_positions::set_from_lexer(value, lexer), which both classes call with their lexer pointer. basic_json befriends the new struct (only when JSON_DIAGNOSTIC_POSITIONS is enabled), as the position members are private. The discarded case is reached through the callback parser, so the LCOV_EXCL markers that only the dom parser's copy had are gone. The NOLINT on the unreachable default case loses the stray "-warnings-as-errors", which is not a check name. The start-position setup in start_object()/start_array() is left alone, as #5706 is editing the callback parser's versions. Behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Correct the parser comments on recursion and skip_to_state_evaluation The class documentation called the parser a recursive descent parser, but sax_parse_internal() is a loop that keeps the open containers on an explicit stack. The comment at the end of an array and of an object said the flag is set to false while the code below it sets it to true. Describe what the code does instead. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Update the discard_number_values comments to the current number path The comments explaining the accept() shortcut in convert_number() and the member documentation still argued in terms of strtoull()/strtoll() and errno, which #5283 replaced with convert_integer(), and pointed at scan_number() instead of convert_number(). They also did not say that scan_number_bulk_contiguous() converts integers itself, so the shortcut is only reached for input without bulk access, with JSON_DIAGNOSTIC_POSITIONS, or when the bulk scanner falls back. Rewrite both comments to describe the digit-count check in front of convert_integer(), keeping the 18-digit bound and the json_sax_acceptor argument. The stale <cstdlib> comment is left for after #5616, which edits that include block. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * List the UTF-8 validators instead of calling the DFA the only one The documentation of decode() called the Hoehrmann DFA the single source of truth for UTF-8 validation. It is used only by the serializer and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text strings). The lexer's scan_string() switch, validate_one_utf8() / valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer) and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges on their own. Replace the sentence with a list of the four validators, what each is used for, and a note that they must accept the same sequences. Sharing code between them was considered and dropped: it would save a few lines in a validator that is entangled with BON8 pushback, and #5677 is editing the BON8 byte path. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix stale doc comments and include lists in the input headers input_adapters.hpp included <memory> and <numeric> for the removed shared_ptr-based adapter design but used neither; it called (std::min) without including <algorithm>. json_sax.hpp used std::numeric_limits without including <limits>. Also corrected comments that no longer matched the code: input_stream_adapter does not skip the input's BOM (the lexer's skip_bom() does), the span_input_adapter comment named the no-longer-existing input_buffer_adapter type, lexer::get_string() does not reset the token, binary_reader's get_number() doc opened with /* instead of /*! (so Doxygen skipped it) and omitted BON8 from its endianness note, and the UBJSON-binary-types note did not mention that BJData 'B' arrays are read as binary. Left out: the lgtm suppression on lexer.hpp's scan_number() (in #5616's hunk) and the "-1 if unknown" wording in json_sax.hpp's start_object/start_array docs (in draft #5267's hunk), per the verdict's conflict list. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the strict-EOF/release_lookahead/error block in parser::parse() json_sax_dom_callback_parser and json_sax_dom_parser branches of parser::parse() ran the same ~25 lines after sax_parse_internal(): the strict-mode EOF check (raising parse_error.101 through the SAX parser), release_lookahead() in non-strict mode, and mapping an errored SAX parser to a discarded result. The two copies had already drifted apart in formatting and in the second copy's "see above" comment. Add a private parse_dom(DomSax&, strict) member that runs this shared sequence once and returns whether the SAX parser did not error; both branches of parse() now only construct their DOM SAX parser, call parse_dom(), and (for the callback parser) map a discarded top-level value to null. sax_parse() is left untouched, since it only runs the EOF check and release_lookahead() when sax_parse_internal() succeeded, unlike parse(), which runs them unconditionally. Behavior-preserving: same operations in the same order for both SAX parser kinds. Verified with unit-class_parser (strict/non-strict, callback and non-callback), unit-deserialization and unit-disabled_exceptions (JSON_NOEXCEPTION), plus a clean make amalgamate / make check-amalgamation diff. Overlaps #5601, which touches the same lines. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5712 item 2 * Share the code point to UTF-8 encoding between the wide-string helpers and the lexer The 1/2/3/4-byte UTF-8 encoding ladder was written out by hand three times: in wide_string_input_helper<..., 4>::fill_buffer() for a UTF-32 code point, in the UTF-16 helper for both a BMP code unit and a valid surrogate pair, and in the lexer's \uXXXX/\uXXXX\uYYYY handling. The copies had drifted: the UTF-32 helper masked the leading bits of each byte (& 0x1Fu, & 0x0Fu, & 0x07u) where the others relied on the shift alone, even though both give the same result for a code point that is already known to be in range. Add detail::encode_utf8(cp, out) in string_utils.hpp, a single encoder that invokes a callable once per output byte, most significant byte first. Use it in the three valid-code-point branches (UTF-32 code points up to U+10FFFF, UTF-16 code units outside the surrogate range, and valid UTF-16 surrogate pairs) and in the lexer's \u handling, where out forwards to add(). The UTF-16 helper's deliberate pass-through of malformed surrogate units and the UTF-32 helper's 0xFF sentinel for code points above U+10FFFF are untouched, since neither reaches the new helper. Behavior-preserving: same bytes in the same order for every valid code point, verified with unit-class_lexer, unit-class_parser, unit-deserialization, unit-wstring and the non-test-data parts of unit-unicode1..5 (ASan/UBSan, C++11/17/20), and an escape-heavy parse microbenchmark that shows no change (about 73 ms either way, median of 3, 1M escape sequences). single_include/ regenerated with make amalgamate; make check-amalgamation leaves a clean tree. Overlaps #5704, which rewrites the wide_string_input_helper specializations touched here. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5712 item 6 --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
gregmarr
approved these changes
Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bug
Inside a BON8 string, a UTF-8 lead byte (0xC2..0xF7) begins either another character, if a continuation byte follows, or an integer, if not. When the input ended right after the lead byte, the reader took the missing byte as "not a continuation byte". It ended the string before the lead byte and read the lead byte as the start of the next value.
With
strict = false, a message cut off at such a point was therefore read as a shorter value without an error:"😀😀é"cut after 9 of its 11 bytes"😀😀"81 61 C3(["aé"]cut after 3 of 5 bytes)["a"]81 C3 A9 C3["é"]With
strict = true, the same inputs were rejected, but with a misleading message ("expected end of input; last byte: 0xC3"). A key cut off this way (87 C3) gave parse_error.112 "expected a string" instead of 110.Why this is wrong
get_bon8_string()already says a string "must not end at the end of the input: the last string of a message is always terminated by 0xFF".Evidence
to_bon8and cut each message at every byte (680,617 proper prefixes). The reference decoder rejects all of them.from_bon8(…, strict = false)returned a value for 13,664 of them before this change, and for none after it.4D F5 2A). Here the second byte already identifies an integer, and such input cannot be a prefix of a valid message.Fix
In
get_bon8_string()andget_bon8_key(), report parse_error.110 ("unexpected end of input") if the input ends right after a lead byte. This covers both the contiguous (bulk) and the byte-by-byte (stream) paths, because the bulk scan stops at an incomplete sequence and hands it to this code.Tests
unit-bon8has two new sections:Both sections fail on
develop.unit-bon8passes for C++11, C++17, and C++20 with ASan/UBSan against the modular headers, and for C++17 againstsingle_include. There are no new-Weverythingwarnings in the changed lines.Public API
No breaking change to a released API: BON8 support (#2998) is not in a release yet. Within
develop,from_bon8(…, strict = false)now throws (or returns a discarded value withallow_exceptions = false) for truncated input that it used to read as a shorter value. Some error messages for such input also change: a cut-off key now gives parse_error.110 instead of parse_error.112. Valid messages decode exactly as before.This PR was written by Claude Code on behalf of @nlohmann.
🤖 Generated with Claude Code