Repository navigation
Handle numbers that do not fit narrow number types in the binary readers - #5607
Merged
Merged
Conversation
With custom number types narrower than the values in a binary document, for example basic_json<..., std::int32_t, std::uint32_t, float>, every binary reader (CBOR, MessagePack, UBJSON, BJData, BSON, BON8) passed the decoded number to the SAX interface with an implicit conversion: the integer 5000000000 silently became 705032704, and a finite double such as 1e300 became infinity. The lexer handles the same values in JSON text: an integer that fits neither integer type is stored as number_float_t, and a finite number that overflows number_float_t is rejected with out_of_range.406. Pass every number read from binary input through three helpers that apply the lexer's rules: - emit_signed(): number_integer_t, else number_unsigned_t for a non-negative value, else number_float_t - emit_unsigned(): number_unsigned_t, else number_float_t - emit_float(): out_of_range.406 if a finite value overflows number_float_t; infinity and NaN are passed on For consistency, a CBOR negative integer below the range of number_integer_t is now stored as number_float_t, like a too small integer in JSON text, instead of being rejected with parse_error.112. With the default number types, this is the only change in behavior. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
gregmarr
reviewed
Sep 28, 2026
MSVC types 3000000000 and 5000000000 as unsigned long, so json(-3000000000) triggered C4146 (unary minus on an unsigned type), which /WX turns into an error. Use LL literals, as elsewhere in the tests. clang 3.5 cannot convert the lambdas in the braced initializer of the format table to function pointers. Use named functions instead. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Owner
Author
|
CI fix: MSVC and
Fixed in 0c630d4. The narrow-number test case passes with clang 3.5 (CI image), GCC 14 (strict warnings, — posted by Claude Code on behalf of @nlohmann |
emit_signed, emit_unsigned, and the CBOR negative integer fallback now pass their number_float_t fallback through emit_float, so a value that overflows number_float_t is rejected with out_of_range.406 like a floating-point value, instead of silently becoming infinity. This only matters for a number_float_t that cannot represent 2^64, such as a half-precision type. The CBOR value -1 - n is computed as long double so that emit_float sees a finite value. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This was referenced Sep 29, 2026
Closed
nlohmann
added a commit
that referenced
this pull request
Sep 30, 2026
parse_cbor_internal() hand-wrote the same "read a 1/2/4/8-byte big-endian unsigned integer" ladder four times over: - twice for tag numbers 0xD8-0xDB, once in the tag_handler::ignore branch and once, nearly identically, in the ::store branch (~90 lines to read one integer); - twice more for container lengths, once for array heads 0x98-0x9B and once for map heads 0xB8-0xBB, where the 1/2-byte forms called enter_array()/enter_object() directly and the 4/8-byte forms additionally went through get_cbor_container_size(). Add get_cbor_argument(std::uint64_t&), reading the width selected by current & 0x1F via the same get_number() calls as before (so EOF is reported exactly as before), and route all four sites through it: - 0xD8-0xDB now read the argument once per branch instead of switching on `current` a second time; behavior split cleanly from embedded tags 0xC0-0xD7 (tag value in the head, no argument to read), which is now its own case block that no longer has to fall into the ::store switch's "default" case to reach the same tag_pending = true; return true; outcome. - 0x98-0x9B and 0xB8-0xBB collapse into one case block each, always going through get_cbor_container_size() (harmless for 1/2-byte lengths, which already always fit). Verified byte-for-byte identical behavior before/after with a standalone probe covering embedded and multi-byte tags under all three tag_handler_t settings, a tag over a byte string (subtype path), truncated tag/length arguments of every width, and array/map lengths of every width, including the out_of_range.408 "excessive size" case: same exceptions, same messages, same chars_read, same successful results. Left the string/byte-string length ladders in get_cbor_string()/ get_cbor_binary() untouched, as noted in #5711 item 2, since #5325 is expected to touch them separately. Overlaps #5601 (adds a branch right above the embedded-tag case) and #5607 (touches the integer cases 0x18-0x1B, which share this ladder's shape in separate hunks). #5711 item 2 Signed-off-by: Niels Lohmann <mail@nlohmann.me>
nlohmann
added a commit
that referenced
this pull request
Sep 30, 2026
…ecks get_ubjson_size_value()'s 'i'/'I'/'l'/'L' cases each read a differently sized signed integer and then repeated the same "reject negative with error 113" check; only 'L' additionally checked value_in_range_of for the out_of_range.408 case. Any change to that error path had to be made four times. Add get_ubjson_signed_count<SignedType>(std::size_t&), doing the read, the negative check and the range check once, and route all four markers through it. The range check is a no-op for 'i'/'I'/'l' (their values always fit std::size_t) and only live for 'L' on a 32-bit std::size_t target, matching today's behavior exactly. In the ndarray dimension-product loop, the preceding loop already returns early on any zero dimension and result starts at 1, so `i > 0` in the pre-multiplication overflow check was always true, and `result == 0` in the post-multiplication check could not be reached either: two positive factors whose product does not overflow (as the pre-check already guarantees) cannot be zero. Drop the dead `i > 0 &&` and narrow the post-check to `result == npos`, the one case the pre-check cannot rule out (an exact, non-overflowing match with the sentinel reserved for unknown-size containers), with a comment explaining why. Verified byte-for-byte identical behavior before/after with a standalone probe covering negative counts for every marker, a matching positive count, and ndarray inputs, plus the full unit-ubjson and unit-bjdata suites (same assertion counts as before this change). Overlaps #5601 (rewrites the four parse_error calls and the overflow checks touched here) and #5607/#5707 (touch neighboring lines in the same functions). #5711 item 6 Signed-off-by: Niels Lohmann <mail@nlohmann.me>
nlohmann
added a commit
that referenced
this pull request
Sep 30, 2026
* Fix stale and missing comments in binary_writer The doc block of write_number() ended up above the byte_swap() helpers added in #5286, about 80 lines from the function. It was also a plain comment that Doxygen skips, said "write a number to output input", and left BON8 out of the big-endian formats. Move it back onto write_number() as a /*! block and fix the text. write_bson() documented "@pre j.type() == value_t::object", but it throws type_error.317 for every other type, and to_bson() relies on that. Document the exception instead. Explain why the CBOR binary subtype is always written with a 0xD8..0xDB head and never in the one-byte tag form: binary_reader with cbor_tag_handler_t::store only keeps those heads as a subtype, so switching to write_cbor_head() would break round trips for subtypes 0..23. Also fix the grammar of the to_char_type comment. Comments only; no change in behavior, API or ABI. Part of #5710 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Merge the duplicated UBJSON/BJData integer marker ladders write_number_with_ubjson_prefix() (unsigned and signed overloads) and ubjson_prefix() (number_integer and number_unsigned cases) each picked the UBJSON/BJData integer marker (i, U, I, u, l, m, L, M, H) with their own independent if/else ladder, and the values beyond 64 bits were handled by a second, tag-dispatched pair of ladders. An optimized container announces the marker of its first element via ubjson_prefix() and then writes every element through write_number_with_ubjson_prefix(), so the two had to be kept in lockstep by hand across four call sites. Replace all of that with one ubjson_integer_prefix() built on value_in_range_of<T>, and one write_ubjson_integer_payload() that writes the value (or, for 'H', the decimal digits) for a given marker. write_number_with_ubjson_prefix() and ubjson_prefix() keep their signatures and now just call these two helpers. Behavior, the public API and the ABI are unchanged. Verified with a new regression test covering scalars and $-optimized arrays/objects at every int8/uint8/int16/uint16/int32/uint32/int64/uint64 boundary for to_ubjson/to_bjdata (both use_size/use_type settings), and by diffing to_ubjson/to_bjdata output before and after over the json_test_data corpus (bit-identical). Part of #5710 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove dead get_char parameters in binary_reader The non-recursive rewrite of the binary readers (#5505, #5506, #5507) left parse_cbor_internal()'s and parse_ubjson_internal()'s get_char parameters dead: parse_cbor_internal() has one caller and it always passes true, and parse_ubjson_internal() has one caller and it always uses the true default. Both parameters, and the @PARAM docs describing the "reuse the last character" mode they used to select, no longer correspond to anything. Drop both parameters, initialise fetch/prefix unconditionally, and update the two call sites in sax_parse(). parse_cbor_value()'s and get_ubjson_string()'s own get_char parameters are unrelated and are left alone; both still have a false caller. Also delete a stray `@return whether a valid MessagePack value was passed to the SAX parser` doxygen block that sits directly above parse_msgpack_value()'s real doc comment, a leftover of the same rewrite. Behavior, the public API and the ABI are unchanged; these are private members of detail::binary_reader. Verified by compiling with -Wunused-parameter and running unit-cbor, unit-ubjson, unit-bjdata and unit-msgpack (offline, against the stubbed test_data.hpp). Part of #5711 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Share the IEEE half-precision decoder between CBOR and BJData binary_reader had two ~45-line copies of the IEEE 754 half-precision decoder: CBOR's case 0xF9 and BJData's case 'h'. Once formatting is normalised, the two blocks were identical except for the byte order used to assemble the 16-bit half (CBOR is big endian, BJData is little endian). Any future change to half-float decoding had to be made and kept in sync in both places. Add one get_half_float(format, little_endian) helper that does the two get()/unexpect_eof() reads, assembles the half in the requested byte order, decodes it per RFC 8949 Appendix D, and calls sax->number_float. Both cases now just call it with their byte order; the BJData case keeps its bjdata-only guard. Behavior, the public API and the ABI are unchanged. Verified with a scratch probe comparing the old and new decoders bit-for-bit (NaN by isnan()) over all 65536 wire byte pairs, in both formats, and by running unit-cbor and unit-bjdata (offline, against the stubbed test_data.hpp). Part of #5711 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the MessagePack unsigned-integer writer ladder The number_integer (non-negative branch) and number_unsigned cases in write_msgpack() each held their own copy of the fixint/uint8/16/32/64 ladder, kept in lockstep only by a comment ("we used the code from the value_t::number_unsigned case here"). Both copies mixed union members: the signed copy compared number_unsigned but wrote number_integer, and vice versa. Extract write_msgpack_unsigned(std::uint64_t), mirroring how write_cbor_head() already avoids the same duplication for CBOR, and call it from both cases. Each case now reads only its own active union member. Output bytes are unchanged for the default 64-bit number types. #5710 item 3 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Unify float marker selection and fix the long double compile error Four formats picked between a float32 and float64 marker through four different helper styles: dummy-argument overloads for CBOR and MessagePack, an std::is_same template for BON8, and a runtime if-chain on input_format_t for write_compact_float(). With number_float_t set to long double, to_cbor, to_msgpack and to_ubjson failed inside the library with "call to 'get_cbor_float_prefix' is ambiguous", while to_bson kept working because write_bson_double() takes a plain double. Change write_compact_float() to take the two marker bytes directly (each of its three callers already knows them at compile time) instead of an input_format_t it only forwarded, and delete the now-unused get_cbor_float_prefix(), get_msgpack_float_prefix(), get_bon8_float_prefix() and get_compact_float_prefix() helpers. Turn the two get_ubjson_float_prefix() overloads into one template. Both write_compact_float() and get_ubjson_float_prefix() now report an unsupported number_float_t with a static_assert naming the requirement, rather than an ambiguous-overload error; the assert lives in the function body, not the class scope, so to_bson with long double is unaffected. Verified with a probe basic_json<..., long double>: to_bson still compiles and round-trips, while to_cbor/to_msgpack/to_ubjson now fail to compile with the new static_assert message. This changes the text of an existing compile error for users with an unsupported number_float_t (documented as a public-API-visible change in #5710). #5710 item 1 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the BJData ndarray writer's dtype dispatch and drop <map> write_bjdata_ndarray() built a 12-entry std::map<string_t, CharType> on every call just to translate the _ArrayType_ name to a dtype marker (the only reason binary_writer.hpp included <map>), then mapped dtype to C++ type twice more: once as a switch for the range-check pass and once as a separate if/else chain for the write pass, with nothing checking that the two agreed. The caller also ran three at() lookups, and the callee called value.at(key) about ten more times for the same three members. Replace the map with bjdata_ndarray_type_marker(), a plain string comparison chain (a C++11 constexpr function cannot contain a switch, so this mirrors binary_reader's own static table style). Replace the switch/if-chain pair with one write_bjdata_ndarray_elements() that switches on dtype once and calls a per-type helper - write_bjdata_ndarray_element<T>() for the eight integer dtypes and write_bjdata_ndarray_float_element() for 'd' - with a dry_run flag selecting the range check or the actual write, so the two passes can no longer disagree on the type. _ArrayType_, _ArraySize_ and _ArrayData_ are now looked up once into references, and the four header marker bytes ('[', '$', '#') are written through to_char_type() like the rest of the UBJSON/BJData writer. The 'd' (single-precision) rule is left exactly as before, since #5707 is expected to change it separately. Verified byte-for-byte identical output before/after for every dtype (including the Draft 2/Draft 3 'byte' fallback and the use_count/ use_type combinations) via a standalone probe, plus round-tripping through from_bjdata(). Overlaps #5707, which is expected to touch the 'd' dtype case, and #5518, which is expected to move the write_bjdata_ndarray() call site. #5710 item 4 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Assert that write_bson_document() consumes every calc_bson_sizes() entry calc_bson_sizes() and write_bson_document() are a hand-synchronized pair of passes over the same object/array tree, introduced by #5553: the size pass appends to nested_sizes in visiting order, and the write pass consumes the table by position with nested_sizes[next_size++]. Nothing checked that the write pass consumed the whole table. If a future change touched only one of the two passes - for example to skip or reject an entry - every later size prefix in the document would be silently wrong. Add JSON_ASSERT(next_size == nested_sizes.size()) where write_bson_document() returns, so such a future drift between the two passes is caught immediately (JSON_ASSERT expands to nothing in release builds using assert(), and the fuzzers/tests already build with it enabled). The two passes agree today, so this changes nothing observable; it only guards against the risk described in #5710 item 5. Extracting a shared stepper for the two passes (the second half of the proposed change) is left for a follow-up: it only saves ~30 lines and the issue asks for it only if the result reads clearly, which needs more room to get right than a mechanical cleanup pass allows. #5710 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Make the BJData lookup tables static functions instead of members binary_reader held bjd_optimized_type_markers and bjd_types_map as non-static const members (12 string_t objects for the type-name table), built and destroyed on every from_cbor/from_msgpack/from_bson/ from_ubjson/from_bon8/from_bjdata call even though only from_bjdata ever reads them. They also needed the #define/decltype/#undef workaround from #3637 and two NOLINTNEXTLINE suppressions, and binary_writer already carries the same two lists in another form (is_bjdata_excluded_type_marker() and a local std::map in write_bjdata_ndarray(), the latter removed by the item-4 commit), so the excluded-marker lists could drift apart. Replace bjd_optimized_type_markers with static constexpr is_bjd_excluded_optimized_type(char_int_type), using the same ||-chain as binary_writer's is_bjdata_excluded_type_marker(). Replace bjd_types_map with a non-constexpr static bjd_type_name(char_int_type) switch returning nullptr for an unknown marker (a C++11 constexpr function cannot contain a switch). Delete both JSON_BINARY_READER_MAKE_* macros, the bjd_type pair alias, the NOLINTNEXTLINE suppressions, detail::make_array() (no longer used anywhere), and the now-unused <algorithm> and <array> includes. Update the two call sites (the ND-array excluded-type check and the _ArrayType_ lookup) accordingly, and replace unit-bjdata.cpp's "LUT arrays are sorted" section, which only checked the two tables' internal ordering, with a check of all 12 type names and all 8 excluded markers against both new functions. #5711 item 1 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read CBOR's 1/2/4/8-byte argument through one helper parse_cbor_internal() hand-wrote the same "read a 1/2/4/8-byte big-endian unsigned integer" ladder four times over: - twice for tag numbers 0xD8-0xDB, once in the tag_handler::ignore branch and once, nearly identically, in the ::store branch (~90 lines to read one integer); - twice more for container lengths, once for array heads 0x98-0x9B and once for map heads 0xB8-0xBB, where the 1/2-byte forms called enter_array()/enter_object() directly and the 4/8-byte forms additionally went through get_cbor_container_size(). Add get_cbor_argument(std::uint64_t&), reading the width selected by current & 0x1F via the same get_number() calls as before (so EOF is reported exactly as before), and route all four sites through it: - 0xD8-0xDB now read the argument once per branch instead of switching on `current` a second time; behavior split cleanly from embedded tags 0xC0-0xD7 (tag value in the head, no argument to read), which is now its own case block that no longer has to fall into the ::store switch's "default" case to reach the same tag_pending = true; return true; outcome. - 0x98-0x9B and 0xB8-0xBB collapse into one case block each, always going through get_cbor_container_size() (harmless for 1/2-byte lengths, which already always fit). Verified byte-for-byte identical behavior before/after with a standalone probe covering embedded and multi-byte tags under all three tag_handler_t settings, a tag over a byte string (subtype path), truncated tag/length arguments of every width, and array/map lengths of every width, including the out_of_range.408 "excessive size" case: same exceptions, same messages, same chars_read, same successful results. Left the string/byte-string length ladders in get_cbor_string()/ get_cbor_binary() untouched, as noted in #5711 item 2, since #5325 is expected to touch them separately. Overlaps #5601 (adds a branch right above the embedded-tag case) and #5607 (touches the integer cases 0x18-0x1B, which share this ladder's shape in separate hunks). #5711 item 2 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Add leave_container() to match enter_container() Every container is opened through enter_container(), whose docs promise that a check placed there runs before every start event. The close side had no equivalent: the same "container_stack.pop_back(); dispatch to end_object() or end_array()" sequence was written out separately in BSON, CBOR, MessagePack, UBJSON/BJData and BON8, each copying the pattern of keeping an is_object flag around the pop_back() that would otherwise invalidate a reference to it. A check needed on close would have had to be added in five places, and a sixth copy could go unnoticed. Add leave_container() next to enter_container(), doing the same pop-then-dispatch, and replace the five sites with it. Each site keeps its own surrounding logic (BSON's check_bson_document_size() call before popping, MessagePack's is_object copy used again below, UBJSON/BJData's remaining-container handling after popping, BON8's top used again below); only the repeated pop/dispatch line pair is now shared. Verified all six binary-format unit suites and unit-regression2's deep-nesting tests (dependent count/reuse count and the bjdata ndarray depth cases) still pass, compiled with -Wall -Wextra and ASan/UBSan. Overlaps #5601, which is expected to add a sixth close site in its own skip loop; that site can route through leave_container() too once it lands. #5711 item 4 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Stop passing the input format to sax_parse() when the reader already has it binary_reader's constructor stores the format in the input_format member, and sax_parse(format, sax_, strict, tag_handler) took the same value again purely to dispatch on it. Every in-tree caller passed the same value both times (all 16 from_cbor/from_msgpack/from_ubjson/ from_bjdata/from_bon8/from_bson call sites in json.hpp, and the three public basic_json::sax_parse() overloads), so nothing was broken today, but a caller of the detail class directly (only reachable via JSON_PRIVATE_UNLESS_TESTED, as unit-bjdata.cpp already does) could pass a mismatched pair - say bjdata to the constructor and ubjson to sax_parse - and dispatch on one format while applying the other format's rules; the default-constructed input_format_t::json reader would additionally hit JSON_ASSERT(false) in exception_message() on its first error. Add sax_parse(json_sax_t*, bool, cbor_tag_handler_t) forwarding to the existing overload with the stored input_format, and switch every caller to it: the 16 from_*() sites (keeping their `// cppcheck-suppress[accessMoved]` comments) and the three basic_json::sax_parse() overloads, all of which already had the format available from their own `format` parameter. The four-argument overload is kept for anyone still calling it, now with JSON_ASSERT(format == input_format) so a mismatch fails immediately in a debug build (assert-enabled binaries, including the fuzzers and test suite) instead of misbehaving; verified with a probe that constructs a reader for one format and calls the explicit overload with another, which aborts on that assertion as expected. Removing or asserting against the constructor's input_format_t::json default, which would affect direct detail users, is left as a separate decision per #5711 item 5. Overlaps #5601, which is expected to add an AllowRecovery template parameter to sax_parse() and touch these same call sites in json.hpp. #5711 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate UBJSON/BJData signed-count handling, drop dead ndarray checks get_ubjson_size_value()'s 'i'/'I'/'l'/'L' cases each read a differently sized signed integer and then repeated the same "reject negative with error 113" check; only 'L' additionally checked value_in_range_of for the out_of_range.408 case. Any change to that error path had to be made four times. Add get_ubjson_signed_count<SignedType>(std::size_t&), doing the read, the negative check and the range check once, and route all four markers through it. The range check is a no-op for 'i'/'I'/'l' (their values always fit std::size_t) and only live for 'L' on a 32-bit std::size_t target, matching today's behavior exactly. In the ndarray dimension-product loop, the preceding loop already returns early on any zero dimension and result starts at 1, so `i > 0` in the pre-multiplication overflow check was always true, and `result == 0` in the post-multiplication check could not be reached either: two positive factors whose product does not overflow (as the pre-check already guarantees) cannot be zero. Drop the dead `i > 0 &&` and narrow the post-check to `result == npos`, the one case the pre-check cannot rule out (an exact, non-overflowing match with the sentinel reserved for unknown-size containers), with a comment explaining why. Verified byte-for-byte identical behavior before/after with a standalone probe covering negative counts for every marker, a matching positive count, and ndarray inputs, plus the full unit-ubjson and unit-bjdata suites (same assertion counts as before this change). Overlaps #5601 (rewrites the four parse_error calls and the overflow checks touched here) and #5607/#5707 (touch neighboring lines in the same functions). #5711 item 6 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Drop redundant format parameter and dummy float argument (review) binary_reader::sax_parse(format, ...) only ever had to equal the format given to the constructor, which it asserted. With every caller already on the format-less overload, remove the four-argument overload and dispatch on the stored input_format directly. binary_reader is a detail class, so this is not a public API change. get_ubjson_float_prefix() took a value only to deduce its type; make the type an explicit template argument instead. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
gregmarr
reviewed
Oct 4, 2026
…eader helpers The helpers (get_number, get_to, get_string, get_binary, get_bytes, emit_signed, emit_unsigned, emit_float, unexpect_eof, exception_message) are members of binary_reader, which already stores the format it was constructed with, so the parameter was redundant. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
gregmarr
approved these changes
Oct 4, 2026
nlohmann
added a commit
that referenced
this pull request
Oct 5, 2026
- binary_reader: emit_signed/emit_unsigned pass integers that do not fit the number types to emit_float as long double, because MSVC's <cmath> has no integer overloads of std::isfinite (C2668 'fpclassify'), from #5607 - scalar comparisons: the friend operators take the JSON type for their noexcept from their parameter, because MSVC 2015/2017 take basic_json as the class template there (C3203) and MSVC 2019 16.0 does not see member types or template parameters, from #5751 - unit-conversions2: skip the !is_nothrow_constructible static_assert for std::optional on MSVC 2017, which evaluates the conditional noexcept as true (C2607), from #5754 Signed-off-by: Niels Lohmann <mail@nlohmann.me>
nlohmann
added a commit
that referenced
this pull request
Oct 6, 2026
* Fix CI on develop after #5600, #5607, and #5755 - binary_reader: cast the result of -1 - number back to number_integer_t, because a number_integer_t narrower than int is promoted to int, which GCC's -Warith-conversion rejects (ci_test_gcc, ci_test_standards_gcc) - JSON_DELETE_DEPRECATED_FUNCTIONS: declare the deleted stream operators as function templates at namespace scope, because GCC < 5 rejects deleted friend functions ("can't initialize friend function") and Clang 7-9 report a redefinition when a class template has a deleted friend function - docs: give the examples of JSON_USE_OBJECTS_FOR_ENUM_KEYED_MAPS "Example:" titles and add the page to the docset (style_check) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Ignore libstdc++'s <format> sign change in the sanitizer job libstdc++ 14's <format> initializes a size_t parameter with -1 (GCC bug 119429), so every std::format call fails ci_test_clang_sanitizer under -fsanitize=integer (test-std-format_cpp20). Exclude only the implicit-integer-sign-change check and only that header via -fsanitize-ignorelist. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the AppVeyor (MSVC 2015-2019) build - binary_reader: emit_signed/emit_unsigned pass integers that do not fit the number types to emit_float as long double, because MSVC's <cmath> has no integer overloads of std::isfinite (C2668 'fpclassify'), from #5607 - scalar comparisons: the friend operators take the JSON type for their noexcept from their parameter, because MSVC 2015/2017 take basic_json as the class template there (C3203) and MSVC 2019 16.0 does not see member types or template parameters, from #5751 - unit-conversions2: skip the !is_nothrow_constructible static_assert for std::optional on MSVC 2017, which evaluates the conditional noexcept as true (C2607), from #5754 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Re-amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Avoid MSVC 2015's C4800 for enums with underlying type bool MSVC 2015 warns about any conversion to bool (C4800), even with an explicit cast, so the enum conversions from #5754 (#5671) failed the AppVeyor build with /WX. Convert to bool by comparing with zero via the new detail::bool_aware_static_cast, and keep doctest from printing the enum in the test. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Skip the #5650 allocator test on MSVC 2015 debug builds MSVC 2015's debug STL constructs the containers' debug proxies through the allocator in noexcept constructors, so countdown_allocator's failing construction crashes test-allocator (SIGSEGV) instead of throwing std::bad_alloc. Use the guard #5585 uses for the same reason. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Skip deleted-function detection checks on MSVC 2015 MSVC 2015 does not treat selecting a deleted function in decltype as a substitution failure, so the detection traits in unit-delete_deprecated_functions (#5755) and the integral-key checks in unit-element_access2 (#5657) report deleted overloads as callable there. Calling them still fails to compile. Skip those checks for _MSC_VER < 1910, and use the stream operators for real in the runtime section, so MSVC 2015 still compiles them with the macro set. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Split the Visual Studio 2017 AppVeyor jobs in two The VS 2017 jobs hit AppVeyor's 60-minute limit per job while still compiling the tests (77 of about 108 test targets after 58 minutes). They pass /std:c++17 for everything anyway, so build only the C++17 variant of each test (JSON_TestStandards=17), and split the unit test files across two jobs each with the new JSON_TestShard=<index>/<count> option, which keeps every <count>-th test file starting at <index>. The extra variants of single test files are built in shard 0 only. CMAKE_OPTIONS is no longer quoted in appveyor.yml, so that it can hold more than one option. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix std::terminate when converting std::optional with MSVC 2017 MSVC 2017 evaluates std::is_nothrow_assignable<json&, const T&> as true even if T's to_json throws, so to_json(json&, const std::optional<T>&) was noexcept there and the exception from #5642's test called std::terminate instead of propagating. Make that conversion never noexcept on MSVC 2017; all other compilers keep the exact condition. The static_asserts on the condition are skipped for MSVC 2017; the runtime check that the exception propagates still runs there. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
With number types narrower than the values in a binary document, every binary reader converted the decoded number implicitly:
The same happens for MessagePack, UBJSON, BJData, BSON and BON8, and already in 3.12.0. JSON text doesn't have this problem: the lexer stores an integer that fits neither integer type as
number_float_t(tested for #89 inunit-regression1.cpp, documented innumber_integer_t), and rejects a finite number that overflowsnumber_float_twithout_of_range.406.Changes
Three helpers in
binary_readerapply the lexer's rules to every number read from binary input:emit_signed():number_integer_tif the value fits, elsenumber_unsigned_tfor a non-negative value, elsenumber_float_temit_unsigned():number_unsigned_tif the value fits, elsenumber_float_temit_float():out_of_range.406if a finite value overflowsnumber_float_t; infinity and NaN are passed onThe range checks use the existing
detail::value_in_range_of, which is a compile-timetruewhen the types can't overflow, so the default types pay nothing.For consistency, a CBOR negative integer below the range of
number_integer_t(for example0x3Bfollowed by eight0xFFbytes, i.e. -2^64) is now stored asnumber_float_t, like-18446744073709551616in JSON text, instead of being rejected withparse_error.112(#5039). With the default number types, this is the only change in behavior.Docs:
cbor.md(negative integer note),exceptions.md(a binary-format example forout_of_range.406; theparse_error.112example removed),number_integer_t,number_unsigned_t,number_float_t.Tests
unit-binary_formats.cpp, for all six formats: integers that fit and that fit neither type, floats that fit, round toFLT_MAXor overflow (exactout_of_range.406message, anddiscardedwithallow_exceptions=false), infinity and NaN. 64 of its 112 assertions fail ondevelop.unit-cbor.cpp: the fix(cbor): reject overflowing negative integers #5039 overflow cases now expect floats, compared with the JSON text parse of the same numbers.developwith the default types: 1.08 million decodes of random and randomly mutated documents in all six formats. All 803 differences are CBOR inputs thatdeveloprejected and that now decode with a negative integer below INT64_MIN stored as a float.test-bon8cases, which need test data v3.2.0 that wasn't available locally. The new and changed tests compile without warnings with the CI's GCC 16 and clang flags at C++11 and C++20; before the fix, the new test failed the GCC flags with-Wconversionerrors in the readers.Public API
No signature changes. Behavior changes:
number_float_tinstead of wrapping around, and overflowing floats throwout_of_range.406(or yielddiscarded) instead of becoming infinity.number_integer_tdecodes asnumber_float_tinstead of throwingparse_error.112. fix(cbor): reject overflowing negative integers #5039 is not released yet, so compared with 3.12.0 this replaces a silently wrong value with the (rounded) right one. The release note for fix(cbor): reject overflowing negative integers #5039 needs to change accordingly.This PR was written by Claude Code.
🤖 Generated with Claude Code