Repository navigation
Fix NUL bytes in UBJSON/BJData high-precision numbers - #5760
Conversation
| exception_message(concat("invalid number text: ", number_lexer.get_token_string()), "high-precision number"), nullptr)); | ||
| } | ||
|
|
||
| // 2026-10-05: A NUL must not terminate a length-delimited number payload; preserve the lexer diagnostic. |
There was a problem hiding this comment.
Would it be better performance-wise to catch this before pushing it into number_vector?
for (std::size_t i = 0; i < size; ++i)
{
get();
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof("number")))
{
return false;
}
if (JSON_HEDLEY_UNLIKELY(current == '\0')
{
return sax->parse_error(chars_read, ...
}
number_vector.push_back(static_cast<char>(current));
}
There was a problem hiding this comment.
Moved NUL detection into the read loop in e9b4643, so there's no separate payload scan. I kept error reporting after the full read to preserve the byte offset and lexer message; truncated payloads still report EOF (110). Added a truncated-NUL case to both suites. The UBJSON/BJData tests pass with both NUL macro settings.
| // 2026-10-05:长度限定的数字载荷不能把 NUL 当作真实结尾,沿用词法器的错误文本。 | ||
| if (JSON_HEDLEY_UNLIKELY(std::find(number_vector.begin(), number_vector.end(), '\0') != number_vector.end())) | ||
| // Reject NULs after a complete read to preserve EOF errors and lexer diagnostics. | ||
| if (JSON_HEDLEY_UNLIKELY(contains_nul)) |
There was a problem hiding this comment.
I kept error reporting after the full read to preserve the byte offset and lexer message; truncated payloads still report EOF (110). Added a truncated-NUL case to both suites. The UBJSON/BJData tests pass with both NUL macro settings.
It's up to @nlohmann, but I don't see the utility in preserving the lexer message. It causes extra work before reporting the error, and could change the error from NUL to EOF. I think the byte offset should be the exact location of the NUL byte.
There was a problem hiding this comment.
Agreed. In fecb084 the read loop returns the parse error as soon as it reads the NUL, so the byte offset is the NUL's position and the payload is no longer lexed first. The message follows the other binary reader errors: invalid number text; last byte: 0x00. A payload cut off after a NUL now reports that NUL (115) instead of end of input. The tests cover that case alongside the trailing and nested ones.
The UBJSON and BJData tests pass with both the split and amalgamated headers, and with JSON_STRICT_NUL_HANDLING=1. The previous commit was missing its sign-off, which is now added.
|
Your most recent commit did not contain the DCO signoff. |
a050f87 to
fecb084
Compare
|
The GCC/C++26 failure is the existing CBOR narrowing warning at binary_reader.hpp:668, already covered by #5764. I checked that get_cbor_negative_integer() is unchanged from this PR's base (4f69be8). The other five GCC standard jobs were cancelled by fail-fast. I'll keep this PR focused on the NUL fix and pick up the develop fix if a branch update is needed. |
|
The sanitizer job has a separate failure: 137/138 tests passed, with |
|
Thanks for tracking this down! Your analysis is right: libstdc++ 14's #5764 now also fixes this. It adds (This comment was written by Claude Code on my behalf.) |
| { | ||
| auto last_token = get_token_string(); | ||
| return sax->parse_error(chars_read, last_token, parse_error::create(115, chars_read, | ||
| exception_message(concat("invalid number text; last byte: 0x", last_token), "high-precision number"), nullptr)); |
There was a problem hiding this comment.
You know that the last byte is 0 because that's why you're throwing the error. No need to get it and convert it to a string.
There was a problem hiding this comment.
Changed in f2c82f7: the SAX token is now "00" and the error text is fixed, so this branch no longer calls get_token_string() or concat(). The error message and byte offsets are unchanged. The UBJSON/BJData suites pass with split and single headers, with both NUL macro settings.
gregmarr
left a comment
There was a problem hiding this comment.
One minor nit for some unnecessary work in an error path, but otherwise looks good.
|
The develop branch is green now. Please rebase a last time and we're should be good to go! |
Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>
Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>
Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>
Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>
Return the parse error from the read loop so the byte offset points at the NUL, and drop the separate check after lexing. Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>
2026-10-06: Use the known zero byte directly instead of formatting and concatenating it. Preserve the SAX token, error message, and byte offset. Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>
f2c82f7 to
bdb821a
Compare
|
The I reproduced this with MinGW GCC 15.1/C++11 on unmodified sources from develop at g++ -std=c++11 -Wswitch-enum -Werror=switch-enum -Iinclude -fsyntax-only tests/abi/diag/diag_off.cppThe split |
|
Opened #5770 for the remaining |
|
Thanks! |
Reject NUL bytes inside length-delimited UBJSON/BJData high-precision (
H) number payloads.H i 3 1 NUL xused to be accepted as the unsigned integer1, because the lexer stops at the NUL and never sees the rest of the payload. It now producesparse_error.115.The payload read loop returns the error as soon as it reads a NUL, so the byte offset points at the NUL and the payload is not lexed. The message follows the other binary reader errors:
invalid number text; last byte: 0x00. A payload that is cut off after a NUL reports the NUL rather than end of input.Regression cases in both format suites cover hidden bytes after a NUL, a trailing NUL, a truncated payload containing a NUL, a nested array, non-throwing parsing, and a valid number.
single_include/nlohmann/json.hppcarries the same change.Fixes #5753.
Validation on Windows x86_64, MinGW GCC 15.1.0, C++11:
The new cases fail without the fix, because the malformed payloads are accepted.
The UBJSON and BJData suites pass with both
include/andsingle_include/, and withJSON_STRICT_NUL_HANDLING=1.Coverage was not measured locally.
The changes are described in detail, both the what and why.
An existing issue is referenced.
Code coverage remained at 100% (not measured locally; awaiting CI).
Source code is amalgamated.
OSS-Fuzz and documentation checklist items are not applicable to this change.