Skip to content

Reduce test suite runtime and run the Unicode tests everywhere - #5605

Merged
nlohmann merged 1 commit into
developfrom
claude/issue-5418-5426-94fef8
Oct 6, 2026
Merged

nlohmann merged 1 commit into
developfrom
claude/issue-5418-5426-94fef8

Conversation

@nlohmann

@nlohmann nlohmann commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Description

Fixes #5418 by addressing everything that #5519 did not already cover. All changes are test-only.

Credit: the first commit is @22elix3r's change from #5426 (pin unrelated bytes in the wrong-2nd/3rd-byte UTF-8 sweeps), cherry-picked unchanged with authorship preserved. Thanks! This PR supersedes #5426.

Changes

  1. Unicode ill-formed sweeps (unit-unicode2..5.cpp): the "wrong Nth byte" sections swept the invalid byte through all 256 values and every valid value of the other continuation bytes, which multiplied the work.

    • @22elix3r's commit from Cut Unicode ill-formed byte sweeps to one representative prefix #5426 (cherry-picked) fixed the other bytes to a single value in the wrong-2nd/3rd-byte sections of the 4-byte sequences. Since then, Fix dead ill-formed-fourth-byte UTF-8 test sections (byte3/byte4 typo) #5499 fixed the byte3/byte4 typo, so the wrong-4th-byte sections became live again as well.
    • A single value would lose coverage. With error_handler_t::ignore/replace, the serializer re-reads the invalid byte and decodes the bytes after it. Its decoder (detail::decode) tells three continuation classes apart: 0x80–8F, 0x90–9F and 0xA0–BF. The lexer's range checks use the same boundaries.
    • So the other positions now iterate over the first and last byte of each class within their valid range (utils::utf8_continuation_bytes). The invalid byte still takes all 256 values. This applies to all wrong-byte sections of unicode2–5.
    • Defining JSON_TEST_UTF8_EXHAUSTIVE switches back to the full Cartesian product. I checked that this reproduces the exact assertion counts from develop (unicode2: 5,400,433; unicode5: 17,248,461).
  2. unit-unicode1.cpp: \uxxxx escapes are now formatted by hand instead of with a std::stringstream per code point (~1.1 M calls). The JSON Pointer check still runs over every code point in all_unicode.json.

  3. unit-binary_formats.cpp: jeopardy.json (52 MB, ~20 s) moves into its own doctest::skip() test case. The four small corpus files (~1.5 s) now also run in JSON_FastTests configurations.

  4. unit-msgpack.cpp: the file mentioned JSON_HAS_CPP_17 for one std::byte test case, so the whole file was also compiled and run as test-msgpack_cpp17. That test case moves to unit-msgpack-cpp17.cpp, following the unit-items-cpp17.cpp precedent.

  5. Merge the Unicode tests into unit-unicode.cpp. They were split into five files in Refactor Unicode tests #2889 so they could run in parallel. They now take about 20 s together, so the split no longer pays off.

    • The merge removes four copies of the check helpers.
    • It removes the progress output, whose hard-coded totals had become wrong.
    • It removes the raised test-unicode4 timeout.
    • The merged test runs exactly the same 17,389,849 assertions as the five files together (minus the four duplicate "test data downloaded" checks).
  6. Stop excluding the Unicode tests. They were excluded in two ways:

    • doctest::skip() meant no JSON_FastTests configuration ran them. It is dropped.
    • --exclude-regex "test-unicode" kept them out of the ci.cmake compiler matrix (g++-4.8 … clang++-20), the default, icpc, icpx and nvhpc jobs, the Windows MinGW and clang-cl jobs, and AppVeyor Debug builds. This is removed everywhere except the Valgrind job, which now excludes them explicitly; before, skip() kept them out there.

What is still tested exhaustively

  • Every code point U+0000..U+10FFFF, excluding surrogates, as a \uxxxx escape or surrogate pair.
  • Every code point from all_unicode.json: parsing and the JSON Pointer escape roundtrip.
  • Every well-formed 1–4 byte UTF-8 sequence: parsing and all 9 dump() variants. These sections are unchanged.
  • Every truncated ("missing Nth byte") sequence. These sections are unchanged.
  • Every invalid byte value at every position after every lead byte, combined with every byte class that the lexer and serializer tell apart at the other positions.

Timings

The per-file rows are for the state before the merge. Unoptimized clang, --no-skip, macOS arm64, wall time:

test before after assertions before → after
unicode1 5 s 4.7 s 3.3 M → 3.3 M
unicode2 19.7 s 2.7 s 5.4 M → 1.2 M
unicode3 99 s 2.6 s 26.6 M → 2.4 M
unicode4 352 s* 9.4 s 93.7 M → 9.5 M
unicode5 66 s 1.4 s 17.2 M → 0.9 M
merged test-unicode ~540 s 19.4 s (ASan+UBSan: ~25 s) 149 M → 17.4 M
msgpack_cpp17 full suite a second time 5 assertions

* This run was in parallel with the others. The rest were measured serially.

Breaking changes

None. Only tests changed; the library and the amalgamation are unaffected.

Checklist

  • The changes are described in detail, both the what and why.
  • An existing issue is referenced.
  • Code coverage: unaffected (no library changes).
  • Documentation: n/a.
  • Amalgamation: n/a (tests only). astyle 3.4.13 is clean.

This PR was written by Claude Code on behalf of @nlohmann.

🤖 Generated with Claude Code

@nlohmann nlohmann changed the title Reduce test suite runtime: Unicode sweeps, unicode1, binary formats, msgpack C++17 Reduce test suite runtime and run the Unicode tests everywhere Sep 27, 2026
@nlohmann nlohmann added the review needed It would be great if someone could review the proposed changes. label Sep 27, 2026
Squashed onto develop from:
- Cut Unicode ill-formed byte sweeps to one representative prefix
- Pin unrelated bytes in the remaining ill-formed UTF-8 sweeps
- Speed up unit-unicode1
- Run the cheap binary format size tests unconditionally
- Compile unit-msgpack.cpp only once
- Check the JSON Pointer roundtrip for every code point again
- Cover every byte class in the ill-formed UTF-8 sweeps
- Merge the Unicode tests into unit-unicode.cpp
- Stop excluding the Unicode tests in CI

Co-authored-by: elix3r <157088510+22elix3r@users.noreply.github.com>
Signed-off-by: elix3r <157088510+22elix3r@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
@nlohmann
nlohmann force-pushed the claude/issue-5418-5426-94fef8 branch from 2c30264 to b872c0a Compare October 6, 2026 05:30
@nlohmann nlohmann added this to the Release 3.13.0 milestone Oct 6, 2026
@nlohmann nlohmann removed the review needed It would be great if someone could review the proposed changes. label Oct 6, 2026
@nlohmann
nlohmann merged commit 0c2329d into develop Oct 6, 2026
166 of 167 checks passed
@nlohmann
nlohmann deleted the claude/issue-5418-5426-94fef8 branch October 6, 2026 20:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Test suite: reduce runtime — 52 % is Unicode byte sweeps, plus redundant re-parsing in binary roundtrip tests

1 participant