Skip to content

build: C++26 by default (amdclang 23 / g++ 15), no module scanning - #91

Merged
bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/cxx26
Oct 5, 2026
Merged

bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/cxx26

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

The fork builds at C++26 by default, like the engine (CMAKE_CXX_STANDARD 26 in the top level and ggml/, CMAKE_CXX_STANDARD_REQUIRED kept, CMAKE_CXX_SCAN_FOR_MODULES OFF since TheRock has no clang-scan-deps). Per-target cxx_std_17 lines stay (they are minimums); HIP code objects keep their explicit -std=c++17. 16 files, +95/−36, upstream-style edits, no new files.

Source fixes the standard needed

  • src/models/models.h: declares the template <> specializations of the eagle3/dflash/t5 graph<...> constructors (and build_inp_embd_enc) after each struct. Since C++23 libstdc++'s make_unique is constexpr, so clang instantiates it inside build_arch_graph before the specialization was declared — the code was ill-formed NDR in every standard; this also fixes t5encoder.cpp.
  • common/chat.cpp: std::optional<json>(std::in_place, …) — the json → optional<json> conversion is ambiguous at C++26 with libstdc++ 16.
  • enum | enum of different types is ill-formed since C++26: casts in tools/mtmd/clip-graph.h, tools/mtmd/models/qwen3vl.cpp, tests/test-backend-ops.cpp (6 sites incl. var_to_str).
  • ggml/src/ggml-hrx/runtime/kernel-executable-cache.cpp: std::atomic<std::shared_ptr<…>> instead of the deprecated atomic_load_explicit on a shared_ptr.
  • ggml/src/ggml-backend-reg.cpp: path_from_u8() (from std::u8string when __cpp_lib_char8_t) replaces the 8 deprecated fs::u8path calls.
  • vendor/nlohmann/json.hpp:17557: is_trivial (deprecated) → is_trivially_default_constructible && is_trivially_copyable (removes 43 g++ warnings).

Silent behaviour change caught and neutralised. In C++26 std::to_string(float/double) formats like std::format("{}") instead of %f. It compiled, but test-jinja failed 5 cases ({{ -1.0 }} → -1, 42|float → 42; through the jinja trailing-zero trim 100.0 would print 1 and 0.0 an empty string followed by out.back() on it). All 17 call sites in 7 files now print %f as before (common/jinja/{string,value}.h, src/llama-impl.cpp, tools/mtmd/clip-impl.h, tools/llama-bench/llama-bench.cpp, tools/server/server-schema.cpp, tests/test-backend-ops.cpp), found by a temporary [[deprecated]] to_string(float) overload compiled through every TU with both compilers.

Checks (Strix Halo, performance mode; amdclang 23.0.0git TheRock + libstdc++ 16, g++ 15.2.0, cmake 4.2.3)

  • 0 warnings from fork sources on both compilers (same as the C++17 baseline).
  • test-backend-ops -b HRX0 ×2, C++17 and C++26 builds identical: MUL_MAT 325, MUL_MAT_ID 108, MUL_MAT_VEC_FUSION 57, ROPE 42, RMS_NORM 6, SOFT_MAX 8, CONCAT 6, SSM_CONV 47.
  • KLD Qwen3-0.6B c512 ×16: the saved logits of the C++17 and C++26 builds are byte-identical (PPL 26.1766 both).
  • 8 identical greedy requests: 1/8 distinct on Qwen3-0.6B and ZAYA1-8B, byte-identical to the C++17 build.
  • ctest (g++ CPU): same results at 17 and 26; test-jinja passes.
  • llama-bench HRX0, pp512/tg128 separate processes, 3 interleaved rounds × -r 3, medians of 9: Qwen3-0.6B −0.31% / −0.08%, Qwen3-8B +0.42% / +0.11%, GLM-4.7-Flash −0.12% / +0.05%, ZAYA1-8B −0.05% / +0.15% — all within noise.

Not in this PR: the private add-on's two GLM HIP harnesses needed C++23 (clang 23 cuda_wrappers/new vs libstdc++ 16 constexpr placement delete) — fixed in that repo. The -MD -MF on the fork's --cuda-device-only code-object rule gives 35 unused-argument warnings and probably never writes the depfiles — separate follow-up.

🤖 Generated with Claude Code

CMAKE_CXX_STANDARD 17 -> 26 in ggml/CMakeLists.txt and the top level
(CMAKE_CXX_STANDARD_REQUIRED kept); CMAKE_CXX_SCAN_FOR_MODULES OFF, so no
clang-scan-deps is needed (TheRock amdclang has none).

Source fixes for what C++26 rejects or changes:
- src/models/models.h: declare the explicit specializations of the
  eagle3/dflash/t5 graph members before first use. make_unique is constexpr
  since C++23, so clang instantiates it where build_arch_graph calls it,
  before the specialization was declared ("explicit specialization after
  instantiation"; ill-formed NDR in every standard).
- common/chat.cpp: build std::optional<json> in place (json -> optional<json>
  is ambiguous with libstdc++ 16 at C++26).
- tools/mtmd/clip-graph.h, tools/mtmd/models/qwen3vl.cpp,
  tests/test-backend-ops.cpp: cast ggml_scale_flag to an integer before
  combining it with ggml_scale_mode (enum|enum of different types is
  ill-formed since C++26, P2864).
- std::to_string(float/double) prints std::format("{}") since C++26 (P2587):
  test-jinja failed 5 cases ({{ -1.0 }} -> "-1", 42|float -> "42"; 100.0
  would print "1" and 0.0 an empty string through the trailing-zero trim).
  Every floating-point to_string call site (found by a deprecated-overload
  scan of all TUs with both compilers) now prints "%f" as before:
  common/jinja/{string,value}.h, src/llama-impl.cpp, tools/mtmd/clip-impl.h,
  tools/llama-bench/llama-bench.cpp, tools/server/server-schema.cpp,
  tests/test-backend-ops.cpp.
- ggml-hrx kernel-executable-cache.cpp: std::atomic<std::shared_ptr> instead
  of the deprecated atomic_load/store_explicit(shared_ptr*).
- ggml-backend-reg.cpp: path from a UTF-8 string via std::u8string instead of
  the deprecated fs::u8path (C++17 keeps u8path).
- vendor/nlohmann/json.hpp: std::is_trivial (deprecated in C++26) spelled as
  is_trivially_default_constructible && is_trivially_copyable.

Builds (strixhalo): amdclang 23.0.0git (TheRock) HRX build, with and without
the HIP add-on, and g++ 15.2.0 CPU build: 0 errors, 0 warnings from fork
sources (the add-on's own warnings and its two glm harness targets, which
hit clang 23 cuda_wrappers/new vs libstdc++ 16 constexpr placement delete at
C++26, are left to the add-on). ctest (g++): same results as C++17
(test-tokenizers-ggml-vocabs fails in both: LFS pointer vocab files).

Gates (performance mode, HRX0, C++17 tip f95f2db = A vs this = B):
- test-backend-ops -b HRX0 MUL_MAT 325/325, MUL_MAT_ID 108/108,
  MUL_MAT_VEC_FUSION 57/57, ROPE 42/42, RMS_NORM 6/6, SOFT_MAX 8/8,
  CONCAT 6/6, SSM_CONV 47/47; x2, both builds.
- Qwen3-0.6B Q4_K_M c512 16 chunks: saved logits byte-identical A vs B,
  PPL 26.1766 both, KLD 0.
- 8 identical greedy requests: 1/8 distinct on Qwen3-0.6B and ZAYA1-8B,
  output identical to the C++17 build.
- llama-bench medians of 9 (3 interleaved runs x -r 3), B vs A:
  Qwen3-0.6B pp512 -0.31% tg128 -0.08%; Qwen3-8B +0.42% / +0.11%;
  GLM-4.7-Flash -0.12% / +0.05%; ZAYA1-8B -0.05% / +0.15%.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit 6150ff0 into 1bit/hrx-vulkan-patched Oct 5, 2026
3 of 25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant