From 750c266a7f9433e50343340e5e1c02780f1ef2a2 Mon Sep 17 00:00:00 2001 From: 1bit release bot Date: Tue, 6 Oct 2026 04:45:38 -0300 Subject: [PATCH 1/8] Bump the HRX pair to the full AMD sync (llama.cpp + hrx-system) Moves both pinned submodules onto AMD's tested pair, merged with our commits. third_party/llama.cpp 522dab47 -> b04c4e95 = AMD b802a507 merged with our 159 commits, then llama.cpp#94 on top, so it carries the engine#315 NaN fix as well as AMD's kernel-corpus refactor. third_party/hrx-system 98d05d94 -> 563afce1 = AMD 40a1a36c merged with our libhrx device knobs (HRX_AQL_BLOCK_SIZE, HRX_COMMAND_BUFFER_MODE). The three Loom target-info conflicts resolved to AMD's side: our commit there was a backport of a fix AMD has since landed. DO NOT MERGE UNTIL VALIDATED ON STRIX HALO. GitHub-hosted CI builds this WITHOUT HRX, so CI cannot catch an HRX breakage: cmake -B build -G Ninja -DONEBIT_HRX=ON && cmake --build build --target onebit tests/serve_e2e.sh build/1bit hrx tests/serve_e2e.sh build/1bit cpu test-backend-ops -b HRX0 --- third_party/hrx-system | 2 +- third_party/llama.cpp | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/third_party/hrx-system b/third_party/hrx-system index 98d05d9..563afce 160000 --- a/third_party/hrx-system +++ b/third_party/hrx-system @@ -1 +1 @@ -Subproject commit 98d05d94a9f9e405683f8aeafc037f5ace0abe4a +Subproject commit 563afce154d233c82361bbd335b489f20709cdcd diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 522dab4..b04c4e9 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 522dab4789ce84d32f68ecee722f50b1fe4a3a84 +Subproject commit b04c4e95cd607e4216d8a189aa6261dc2c526262 From e9435774ea911133b8123e38d7d8a2d45cfe5338 Mon Sep 17 00:00:00 2001 From: bong-water-water-bong <277547417+bong-water-water-bong@users.noreply.github.com> Date: Tue, 6 Oct 2026 04:49:37 -0300 Subject: [PATCH 2/8] registry: record the bumped llama.cpp pin (engine#329) The bump moved third_party/llama.cpp to b04c4e95 but left registry/architectures.json recording 522dab4, so ctest registry_pins - which the build job runs, not gated on ONEBIT_HRX - fails on the pin mismatch. The registry's inputs are byte-identical between the two pins (convert_hf_to_gguf.py, conversion/, gguf-py/gguf/constants.py, src/llama-arch.cpp, src/llama-model.cpp), so only sources["llama.cpp (hrx)"] moves; the architecture set and counts are unchanged. Co-authored-by: agent --- registry/architectures.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/registry/architectures.json b/registry/architectures.json index 0aaac4f..4af8413 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -1,7 +1,7 @@ { "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { - "llama.cpp (hrx)": "522dab4789ce84d32f68ecee722f50b1fe4a3a84" + "llama.cpp (hrx)": "b04c4e95cd607e4216d8a189aa6261dc2c526262" }, "counts": { "hrx": 292, From 4847858fb90af33e842bd474812f5f079d8983b7 Mon Sep 17 00:00:00 2001 From: 1bit release bot Date: Tue, 6 Oct 2026 04:50:30 -0300 Subject: [PATCH 3/8] Bump HRX: llama.cpp 51e538fa046d (adds the decode-split 2048 cap) AMD's refactor removed the multipass reduce provider, so the merged corpus only covers key_value_token_capacity up to 2048. Our inherited constant was still 32768, which would offer the decode-split dispatch for capacities with no provider, making the selector reject every candidate ("all_rejected") and failing the whole decode. Capped at 2048. --- registry/architectures.json | 2 +- third_party/llama.cpp | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/registry/architectures.json b/registry/architectures.json index 4af8413..8fb7050 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -1,7 +1,7 @@ { "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { - "llama.cpp (hrx)": "b04c4e95cd607e4216d8a189aa6261dc2c526262" + "llama.cpp (hrx)": "51e538fa046d6fcb09c3ecadd68d2103616055d9" }, "counts": { "hrx": 292, diff --git a/third_party/llama.cpp b/third_party/llama.cpp index b04c4e9..51e538f 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit b04c4e95cd607e4216d8a189aa6261dc2c526262 +Subproject commit 51e538fa046d6fcb09c3ecadd68d2103616055d9 From 4e775e3b5d3b6c7fea16014627f8c2858f068d44 Mon Sep 17 00:00:00 2001 From: bong-water-water-bong <277547417+bong-water-water-bong@users.noreply.github.com> Date: Tue, 6 Oct 2026 07:50:12 -0300 Subject: [PATCH 4/8] fix(hrx): pin the repaired llama.cpp tree and the planner fix Engine#329 as it stands does not build: the AMD sync auto-merged our kernels into AMD's refactored corpus, which left duplicate SSA names in shared motifs, 47 manifest paths with no file, 19 exports with no source, and dispatch-mul-mat.cpp with our body and AMD's include block. Point third_party/llama.cpp at 1bit/amd-sync-corpus-repair (420b4e6f) and third_party/hrx-system at 1bit/amd-sync-wait-plan (3d0cca89a0), and move the registry source pin with them. Both are one commit ahead of what this PR pinned, so the pin check still only moves forward. With these, cmake --build build --target onebit succeeds (0 errors) and serve_e2e.sh passes on HRX for Qwen3-0.6B. test-backend-ops -b HRX0 still reports 989 OK / 84 FAIL and gpt-oss/GLM do not run, so this makes the gate measurable, not green. --- registry/architectures.json | 2 +- third_party/hrx-system | 2 +- third_party/llama.cpp | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/registry/architectures.json b/registry/architectures.json index 8fb7050..825ceae 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -1,7 +1,7 @@ { "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { - "llama.cpp (hrx)": "51e538fa046d6fcb09c3ecadd68d2103616055d9" + "llama.cpp (hrx)": "420b4e6fe6a9bfce0ee31cb671933a1fee00ca44" }, "counts": { "hrx": 292, diff --git a/third_party/hrx-system b/third_party/hrx-system index 563afce..3d0cca8 160000 --- a/third_party/hrx-system +++ b/third_party/hrx-system @@ -1 +1 @@ -Subproject commit 563afce154d233c82361bbd335b489f20709cdcd +Subproject commit 3d0cca89a00d27d58e5d2345eb35c9f263ed83b1 diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 51e538f..420b4e6 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 51e538fa046d6fcb09c3ecadd68d2103616055d9 +Subproject commit 420b4e6fe6a9bfce0ee31cb671933a1fee00ca44 From 99a7a751594021e78279ccb52593fa99ce0ce5e1 Mon Sep 17 00:00:00 2001 From: bong-water-water-bong <277547417+bong-water-water-bong@users.noreply.github.com> Date: Tue, 6 Oct 2026 08:09:24 -0300 Subject: [PATCH 5/8] fix(hrx): pin llama.cpp 39588ef9, which stops the JIT cache discarding launch programs The previous pin (420b4e6f) still had runtime/loom-jit-disk-cache.cpp storing entries without the host-side launch program, so every cache hit was refused as "compiled ABI does not match manifest". 39588ef9 skips caching those results and bumps the cache version. With it, GLM-4.7-Flash-Q4_K_M runs on HRX end to end: pp512 896.84 t/s, tg32 28.05 t/s. serve_e2e on Qwen3-0.6B still passes. test-backend-ops -b HRX0 is unchanged at 989 OK / 84 FAIL. --- registry/architectures.json | 2 +- third_party/llama.cpp | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/registry/architectures.json b/registry/architectures.json index 825ceae..eda5a4e 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -1,7 +1,7 @@ { "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { - "llama.cpp (hrx)": "420b4e6fe6a9bfce0ee31cb671933a1fee00ca44" + "llama.cpp (hrx)": "39588ef99ebc852b7bf7c22f6d6ae7754b0f3002" }, "counts": { "hrx": 292, diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 420b4e6..39588ef 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 420b4e6fe6a9bfce0ee31cb671933a1fee00ca44 +Subproject commit 39588ef99ebc852b7bf7c22f6d6ae7754b0f3002 From c7184973489da520cb9ee824e6add5df1ed40aef Mon Sep 17 00:00:00 2001 From: bong-water-water-bong <277547417+bong-water-water-bong@users.noreply.github.com> Date: Tue, 6 Oct 2026 08:39:15 -0300 Subject: [PATCH 6/8] fix(hrx): pin llama.cpp 56c3c8a3, which routes mul_mat_id through our kernel again 39588ef9 dropped our MoE matmul in favour of AMD's tiled/skinny pair, which mishandles n = 5, 17, 32 and 129. 56c3c8a3 points both matcher constants back at our ggml_mul_mat_id_f32_f32_wmma. test-backend-ops -b HRX0 -o MUL_MAT_ID 78/108 -> 108/108 GLM-4.7-Flash, default -ub 512 decode failure -> pp512 902.21 t/s, tg32 27.40 t/s The GLM row is engine#315: the all-NaN logits are gone and the -ub 128 workaround is no longer needed. --- registry/architectures.json | 2 +- third_party/llama.cpp | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/registry/architectures.json b/registry/architectures.json index eda5a4e..92cd7ce 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -1,7 +1,7 @@ { "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { - "llama.cpp (hrx)": "39588ef99ebc852b7bf7c22f6d6ae7754b0f3002" + "llama.cpp (hrx)": "56c3c8a3547712f373ce6b89f8698f31d16bef80" }, "counts": { "hrx": 292, diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 39588ef..56c3c8a 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 39588ef99ebc852b7bf7c22f6d6ae7754b0f3002 +Subproject commit 56c3c8a3547712f373ce6b89f8698f31d16bef80 From 3758760c649341f82c267d7857fd2486d33d9b38 Mon Sep 17 00:00:00 2001 From: bong-water-water-bong <277547417+bong-water-water-bong@users.noreply.github.com> Date: Tue, 6 Oct 2026 08:44:06 -0300 Subject: [PATCH 7/8] fix(hrx): pin llama.cpp 89d4f17c, which also keeps our binary_f32 kernel ggml_binary_f32 is a name both trees use and the merged corpus held AMD's file under it, so the ADD path ran AMD's implementation. 89d4f17c repoints that export at our pre-merge source. test-backend-ops -b HRX0 -o ADD 5 failures -> 22/22 passed GLM-4.7-Flash pp512/tg32 885.17 / 27.30 t/s, unchanged Not fixed this way, and recorded as such: ggml_get_rows_f32 (the same repoint makes GET_ROWS worse, 7 -> 24) and MUL_MAT (40 -> 42 alone). --- registry/architectures.json | 2 +- third_party/llama.cpp | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/registry/architectures.json b/registry/architectures.json index 92cd7ce..4664eec 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -1,7 +1,7 @@ { "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { - "llama.cpp (hrx)": "56c3c8a3547712f373ce6b89f8698f31d16bef80" + "llama.cpp (hrx)": "89d4f17c866e658f444c91fc906b76a2b2d2ca23" }, "counts": { "hrx": 292, diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 56c3c8a..89d4f17 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 56c3c8a3547712f373ce6b89f8698f31d16bef80 +Subproject commit 89d4f17c866e658f444c91fc906b76a2b2d2ca23 From fd6bed003a66fcaaffd826bf72abc734d8d4b2d3 Mon Sep 17 00:00:00 2001 From: bong-water-water-bong <277547417+bong-water-water-bong@users.noreply.github.com> Date: Tue, 6 Oct 2026 08:44:54 -0300 Subject: [PATCH 8/8] revert(hrx): pin llama.cpp 56c3c8a3 again - the binary_f32 change is neutral The full gate is 1019 passed / 54 failed both before and after 89d4f17c, and my '5 -> 22/22' comparison was invalid: -o ADD selects 22 cases of a different class while the 5 [ADD]-class failures are the MUL_MAT_VEC_FUSION ones, which still fail. Keeping the pin on the two changes that have a measured effect (the JIT-cache fix and the mul_mat_id routing fix). --- registry/architectures.json | 2 +- third_party/llama.cpp | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/registry/architectures.json b/registry/architectures.json index 4664eec..92cd7ce 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -1,7 +1,7 @@ { "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { - "llama.cpp (hrx)": "89d4f17c866e658f444c91fc906b76a2b2d2ca23" + "llama.cpp (hrx)": "56c3c8a3547712f373ce6b89f8698f31d16bef80" }, "counts": { "hrx": 292, diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 89d4f17..56c3c8a 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 89d4f17c866e658f444c91fc906b76a2b2d2ca23 +Subproject commit 56c3c8a3547712f373ce6b89f8698f31d16bef80