Repository navigation
Bump HRX: llama.cpp f2099e9b7 (AMD's core with our ggml-hrx) - #336
Conversation
AMD's core changes plus our ggml-hrx backend (as of 522dab47) are the working configuration, measured 2026-10-06 in fresh build directories with the engine's own ExternalProject arguments (Release, amdclang, HRX_SOURCE_DIR=hrx-system 98d05d94), full test-backend-ops -b HRX0 and a GLM-4.7-Flash Q4_K_M 269-token prompt at -ub 512: AMD core + our ggml-hrx 1073/1073 (0 FAIL) 0 NaN, reply " Paris." AMD core + AMD ggml-hrx 1019/1071 (54 FAIL) 18 NaN, reply "??????" The 40 MUL_MAT failures and the engine#315 all-NaN logits are properties of AMD's ggml-hrx, not of the core sync, so this pins AMD's core with our backend. hrx-system stays 98d05d94: AMD's hrx-system does not export loomc_amdgpu_runtime_global_flags_t, which our loom-jit.cpp needs. Supersedes the staged pin in engine#329, which resolves the 15 conflicting ggml/src/ggml-hrx/ files toward AMD's side.
|
| check | result |
|---|---|
/health is 200 |
ok |
/v1/models names the alias |
ok |
| chat answers "Paris" | ok |
| streamed chat is SSE | ok - Content-Type: text/event-stream, 11 data: chunks |
| NaN events in the log | 0 |
Worth recording that the streamed check failed on my first attempt and the build was fine. My check helper embedded a JSON body with single quotes inside a double-quoted eval, so the request itself was mangled. Running the identical check against the synced build as a control showed it failing there too - and inspecting the raw bytes showed 11 SSE chunks with text/event-stream on both builds, which is what identified it as my harness rather than a regression. No difference between hybrid and synced on this path.
Scope, stated plainly: these are the OpenAI-compatible routes on this build's llama-server, not the engine's 1bit serve binary. If the literal 1bit serve path should be gated as well, that needs the engine built against this pin, which I have not done.
|
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
Independent reproduction of the decisive number — 1073/1073, 0 FAILI rebuilt this pin from scratch on a separate box state and got your number exactly. Built The I also re-verified the composition by tree hash rather than by eye, and it holds exactly: One trap worth knowing before anyone re-runs this gateMy first run of this reported The path the engine wants is
This does not affect your numbers — a CPU run would have made the two configurations identical, and they
|
Before merging: this pin does not fix
|
| pin | #315 correctness |
#314 compile |
|---|---|---|
this PR (f2099e9b7, our backend) |
1073/1073, 0 FAIL, 0 NaN | fails at D=256/24qh/4kvh/cap 1024 |
#329 (AMD's backend) |
54 FAIL, 18 NaN | compiles |
So neither candidate is clean, and that is a decision I can't make for you. This PR is still the right call —
a NaN logits failure is silent and corrupts output, a compile failure is loud and geometry-specific — but landing
it does not clear #314, and the release gate should not treat that item as closed.
Worth knowing for anyone testing this: a 1073/1073 test-backend-ops -b HRX0 run does not cover it. I rebuilt
this pin from scratch and got exactly 1073/1073 with HRX0 registered, because the corpus compiles per
specialization and none of those tests request that geometry — which is a GLM-4.7-Flash-shaped decode. A green
full-corpus run is not evidence that this kernel compiles.
Follow-up: the merged pin builds the engine's targets, but its own HRX test suite cannot compileFound while re-pinning the release-gate framework tree to this pin. Reporting it because it means Repro (at Cause. #include "dispatch/command-program-bindings.h"
#include "dispatch/command-program-diagnostics.h"
#include "dispatch/command-program-resolver.h"
#include "dispatch/command-program.h"
#include "dispatch/dispatch-scheduler.h"
#include "dispatch_registration/common/dispatch-gather-add.h"
#include "dispatch_registration/common/dispatch-symmetric-i4.h" // <-- missing
#include "dispatch_registration/dispatch-registry.h"But this PR pairs that core with our So the pairing is internally inconsistent for its own test target: the test file expects AMD's backend Impact, honestly scoped:
Options, in rough order of cost: build the needed targets only (what the engine does — worth documenting so the Not a request to change the pin — this PR is still the right one ( |
Corrected HRX pin bump. Supersedes the staged pin in #329, which takes AMD's
ggml-hrx— and that isrelease-blocking.
Why #329 as staged is not landable
Measured 2026-10-06 in fresh build directories with the engine's own
ExternalProjectarguments(
cmake/hrx.cmake: fresh-B, Release,amdclang,-DHRX_SOURCE_DIR=…), fulltest-backend-ops -b HRX0,and a GLM-4.7-Flash Q4_K_M 269-token prompt at
-ub 512:test-backend-ops -b HRX0ggml-hrx' Paris.'ggml-hrx??????The 40
MUL_MATfailures (iq1_s/iq1_m/iq3_xxs/mxfp4/pq2_0/ptq1_0/q2_*) and the engine#315all-NaN logits are properties of AMD's
ggml-hrx, not of the core sync. Resolving the 15 conflictingggml/src/ggml-hrx/files toward AMD's side imports #315 into the release pin.What this PR pins
third_party/llama.cpp522dab47f2099e9b7=1bit-MONSTER/llama.cpp1bit/amd-core-our-hrx—56c3c8a3's core with522dab47'sggml-hrxthird_party/hrx-system98d05d94loomc_amdgpu_runtime_global_flags_t, which ourloom-jit.cppneedsThe branch's composition is verified by tree, not by eye:
git diff --cached 522dab47 -- ggml/src/ggml-hrxis empty (backend tree hash
919cd999fin both) andgit diff --cached 56c3c8a3 -- . ':(exclude)ggml/src/ggml-hrx'is empty (nothing outside the backend moves).
Checks
Validation gate — do not un-draft until this passes
GitHub-hosted CI builds this without HRX, so CI cannot catch an HRX breakage. On Strix Halo, against this
head (
f2099e9b7/98d05d94):The
test-backend-ops -b HRX0line belongs in the PR body only if run against the llama.cpp subprojectwith
LLAMA_BUILD_TESTS=ON—ninja: error: unknown target 'test-backend-ops'in the engine tree, because theengine builds llama.cpp as an external project and does not build its tests (
PIN329-VALIDATION.md).Also up: #335 (
fix/bump-hrx-resolve-ours) makes the scheduled sync resolveggml-hrxtoward our line, sothis class of defect stops recurring instead of being fixed once by hand.