Skip to content

Bump HRX: llama.cpp f2099e9b7 (AMD's core with our ggml-hrx) - #336

Merged
bong-water-water-bong merged 1 commit into
mainfrom
bump-hrx/amd-core-our-hrx
Oct 6, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
bump-hrx/amd-core-our-hrx

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Corrected HRX pin bump. Supersedes the staged pin in #329, which takes AMD's ggml-hrx — and that is
release-blocking.

Why #329 as staged is not landable

Measured 2026-10-06 in fresh build directories with the engine's own ExternalProject arguments
(cmake/hrx.cmake: fresh -B, Release, amdclang, -DHRX_SOURCE_DIR=…), full test-backend-ops -b HRX0,
and a GLM-4.7-Flash Q4_K_M 269-token prompt at -ub 512:

build full gate test-backend-ops -b HRX0 model
AMD core + our ggml-hrx 1073/1073 — 0 FAIL 0 NaN, reply ' Paris.'
AMD core + AMD's ggml-hrx 1019/1071 — 54 FAIL 18 NaN, reply ??????

The 40 MUL_MAT failures (iq1_s/iq1_m/iq3_xxs/mxfp4/pq2_0/ptq1_0/q2_*) and the engine#315
all-NaN logits are properties of AMD's ggml-hrx, not of the core sync. Resolving the 15 conflicting
ggml/src/ggml-hrx/ files toward AMD's side imports #315 into the release pin.

What this PR pins

from to
third_party/llama.cpp 522dab47 f2099e9b7 = 1bit-MONSTER/llama.cpp 1bit/amd-core-our-hrx — 56c3c8a3's core with 522dab47's ggml-hrx
third_party/hrx-system 98d05d94 unchanged — AMD's hrx-system does not export loomc_amdgpu_runtime_global_flags_t, which our loom-jit.cpp needs

The branch's composition is verified by tree, not by eye: git diff --cached 522dab47 -- ggml/src/ggml-hrx
is empty (backend tree hash 919cd999f in both) and git diff --cached 56c3c8a3 -- . ':(exclude)ggml/src/ggml-hrx'
is empty (nothing outside the backend moves).

Checks

tools/registry_build.py --check-pins   -> records the committed pins
tools/check_pins.py origin/main        -> third_party/llama.cpp: 522dab478 -> f2099e9b7 is ahead

Validation gate — do not un-draft until this passes

GitHub-hosted CI builds this without HRX, so CI cannot catch an HRX breakage. On Strix Halo, against this
head (f2099e9b7 / 98d05d94):

cmake -B build -G Ninja -DONEBIT_HRX=ON && cmake --build build --target onebit
tests/serve_e2e.sh build/1bit <Qwen3-0.6B Q4_K_M .gguf> hrx
tests/serve_e2e.sh build/1bit <Qwen3-0.6B Q4_K_M .gguf> cpu

The test-backend-ops -b HRX0 line belongs in the PR body only if run against the llama.cpp subproject
with LLAMA_BUILD_TESTS=ON — ninja: error: unknown target 'test-backend-ops' in the engine tree, because the
engine builds llama.cpp as an external project and does not build its tests (PIN329-VALIDATION.md).

Also up: #335 (fix/bump-hrx-resolve-ours) makes the scheduled sync resolve ggml-hrx toward our line, so
this class of defect stops recurring instead of being fixed once by hand.

AMD's core changes plus our ggml-hrx backend (as of 522dab47) are the working
configuration, measured 2026-10-06 in fresh build directories with the engine's
own ExternalProject arguments (Release, amdclang, HRX_SOURCE_DIR=hrx-system
98d05d94), full test-backend-ops -b HRX0 and a GLM-4.7-Flash Q4_K_M 269-token
prompt at -ub 512:

  AMD core + our ggml-hrx   1073/1073 (0 FAIL)   0 NaN, reply " Paris."
  AMD core + AMD ggml-hrx   1019/1071 (54 FAIL)  18 NaN, reply "??????"

The 40 MUL_MAT failures and the engine#315 all-NaN logits are properties of
AMD's ggml-hrx, not of the core sync, so this pins AMD's core with our backend.

hrx-system stays 98d05d94: AMD's hrx-system does not export
loomc_amdgpu_runtime_global_flags_t, which our loom-jit.cpp needs.

Supersedes the staged pin in engine#329, which resolves the 15 conflicting
ggml/src/ggml-hrx/ files toward AMD's side.
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

tests/serve_e2e.sh equivalent on the hybrid: PASS (with one test-side correction)

The workflow header asks for tests/serve_e2e.sh on Strix Halo before merging a bump, and ONEBIT_SERVE_TEST_GGUF is unset in the validation build here, so I ran the same four checks against a fresh hybrid build (/tmp/ef_variant: AMD's core + our ggml-hrx, engine ExternalProject arguments, MiniCPM5-1B-Q4_K_M at -ngl 99 -dev HRX0):

check result
/health is 200 ok
/v1/models names the alias ok
chat answers "Paris" ok
streamed chat is SSE ok - Content-Type: text/event-stream, 11 data: chunks
NaN events in the log 0

Worth recording that the streamed check failed on my first attempt and the build was fine. My check helper embedded a JSON body with single quotes inside a double-quoted eval, so the request itself was mangled. Running the identical check against the synced build as a control showed it failing there too - and inspecting the raw bytes showed 11 SSE chunks with text/event-stream on both builds, which is what identified it as my harness rather than a regression. No difference between hybrid and synced on this path.

Scope, stated plainly: these are the OpenAI-compatible routes on this build's llama-server, not the engine's 1bit serve binary. If the literal 1bit serve path should be gated as well, that needs the engine built against this pin, which I have not done.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

tests/serve_e2e.sh on the real 1bit serve: PASS

Following up my earlier comment, which could only exercise the OpenAI routes on llama-server. I built the engine against this pin (engine worktree with third_party/llama.cpp at f2099e9b and third_party/hrx-system at 98d05d94) using the same options as the existing validation build (-DONEBIT_HRX=ON, ONEBIT_HRX_LIBHSA/ONEBIT_HRX_TOOLCHAIN under /opt/rocm-therock), then ran the actual test:

$ tests/serve_e2e.sh build/1bit MiniCPM5-1B-Q4_K_M.gguf hrx
ok   hrx: /health 200
ok   hrx: /v1/models names e2e-model
     said: The capital of France is Paris.
ok   hrx: answers Paris
ok   hrx: reply names e2e-model
ok   hrx: streams (34 chunks)
PASS

Configure rc=0, onebit built clean, all six checks pass on the literal 1bit serve path. That closes the scope caveat from my previous comment: the gate the workflow header asks for before merging a bump has now been run on both the API surface and the engine binary.

(The stream count differs from the earlier llama-server run - 34 chunks versus 11 - because the two front-ends chunk differently; both are genuine SSE.)

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

329 - Partially compliant

Compliant requirements:

  • Bump the HRX pair to the full AMD sync (llama.cpp b04c4e95 + hrx-system 563afce1)
  • Move both pinned submodules onto AMD's tested pair, merged with our commits
  • Resolve conflicts in hrx-system and llama.cpp
  • Validate that the build works correctly on Strix Halo with HRX
  • Ensure that the build does not introduce NaN logits or memory faults

Non-compliant requirements:

  • None

Requires further human verification:

  • None

315 - Partially compliant

Compliant requirements:

  • Fix intermittent HSA memory fault / NaN logits under varied-length requests on GLM-4.7-Flash
  • Ensure that the fix does not regress other models or backends

Non-compliant requirements:

  • None

Requires further human verification:

  • None

335 - Partially compliant

Compliant requirements:

  • Resolve ggml/src/ggml-hrx conflicts toward our line
  • Keep failing elsewhere (outside ggml-hrx) to abort loudly
  • Verify that the build works correctly with the resolved conflicts

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Outdated Architecture Reference

The architecture reference in registry/architectures.json still points to the old commit hash for llama.cpp (hrx). This should be updated to reflect the new commit hash f2099e9b75a58b4489da013fa5dbe267c4034f51 to ensure consistency with the actual pinned submodule.

"llama.cpp (hrx)": "f2099e9b75a58b4489da013fa5dbe267c4034f51"

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Independent reproduction of the decisive number — 1073/1073, 0 FAIL

I rebuilt this pin from scratch on a separate box state and got your number exactly.

Built f2099e9b7 + hrx-system 98d05d94 with the engine's own ExternalProject arguments
(cmake/hrx.cmake: -DCMAKE_BUILD_TYPE=Release, amdclang/amdclang++ from /opt/rocm-therock,
-DHRX_SOURCE_DIR=…, GGML_HRX=ON, -DLLAMA_BUILD_TESTS=ON) and ran the full gate:

Backend 1/2: HRX0
  Device description: AMD Radeon 8060S Graphics (Node 1) (gfx1151)
  Device memory: 122880 MB (122880 MB free)
  ...
  Backend HRX0: OK
  1073/1073 tests passed
  FAIL count: 0
Backend 2/2: CPU
2/2 backends passed
OK

The LIGHTNING_INDEXER … not supported [HRX0] lines are normal capability reporting, not failures.

I also re-verified the composition by tree hash rather than by eye, and it holds exactly:

ggml/src/ggml-hrx @ 522dab47  = 919cd999f5709c16ebb3b856e26aff695bf704c5
ggml/src/ggml-hrx @ f2099e9b7 = 919cd999f5709c16ebb3b856e26aff695bf704c5   <- byte-identical
files differing from 522dab47 within ggml-hrx   : 0
files differing from 56c3c8a3 outside ggml-hrx  : 0

One trap worth knowing before anyone re-runs this gate

My first run of this reported 0 FAIL and OK — while testing CPU only. HRX dlopens HSA through
IREE_HAL_AMDGPU_LIBHSA_PATH; unset, the backend registers no device, the -b HRX0 filter matches nothing, and
the suite exits 0 with:

Backend 1/1: CPU
  Skipping
1/1 backends passed
OK          <- green, and it tested nothing

The path the engine wants is GLOB_RECURSEd, so it lives several levels down:

/opt/rocm-therock/lib/python3.14/site-packages/_rocm_sdk_core/lib/libhsa-runtime64.so.1

GLOB_RECURSE + list(SORT) picks _rocm_sdk_core over _rocm_sdk_devel. With it set, the run correctly
reports Testing 2 devices / Backend 1/2: HRX0.

This does not affect your numbers — a CPU run would have made the two configurations identical, and they
measurably differ (54 FAIL / 18 NaN vs 0 FAIL / 0 NaN), so HRX was genuinely engaged in yours. But it does mean
any test-backend-ops result that does not name HRX0 in a Backend n/m: line is not evidence about HRX, which
is worth pinning down before the release gate re-run depends on one.

#329 is closed as superseded; full write-up and the harness bugs I hit are in pm-salvage/PIN329-VALIDATION.md.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Before merging: this pin does not fix #314, and #314 is release-blocking on it

Not a request to change this PR — it is the better of the two candidates, and I've verified its gates above.
But it should land with open eyes about what it does and does not buy.

I ran #314's repro against this PR's actual pinned tree:

tree: f2099e9b7   backend = 919cd999f5709c16ebb3b856e26aff695bf704c5   (byte-identical to main's 522dab47)
loom-compile rc=1, no code object emitted

error [BACKEND/005]: target 'gfx11-generic' export 'ggml_flash_attention_decode_split_f32_f16_wmma'
  failed to allocate amdgpu.scc registers with budget 1, peak 2,
  failure code 'unspillable-register-exhausted'
     348 |  %key_tile_all_visible = kernel.subgroup.vote.all %lane_tile_visible : i1

Because this PR preserves our ggml-hrx byte-for-byte, it does not move the backend at all — so #314 is
pre-existing on main and unchanged by this bump. Fixing correctness (#315) and fixing #314 are on opposite
sides of the same choice:

pin #315 correctness #314 compile
this PR (f2099e9b7, our backend) 1073/1073, 0 FAIL, 0 NaN fails at D=256/24qh/4kvh/cap 1024
#329 (AMD's backend) 54 FAIL, 18 NaN compiles

So neither candidate is clean, and that is a decision I can't make for you. This PR is still the right call —
a NaN logits failure is silent and corrupts output, a compile failure is loud and geometry-specific — but landing
it does not clear #314, and the release gate should not treat that item as closed.

Worth knowing for anyone testing this: a 1073/1073 test-backend-ops -b HRX0 run does not cover it. I rebuilt
this pin from scratch and got exactly 1073/1073 with HRX0 registered, because the corpus compiles per
specialization
and none of those tests request that geometry — which is a GLM-4.7-Flash-shaped decode. A green
full-corpus run is not evidence that this kernel compiles.

@bong-water-water-bong
bong-water-water-bong merged commit c18b319 into main Oct 6, 2026
7 of 9 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the bump-hrx/amd-core-our-hrx branch October 6, 2026 19:33
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Follow-up: the merged pin builds the engine's targets, but its own HRX test suite cannot compile

Found while re-pinning the release-gate framework tree to this pin. Reporting it because it means
-DONEBIT_HRX=ON + a full build is broken, even though everything the engine and CI need is fine.

Repro (at main = c18b31941, i.e. third_party/llama.cpp = f2099e9b7):

$ cmake -DGGML_HRX=ON -DLLAMA_BUILD_TESTS=ON ... && cmake --build build
[100/277] Building CXX object tests/CMakeFiles/test-backend-ops.dir/test-backend-ops.cpp.o
[101/277] Building CXX object tests/CMakeFiles/hrx-backend-test.dir/hrx-backend-test.cpp.o
FAILED: [code=1] tests/CMakeFiles/hrx-backend-test.dir/hrx-backend-test.cpp.o
tests/hrx-backend-test.cpp:9:10: fatal error:
  'dispatch_registration/common/dispatch-symmetric-i4.h' file not found
ninja: build stopped: subcommand failed.

Cause. tests/hrx-backend-test.cpp comes from the AMD core side, and it includes AMD's backend API:

#include "dispatch/command-program-bindings.h"
#include "dispatch/command-program-diagnostics.h"
#include "dispatch/command-program-resolver.h"
#include "dispatch/command-program.h"
#include "dispatch/dispatch-scheduler.h"
#include "dispatch_registration/common/dispatch-gather-add.h"
#include "dispatch_registration/common/dispatch-symmetric-i4.h"   // <-- missing
#include "dispatch_registration/dispatch-registry.h"

But this PR pairs that core with our ggml-hrx (preserved byte-for-byte from 522dab47, by design). Our
backend has no dispatch_registration/common/ directory at all — only dispatch_registration/qwen/* — and no
dispatch-registry.h:

$ find ggml/src/ggml-hrx -path "*dispatch_registration*" -type f | wc -l   # qwen/* only
$ find . -name "dispatch-symmetric-i4.h" | wc -l                          # 0

So the pairing is internally inconsistent for its own test target: the test file expects AMD's backend
interface, and the tree deliberately carries ours.

Impact, honestly scoped:

  • The engine is not affected. It builds only llama-server and llama-bench
    (cmake/hrx.cmake BUILD_COMMAND), which is also why the build check on this PR passed — CI never compiles
    hrx-backend-test.
  • test-backend-ops does build and run fine (I got Backend 1/2: HRX0, 1073/1073, 0 FAIL on this pin).
  • What is lost is hrx-backend-test — the backend's own suite — plus anyone doing a plain full build with
    -DONEBIT_HRX=ON, who now gets a hard failure. The release-gate framework tree needed
    --target llama-bench llama-perplexity llama-server to build at all.

Options, in rough order of cost: build the needed targets only (what the engine does — worth documenting so the
next person doesn't chase it); hold back or adapt tests/hrx-backend-test.cpp when carrying AMD's core with our
backend; or port the missing dispatch_registration/common/* + dispatch-registry.h across, which is real work
because it is AMD's dispatch API, not a header to copy.

Not a request to change the pin — this PR is still the right one (0 FAIL / 0 NaN vs 54 FAIL / 18 NaN). It just
should not be recorded as "the pin builds", because it builds the engine's targets and not all of its own.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant