Skip to content

HRX: follow AMD's live ggml-hrx, pinned to AMD's tested pair and kept current - #11

Merged
bong-water-water-bong merged 2 commits into
mainfrom
hrx-amd-live
Sep 23, 2026
Merged

bong-water-water-bong merged 2 commits into
mainfrom
hrx-amd-live

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves the engine's llama.cpp to AMD's maintained HRX line, pinned to the pair AMD tests, with a daily job that keeps it current.

Why

The engine pinned bong-water-water-bong/llama.cpp@3b33c8a. That's AMD's hrx-v2 line (ggml-hrx2, device HRX20) plus our 33 commits. AMD abandoned that line in June, and it's 2455 commits behind ggml-org. AMD's live line is AMD-Ecosystem/llama.cpp hrx-graph-develop-v2, which has a single ggml-hrx backend (device HRX0), was updated yesterday, and is 915 behind. It's pinned, together with ROCm/hrx-system, by AMD's integration repo ROCm/ggml-staging-automation.

What

  • The llama.cpp source: third_party/llama.cpp now points at the new org fork 1bit-MONSTER/llama.cpp (fork of ggml-org/llama.cpp), branch 1bit/hrx-vulkan = AMD's f1a0aca.
  • The HRX runtime: third_party/hrx-system = fab1624, AMD's pin. The separate third_party/hrx (ROCm/hrx) submodule and patches/hrx/ are removed.
  • The build (cmake/hrx.cmake): a single external project. llama.cpp's ggml-hrx builds hrx-system itself (HRX_SOURCE_DIR) with TheRock's amdclang.
  • The HSA runtime: the distro's libhsa-runtime64 rejects gfx1151's PM4_EMULATION probe, so HRX then registers no device. CMake now finds TheRock's copy, and 1bit lemonade passes it to llama-server as IREE_HAL_AMDGPU_LIBHSA_PATH (unless you've set it yourself).
  • The device name: HRX20 → HRX0 in the app defaults, tests/hrx_lemonade_e2e.sh and the docs.
  • Keeping current: a new daily workflow, .github/workflows/bump-hrx.yml. When AMD moves its pair, it syncs the fork (master to ggml-org, 1bit/hrx-vulkan to AMD's pin) and opens a PR moving both submodules. It needs a secret HRX_BUMP_TOKEN: a fine-grained token with Contents read/write on 1bit-MONSTER/llama.cpp and 1bit-MONSTER/engine, and Pull requests read/write on 1bit-MONSTER/engine.
  • The previous build is kept as 1bit/hrx2-archive in the fork (our Q4NX kernels, zaya, MoE grouping) for a later port to ggml-hrx.

Verified on Strix Halo

  • End to end: tests/hrx_lemonade_e2e.sh PASS. Qwen3-0.6B Q4_0 answers "Paris" through both recipes, served by this build's llama-server on Vulkan0 and HRX0.
  • Speed: llama-bench, Qwen3-0.6B, pp512 / tg128:
previous build this PR
Vulkan0, Q4_K_M 10627 / 306 13905 / 338
HRX, Q4_K_M 1046 / 68 (HRX20) 21702 / 303 (HRX0)
HRX, UD-Q4_K_XL 837 / 83 18059 / 288

Lost until ported

Q4NX GGUFs on HRX: our Q4NX kernels lived in ggml-hrx2, which ggml-hrx doesn't have. GGUFs still route to Vulkan0.

🤖 Generated with Claude Code

bong-water-water-bong and others added 2 commits September 23, 2026 12:00
… current

third_party/llama.cpp moves to the org fork 1bit-MONSTER/llama.cpp (of ggml-org/llama.cpp),
branch 1bit/hrx-vulkan = AMD's hrx-graph-develop-v2 f1a0aca; third_party/hrx-system moves to
fab1624. That is the pair ROCm/ggml-staging-automation builds and tests. The previous build sat on
AMD's hrx-v2 line, abandoned in June and 2455 commits behind ggml-org; it is kept as
1bit/hrx2-archive in the fork, with our Q4NX kernels and zaya, for a later port to ggml-hrx.

- cmake/hrx.cmake: one external project. llama.cpp's ggml-hrx builds hrx-system itself
  (HRX_SOURCE_DIR) with TheRock's amdclang. The separate ROCm/hrx runtime, its three patches and
  the loom-link-only build are gone.
- The device is HRX0 (ggml-hrx), not HRX20 (ggml-hrx2): app defaults, the e2e test, docs.
- HRX dlopens the HSA runtime; the distro's rejects gfx1151's PM4-emulation probe and HRX then
  registers no device. CMake finds TheRock's libhsa-runtime64 and `1bit lemonade` passes it to
  llama-server as IREE_HAL_AMDGPU_LIBHSA_PATH (unless already set).
- .github/workflows/bump-hrx.yml: daily, when AMD moves its pair, syncs the fork and opens a PR
  moving both submodules (needs the HRX_BUMP_TOKEN secret).

Strix Halo, Qwen3-0.6B: tests/hrx_lemonade_e2e.sh PASS (Paris on Vulkan0 and on HRX0).
llama-bench pp512/tg128, previous -> now: Vulkan0 Q4_K_M 10627/306 -> 13905/338;
HRX Q4_K_M 1046/68 -> 21702/303.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong enabled auto-merge (squash) September 23, 2026 16:19
@bong-water-water-bong
bong-water-water-bong merged commit 5b38461 into main Sep 23, 2026
1 check passed
@bong-water-water-bong
bong-water-water-bong deleted the hrx-amd-live branch September 23, 2026 16:30
bong-water-water-bong added a commit that referenced this pull request Sep 26, 2026
…bit-MONSTER/llama.cpp #11) (#127)

The pin adds sliding-window attention and a second rope base for ZAYA1-74B-preview on top of
1e775cd (#119). ZAYA1-8B is unchanged: Q4_K_M perplexity 21.5731 on Vulkan before and after,
and test-llama-archs -a zaya passes. ZAYA1-74B-preview Q4_K_M passes serve_e2e through
1bit serve --device vulkan.

registry/architectures.json, regenerated: it records the new pin (it still named 5556bf2
after #119 moved the pin to 1e775cd without regenerating it).

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong pushed a commit that referenced this pull request Sep 26, 2026
…oints and three HRX fixes

Our llama.cpp #7 (HRX MoE router route-ids binding length), #9 (decode-split
flash-attention capacity bound), #11 (ZAYA sliding-window attention, ZAYA1-74B),
#12 (MoE router experts-per-wave launch geometry) and #14 (ZAYA legacy
checkpoints; --remote fetches chat_template.jinja).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant