Skip to content

Pin llama.cpp 8c71384: IQ4_NL / IQ4_XS on HRX (Qwen3.8-27B UD: 58% -> 72% of Vulkan) - #234

Merged
bong-water-water-bong merged 2 commits into
mainfrom
pin-hrx-iq4
Sep 29, 2026
Merged

bong-water-water-bong merged 2 commits into
mainfrom
pin-hrx-iq4

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves third_party/llama.cpp c9283cf -> 8c71384 (fork #41). Regenerates the registry and updates docs/hrx.md.

HRX declined IQ4_NL / IQ4_XS because AMD's nibble packers remapped every table index (XOR 12) for a lowering that doesn't apply on gfx1151. So every Unsloth UD GGUF ran its IQ4 matmuls on the CPU. Fork #41:

  • picks the other session's fix for the shared packers;
  • fixes a second copy of the remap in the Q5_K/IQ4_XS prefill kernel;
  • fixes the q8-plane SwiGLU placeholder binding that IQ4 on HRX exposed.

Qwen3.8-27B UD-Q4_K_XL on strixhalo, HRX0:

  • graph splits 91 -> 1;
  • wikitext 4×512 PPL 5.8294 (CPU 5.8288);
  • decode 7.25 -> 9.0 tok/s at 130 W (Vulkan 12.5), flat over a 2.5-minute sustained run;
  • ZAYA1-8B unchanged (90 tok/s).

RFC #213: the 27B is the one model still under the decode gate (72%). The remaining gap is the Q5_K/IQ4_XS decode kernels' memory bandwidth.

🤖 Generated with Claude Code

bong-water-water-bong and others added 2 commits September 29, 2026 16:04
…1 graph splits)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 29, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 4cea856

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Outdated Documentation

The documentation in docs/hrx.md still refers to the old llama.cpp commit and mentions that IQ4_NL and IQ4_XS matmul kernels give wrong values and are left to the CPU. However, the PR updates the llama.cpp commit to 8c71384 which includes fixes for these issues. The documentation should be updated to reflect that these kernels are now correctly supported on HRX.

views), which cannot move. Leaf (`NONE`) nodes count as covered. GET_ROWS for batched IQ3_S
gives wrong values, so those nodes are left to the CPU. IQ4_NL and IQ4_XS were left to the CPU too
until [llama.cpp #41](https://github.com/1bit-MONSTER/llama.cpp/pull/41) (below).
Incomplete Feature Description

The documentation in docs/hrx.md describes IQ4_NL / IQ4_XS on HRX as fixed, but does not fully capture the performance improvements mentioned in the PR description. Specifically, the PR mentions a 72% decode performance improvement for Qwen3.8-27B UD-Q4_K_XL, which is not reflected in the documentation update.

- **IQ4_NL / IQ4_XS on HRX ([llama.cpp #41](https://github.com/1bit-MONSTER/llama.cpp/pull/41)).**
  AMD's IQ4 nibble packers remapped every table index (XOR 12) for a lowering that does not apply on
  gfx1151, in `motifs/dequant.loom` and again in the Q5_K/IQ4_XS prefill kernel, so IQ4 was declined
  and ran on the CPU. #41 fixes both copies and the q8-plane SwiGLU path they exposed. Qwen3.8-27B
  UD-Q4_K_XL: 91 graph splits -> 1, wikitext PPL 5.8294 (CPU 5.8288), decode 7.25 -> 9.0 tok/s
  (Vulkan 12.5); `test-backend-ops -b HRX0` passes the IQ4 MUL_MAT and GET_ROWS cases. The rest of
  the 27B gap is Q5_K / IQ4_XS decode kernels reading at ~105-176 GB/s against Vulkan's ~220.

@bong-water-water-bong
bong-water-water-bong merged commit 8b33745 into main Sep 29, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin-hrx-iq4 branch September 29, 2026 19:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant