Repository navigation
hrx: IQ4_NL / IQ4_XS on HRX, correct in decode and prefill (Qwen3.8-27B UD: 91 -> 1 graph splits) - #41
Merged
bong-water-water-bong merged 2 commits intoSep 29, 2026
Conversation
…emap) The IQ4 nibble packers XORed every table index with 12, for a V_PERM lowering of vector.table.lookup that reverses both source halves. On gfx1151 with our pinned hrx-system the lookup indexes in order, so every IQ4_NL and IQ4_XS weight came out as the wrong table entry (test-backend-ops MUL_MAT error 3.7-5.8), and ggml-hrx.cpp declined those types: every Unsloth UD GGUF ran its IQ4_XS matmuls (4.6 of 12.2 GiB in UD-Q3_K_XL) on the CPU. Drop the remap in both packers (motifs/dequant.loom) and the two declines (IQ4 MUL_MAT / MUL_MAT_ID, IQ4_XS GET_ROWS). The batched IQ3_S GET_ROWS decline stays. test-backend-ops -b HRX0: 915/915 (was 791 with the IQ4 declines); all 22 IQ4_NL / IQ4_XS MUL_MAT cases pass (all failed without the declines). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 8407b1a)
With IQ4 weights on HRX (previous commit), Qwen3.8-27B UD-Q4_K_XL prefill on
HRX0 failed or went wrong in two fused paths:
- ops/mul_mat_q5_k_q8_plane_wmma.loom carries its own IQ4_XS nibble packer,
still with the XOR 12 table remap removed from motifs/dequant.loom: every
IQ4_XS weight in that prefill path decoded as the wrong entry (wikitext PPL
59,530). Same fix: the nibble is the table index.
- common.mul_mat_swiglu_q5_projection binds the SwiGLU output to the q8-plane
gate/up kernel, which publishes only q8_output (publish_f32 false). That
value is never written on this path, so the executor stopped the graph
("reads transient value before write"). Bind the already-written input as
the unused placeholder instead.
Qwen3.8-27B UD-Q4_K_XL, HRX0 (gfx1151), with both IQ4 commits:
- graph splits 91 -> 1 (76 IQ4_XS tensors no longer on the CPU);
- wikitext 4 x 512 PPL 5.8294 (CPU backend 5.8288);
- decode 7.25 -> 9.02 tok/s, 150 -> 130 W (Vulkan 12.49);
- test-backend-ops -b HRX0: MUL_MAT 22/22 and GET_ROWS 1/1 IQ4 cases pass.
ZAYA1-8B decode unchanged (90.0 tok/s).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
merged commit Sep 29, 2026
8c71384
into
1bit/hrx-vulkan-patched
10 of 24 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two commits on the integration head (c9283cf).
1.
ggml-hrx: IQ4_NL / IQ4_XS weights decode correctly (no XOR 12 table remap). Picked from1bit/hrx-iq4-fix(8407b1a), which had no PR.motifs/dequant.loomXORed every table index with 12, for a V_PERM lowering that doesn't apply on gfx1151.ggml-hrx.cppdeclined IQ4 matmuls, so every Unsloth UD GGUF ran its IQ4_XS matmuls on the CPU.2. Follow-up for what (1) exposed on Qwen3.8-27B UD-Q4_K_XL prefill:
ops/mul_mat_q5_k_q8_plane_wmma.loomhas its own IQ4_XS packer with the same XOR 12 remap, so that prefill path decoded every IQ4_XS weight wrong (wikitext PPL 59,530). Applied the same fix.common.mul_mat_swiglu_q5_projectionbound the never-written SwiGLU output to the q8-plane gate/up kernel, which publishes onlyq8_output. The executor stopped the graph with "reads transient value before write". It now binds the already-written input as the unused placeholder.Checked on strixhalo (gfx1151), Qwen3.8-27B UD-Q4_K_XL, HRX0:
test-backend-ops -b HRX0: MUL_MAT 22/22 and GET_ROWS 1/1 IQ4 cases pass.The 27B is still at 72% of Vulkan (12.49 tok/s) on decode. The remaining gap is kernel efficiency: the Q5_K/IQ4_XS decode GEMV is ALU-bound (IntelliKit Metrix: 122K VALU instructions per work-item, memory unit 8% busy). That's the next step.
🤖 Generated with Claude Code