Repository navigation
hrx: Q2_K weights on HRX0 (shared dequantizer, K-quant decode kernels) - #49
Merged
Merged
Conversation
Q2_K had no HRX matmul: llama.cpp's load-time buffer check (a 512-token MUL_MAT per weight) failed, every Q2_K tensor went to the CPU_REPACK buffer, and each use split the graph (Qwen3-4B Q2_K: 709 MiB on the CPU, 217 splits). Unsloth's UD GGUFs all carry some Q2_K. - motifs/dequant.loom: ggml_q2k_f32_vector4 / _f16_vector4 (dequantize_row_q2_K order, the same group/packet addressing as Q3_K), format 12 in the tile and row byte tables and in both decode dispatchers, so the prefill WMMA kernels, generic decode, and GET_ROWS read Q2_K. - dispatch-mul-mat-weight-format.h / -common.h: CommonMulMatWeightFormat::Q2K (ggml Q2_K, config 12). - ops/kquant_decode_f32.loom: Q2_K lane functions for the 1-token and 2-8 token decode kernels (lane = one 16-value sub-block: one scale, one min), block size 84, run shape 16; dispatch-kquant-decode.cpp accepts Q2K. Checked on HRX0 (strixhalo, gfx1151), Qwen3-4B requantized Q8_0 -> Q2_K: - test-backend-ops -b HRX0: 929/929; MUL_MAT q2_K 11/11 (was 9 run, 2 not supported); GET_ROWS q2_K 1 run + OK, batched shapes still declined. - KLD vs the Q8_0 model (wikitext-2, 6 x 512): prefill path 1.1045, decode path (-b 1) 1.1031; CPU 1.1140, previous HRX (Q2_K on the CPU) 1.1224. - llama-bench HRX0: pp512 677 -> 1279 tok/s; tg128 see PR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Author
|
Review (standing in for PR-Agent while the local model is paused) Checked the Q2_K lane functions against Correctness: all of these match.
Nit: the header comment says format 12 is one "which only dispatch-kquant-decode.cpp uses". Format 12 is now also in the dequant motif's tile and row tables (prefill WMMA, generic decode, GET_ROWS), so that clause could go. CI:
Good to merge once the re-run is green. |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
merged commit Sep 30, 2026
d1747cb
into
1bit/hrx-vulkan-patched
10 of 24 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Q2_K had no matmul on HRX0. llama.cpp's load-time buffer check (a 512-token MUL_MAT per weight) failed, so every Q2_K tensor landed in the CPU_REPACK buffer and each use split the graph. On Qwen3-4B Q2_K that was 709 MiB on the CPU and 217 splits. Every Unsloth UD GGUF carries some Q2_K.
Changes
motifs/dequant.loom):ggml_q2k_f32_vector4/_f16_vector4followdequantize_row_q2_K, with the same group/packet addressing as Q3_K. Format 12 is in the tile/row byte tables and both decode dispatchers, so the prefill WMMA kernels, generic decode and GET_ROWS read Q2_K.CommonMulMatWeightFormat::Q2K(ggml Q2_K ↔ config 12).ops/kquant_decode_f32.loom): Q2_K lane functions for the 1-token and 2-8-token kernels. One lane is one 16-value sub-block with one scale and one min; block size 84, run shape 16.dispatch-kquant-decode.cppaccepts Q2K.Checked on strixhalo (gfx1151, HRX0), with Qwen3-4B requantized Q8_0 → Q2_K:
test-backend-ops -b HRX0The tg128 spread is most likely the first repetition paying for new Loom kernel compiles.
KLD against the Q8_0 model (wikitext-2, 6 × 512):
-b 1): 1.1031;GET_ROWS q2_K: 1 shape runs and passes; batched shapes are still declined. MUL_MAT_ID (MoE experts) doesn't take Q2_K yet. IQ2_XS and IQ2_XXS follow the same way in a separate PR.
🤖 Generated with Claude Code