Skip to content

hrx: Q2_K weights on HRX0 (shared dequantizer, K-quant decode kernels) - #49

Merged
bong-water-water-bong merged 2 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-q2k
Sep 30, 2026
Merged

bong-water-water-bong merged 2 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-q2k

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

Q2_K had no matmul on HRX0. llama.cpp's load-time buffer check (a 512-token MUL_MAT per weight) failed, so every Q2_K tensor landed in the CPU_REPACK buffer and each use split the graph. On Qwen3-4B Q2_K that was 709 MiB on the CPU and 217 splits. Every Unsloth UD GGUF carries some Q2_K.

Changes

  • Shared dequantizer (motifs/dequant.loom): ggml_q2k_f32_vector4 / _f16_vector4 follow dequantize_row_q2_K, with the same group/packet addressing as Q3_K. Format 12 is in the tile/row byte tables and both decode dispatchers, so the prefill WMMA kernels, generic decode and GET_ROWS read Q2_K.
  • Format enum: CommonMulMatWeightFormat::Q2K (ggml Q2_K ↔ config 12).
  • K-quant decode kernels (ops/kquant_decode_f32.loom): Q2_K lane functions for the 1-token and 2-8-token kernels. One lane is one 16-value sub-block with one scale and one min; block size 84, run shape 16. dispatch-kquant-decode.cpp accepts Q2K.

Checked on strixhalo (gfx1151, HRX0), with Qwen3-4B requantized Q8_0 → Q2_K:

check before after
weights 709 MiB CPU_REPACK, 217 graph splits 1586 MiB HRX0, 73 splits
test-backend-ops -b HRX0 929/929
MUL_MAT q2_K 9 run, 2 not supported 11/11
pp512 (5 reps) 686 1287–1325
tg128 (5 reps) 47.0 ± 0.4 74.4–74.8 ± 14

The tg128 spread is most likely the first repetition paying for new Loom kernel compiles.

KLD against the Q8_0 model (wikitext-2, 6 × 512):

  • prefill path: 1.1045;
  • decode path (-b 1): 1.1031;
  • CPU: 1.1140;
  • previous HRX with Q2_K on the CPU: 1.1224.

GET_ROWS q2_K: 1 shape runs and passes; batched shapes are still declined. MUL_MAT_ID (MoE experts) doesn't take Q2_K yet. IQ2_XS and IQ2_XXS follow the same way in a separate PR.

🤖 Generated with Claude Code

Q2_K had no HRX matmul: llama.cpp's load-time buffer check (a 512-token
MUL_MAT per weight) failed, every Q2_K tensor went to the CPU_REPACK buffer,
and each use split the graph (Qwen3-4B Q2_K: 709 MiB on the CPU, 217 splits).
Unsloth's UD GGUFs all carry some Q2_K.

- motifs/dequant.loom: ggml_q2k_f32_vector4 / _f16_vector4 (dequantize_row_q2_K
  order, the same group/packet addressing as Q3_K), format 12 in the tile and
  row byte tables and in both decode dispatchers, so the prefill WMMA kernels,
  generic decode, and GET_ROWS read Q2_K.
- dispatch-mul-mat-weight-format.h / -common.h: CommonMulMatWeightFormat::Q2K
  (ggml Q2_K, config 12).
- ops/kquant_decode_f32.loom: Q2_K lane functions for the 1-token and 2-8 token
  decode kernels (lane = one 16-value sub-block: one scale, one min), block size
  84, run shape 16; dispatch-kquant-decode.cpp accepts Q2K.

Checked on HRX0 (strixhalo, gfx1151), Qwen3-4B requantized Q8_0 -> Q2_K:
- test-backend-ops -b HRX0: 929/929; MUL_MAT q2_K 11/11 (was 9 run, 2 not
  supported); GET_ROWS q2_K 1 run + OK, batched shapes still declined.
- KLD vs the Q8_0 model (wikitext-2, 6 x 512): prefill path 1.1045, decode
  path (-b 1) 1.1031; CPU 1.1140, previous HRX (Q2_K on the CPU) 1.1224.
- llama-bench HRX0: pp512 677 -> 1279 tok/s; tg128 see PR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the ggml label Sep 30, 2026
@bong-water-water-bong

Copy link
Copy Markdown
Author

Review (standing in for PR-Agent while the local model is paused)

Checked the Q2_K lane functions against dequantize_row_q2_K.

Correctness: all of these match.

  • Block layout: 84 bytes = scales[16], qs[64], d, dmin. d is i16 word 40 and dmin is word 41; qs starts at word 8.
  • Lane mapping: lane l = 8n + 2j + h owns values 128n + 32j + 16h + 0..15, with scale byte scales[l] and qs bytes 32n + 16h .., i.e. words 8 + 16n + 8h. This matches is++ in the reference loop.
  • 2-bit extraction: the 32-bit shift by 2j followed by the 0x03030303 mask gives the right bits in every byte; bits shifted across byte boundaries are masked out.
  • Dot product: d·sc·Σqx − dmin·m·Σx is the reference formula.
  • Multi-token path: lane_weights gives q·dl − ml. Run shape 2 (16 contiguous) is right for one sub-block per lane.
  • Other tables: ggml_kquant_block_bytes gives 84 for format 12. The dispatcher's kquant_format, the enum, the ggml type map and the config value 12 are consistent.

Nit: the header comment says format 12 is one "which only dispatch-kquant-decode.cpp uses". Format 12 is now also in the dequant motif's tile and row tables (prefill WMMA, generic decode, GET_ROWS), so that clause could go.

CI:

Good to merge once the re-run is green.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit d1747cb into 1bit/hrx-vulkan-patched Sep 30, 2026
10 of 24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant