Skip to content

ggml-hrx: Q2_0 prompt matmuls on the int8 q8_1 x4 WMMA kernel (pp512 156 -> 323) - #79

Closed
bong-water-water-bong wants to merge 3 commits into
1bit/hrx-nvfp4-prefillfrom
1bit/hrx-q2_0-prefill
Closed

bong-water-water-bong wants to merge 3 commits into
1bit/hrx-nvfp4-prefillfrom
1bit/hrx-q2_0-prefill

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

Q2_0 prompt matmuls on the same int8 q8_1 x4 WMMA kernel as NVFP4. Stacked on #78. Q2_0 prefill no longer takes the f16 WMMA kernels (ggml-org#284 can skip it too).

How

  • A Q2_0 block (18 bytes: d f16, qs[16]) is 64 values, one K chunk of the kernel. motifs/nvfp4_q8_1_x4.loom stages it as signed values -1..2 (codes 0..3 minus one), so there is no offset or min term. Both K32 blocks get the block scale d, and the kernel's own contraction runs unchanged.
  • AMD's mul_mat_q5_k_q8_plane_wmma.loom gains only the Q2_0 flag (plumbed like the NVFP4 one), ORed into the signed-weight flag, plus a second argument on the existing staging call. That is +50/-22 against its pre-ggml-hrx: NVFP4 prompt matmuls on the int8 q8_1 x4 WMMA kernel (pp512 107 -> 323) #78 text (ggml-hrx: NVFP4 prompt matmuls on the int8 q8_1 x4 WMMA kernel (pp512 107 -> 323) #78 alone: +46/-22). The staging hook is renamed @ggml_own_x4_stage and handles NVFP4 or Q2_0.
  • common.mul_mat.q2_0_q8_1_x4_prefill shares the NVFP4 matcher (dispatch-mul-mat-nvfp4.cpp): 256..2048 tokens in multiples of 256, K a multiple of 256, N a multiple of 64.

Validation, gfx1151, balanced power mode:

Check Result
test-hrx-q2_0-prefill, known answer: integer activations (exact q8_1, nonzero block sums, so a missing offset term fails), raw blocks with positive/negative/small/large f16 scales; K 256-5120, 256-1024 tokens worst error 5.5e-8 of sum |w x| (limit 1e-5)
same shapes, random weights vs CPU NMSE 4e-15 .. 9e-10
test-hrx-nvfp4-prefill and all other HRX tests pass
test-backend-ops -b HRX0 1129/1129
Qwen3-0.6B Q4_K_M / UD-Q4_K_XL perplexity identical to before #78
27B NVFP4 (cdiamond) KLD / PPL 0.004225 / 6.651781, identical to #78
sdkyuan/qwen3.8-27B-qat-q2_0 pp512 156 -> 323 tok/s
same, KLD vs CPU (wikitext-2 8 x 512) 0.00106 -> 0.00058, same top token 97.9%
same, tg128 / routing 19.8 (unchanged); graph splits = 1, prefill on q2_0_q8_1_x4_prefill

Not for the release pin.

🤖 Generated with Claude Code

bong-water-water-bong and others added 3 commits October 2, 2026 20:53
A Q2_0 block (18 bytes: d f16, qs[16]) is 64 values, one K chunk of the q8_1 x4 kernel.
motifs/nvfp4_q8_1_x4.loom stages it as signed values -1..2 (codes 0..3 minus one, so no offset term) under
the block scale d for both K32 blocks, and the kernel's own contraction runs unchanged. The staging hook
becomes @ggml_own_x4_stage (NVFP4 or Q2_0); AMD's file gains the Q2_0 flag and ORs it into the signed-weight
flag, nothing else.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
common.mul_mat.q2_0_q8_1_x4_prefill (priority 300) shares the NVFP4 matcher's conditions: 256..2048 tokens in
multiples of 256, input sizes in multiples of 256, output sizes in multiples of 64, weight format 42.

sdkyuan/qwen3.8-27B-qat-q2_0, balanced mode: pp512 156 -> 323 tok/s; KLD vs CPU 0.00106 -> 0.00058.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… q8_1 x4 path

Integer activations with |max| 127 per K32 block (exact q8_1, nonzero block sums) against raw Q2_0 blocks with
positive, negative, small and large f16 scales; HRX must match an f64 reference from dequantize_row_q2_0 to 1e-5
of sum |w x| (it is within 5.5e-8). Treating codes 0..3 as unsigned or dropping an offset term is off by
d x (block sum) and fails. Then quantized random weights against the CPU backend.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong

Copy link
Copy Markdown
Author

Review (PR-Agent duty): approved. I'll merge after Sunday's release, stacked after #78. Signed staging (code - 1 under the block d) removes any offset term, and the known-answer test uses nonzero block sums, so a dropped offset would fail it (worst error 5.5e-8). AMD's file grows only by the Q2_0 flag and one argument (+50/-22 vs pre-#78 in total). Other-format PPL is identical; 27B Q2_0 pp512 156 -> 323, KLD 0.00106 -> 0.00058.

@bong-water-water-bong

Copy link
Copy Markdown
Author

Closed under the 2026-10-03 direction: new GPU kernels are HIP in the private add-on, the public fork keeps plumbing and fixes. Q8_0 and MXFP4 prompt GEMMs already run on HIP there (hip.mul_mat_q8_0_prefill, the GEAK MXFP4/Q6_K GEMMs); Q2_0 / NVFP4 / Q4_0-Q5_1 prompt matmuls will be HIP kernels when a model needs them. The branch stays for reference.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant