Repository navigation
ggml-hrx: Q2_0 prompt matmuls on the int8 q8_1 x4 WMMA kernel (pp512 156 -> 323) - #79
bong-water-water-bong wants to merge 3 commits into
Conversation
A Q2_0 block (18 bytes: d f16, qs[16]) is 64 values, one K chunk of the q8_1 x4 kernel. motifs/nvfp4_q8_1_x4.loom stages it as signed values -1..2 (codes 0..3 minus one, so no offset term) under the block scale d for both K32 blocks, and the kernel's own contraction runs unchanged. The staging hook becomes @ggml_own_x4_stage (NVFP4 or Q2_0); AMD's file gains the Q2_0 flag and ORs it into the signed-weight flag, nothing else. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
common.mul_mat.q2_0_q8_1_x4_prefill (priority 300) shares the NVFP4 matcher's conditions: 256..2048 tokens in multiples of 256, input sizes in multiples of 256, output sizes in multiples of 64, weight format 42. sdkyuan/qwen3.8-27B-qat-q2_0, balanced mode: pp512 156 -> 323 tok/s; KLD vs CPU 0.00106 -> 0.00058. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… q8_1 x4 path Integer activations with |max| 127 per K32 block (exact q8_1, nonzero block sums) against raw Q2_0 blocks with positive, negative, small and large f16 scales; HRX must match an f64 reference from dequantize_row_q2_0 to 1e-5 of sum |w x| (it is within 5.5e-8). Treating codes 0..3 as unsigned or dropping an offset term is off by d x (block sum) and fails. Then quantized random weights against the CPU backend. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review (PR-Agent duty): approved. I'll merge after Sunday's release, stacked after #78. Signed staging (code - 1 under the block d) removes any offset term, and the known-answer test uses nonzero block sums, so a dropped offset would fail it (worst error 5.5e-8). AMD's file grows only by the Q2_0 flag and one argument (+50/-22 vs pre-#78 in total). Other-format PPL is identical; 27B Q2_0 pp512 156 -> 323, KLD 0.00106 -> 0.00058. |
|
Closed under the 2026-10-03 direction: new GPU kernels are HIP in the private add-on, the public fork keeps plumbing and fixes. Q8_0 and MXFP4 prompt GEMMs already run on HIP there (hip.mul_mat_q8_0_prefill, the GEAK MXFP4/Q6_K GEMMs); Q2_0 / NVFP4 / Q4_0-Q5_1 prompt matmuls will be HIP kernels when a model needs them. The branch stays for reference. |
Q2_0 prompt matmuls on the same int8 q8_1 x4 WMMA kernel as NVFP4. Stacked on #78. Q2_0 prefill no longer takes the f16 WMMA kernels (ggml-org#284 can skip it too).
How
motifs/nvfp4_q8_1_x4.loomstages it as signed values -1..2 (codes 0..3 minus one), so there is no offset or min term. Both K32 blocks get the block scaled, and the kernel's own contraction runs unchanged.mul_mat_q5_k_q8_plane_wmma.loomgains only the Q2_0 flag (plumbed like the NVFP4 one), ORed into the signed-weight flag, plus a second argument on the existing staging call. That is +50/-22 against its pre-ggml-hrx: NVFP4 prompt matmuls on the int8 q8_1 x4 WMMA kernel (pp512 107 -> 323) #78 text (ggml-hrx: NVFP4 prompt matmuls on the int8 q8_1 x4 WMMA kernel (pp512 107 -> 323) #78 alone: +46/-22). The staging hook is renamed@ggml_own_x4_stageand handles NVFP4 or Q2_0.common.mul_mat.q2_0_q8_1_x4_prefillshares the NVFP4 matcher (dispatch-mul-mat-nvfp4.cpp): 256..2048 tokens in multiples of 256, K a multiple of 256, N a multiple of 64.Validation, gfx1151, balanced power mode:
q2_0_q8_1_x4_prefillNot for the release pin.
🤖 Generated with Claude Code