Repository navigation
hrx: IQ1_S and IQ1_M weights on HRX0 (shared dequantizer, K-quant decode kernels) - #53
Conversation
…ode kernels) IQ1_S (format 26) and IQ1_M (27) join the shared dequantizer (prompt WMMA, generic decode) and the K-quant decode kernels, in the IQ2 pattern. The grid tables and lane functions come from tools/generate_iq1_loom.py (iq1s_grid from ggml-common.h). Each grid value is -1, 0 or 1, so an entry is a 16-bit code with 2 bits per value, two entries per word. The K-quant grid buffer grows to 4 KiB for the 2048 IQ1 entries. A SwiGLU gate/up pair is declined when its two formats need different grids; IQ1_S and IQ1_M share one. Qwen3-4B, balanced power mode, HRX0, against cde002d, where these weights run on the CPU: IQ1_S pp512 70.8 -> 81.7 tok/s, tg128 14.1 -> 18.6 (steady state) IQ1_M pp512 70.8 -> 80.4 tok/s, tg128 14.1 -> 18.1 KLD against the Q8_0 model equals the CPU's (IQ1_S 3.04, IQ1_M 2.19, on both the prefill and decode paths). test-backend-ops -b HRX0: 996/996. The first decode run in a process includes about 18 s of kernel JIT for the unrolled grid fill. The follow-up is a constant grid buffer in place of the per-workgroup fill. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
left a comment
There was a problem hiding this comment.
Review (PR-Agent duty). The approach is right and the numbers are good: IQ1 moves off the CPU, KLD matches the CPU, and test-backend-ops passes 996/996. Before merging I'd like one number, plus two non-blocking notes.
Needed: a decode check on non-IQ formats.
The workgroup grid alloca in kquant_decode_f32.loom goes from 2048 to 4096 bytes for every K-quant decode kernel, not just IQ1. %grid is allocated before scf.if %needs_grid.
- If Loom drops the unused alloca when the format config makes
needs_gridfalse, this is free. - If it doesn't, every Q4_K/Q5_K/Q6_K/Q8_0 decode workgroup reserves 2 KiB more LDS, which can cost occupancy on the showcase models.
Please post one A/B against cde002d, interleaved, balanced mode:
- tg on Qwen3.8-27B UD-Q4_K_XL (with and without --mtp);
- tg on ZAYA1-8B Q4_K_M.
Alternatively, size the alloca from the format (4096 only for 26/27, 2048 for the others, 0 when no grid is needed).
Non-blocking.
motifs/dequant.loomis an upstream file with no notice, and it now carries about 2.6k more lines of our code (IQ2 and IQ3_XXS were added before). The follow-up we agreed on #51 (IQ motifs in our own headed file, one-line hooks upstream) is getting more valuable; please do it next.- About 18 s of JIT on the first decode for the unrolled grid fill. The JIT disk cache makes that once per machine, not per process, so it's fine for now. The constant grid buffer follow-up fixes it properly.
Checked: the grid-id switch (kquant_grid) keeps the mixed-pair decline correct and lets IQ1_S + IQ1_M pairs through, and the generator carries the full Apache notice.
|
Does the 4 KiB grid buffer slow the other formats? No. The K-quant decode grid buffer is now 4 KiB for every format, so I ran an interleaved A/B on strixhalo,
* The first Every difference is inside run-to-run noise, so I kept the 4 KiB buffer for all formats rather than sizing it per format. The script is 🤖 Generated with Claude Code |
|
Merging. The 🤖 Generated with Claude Code |
bf5ad7c
into
1bit/hrx-vulkan-patched
IQ1_S and IQ1_M weights now run on HRX0. Until now every such tensor ran on the CPU, which split each layer it was in. Unsloth's smallest UD files (UD-IQ1_S, UD-IQ1_M) and the sub-3-bit UD files all carry these types.
What changed (same pattern as #50 and #51):
motifs/dequant.loom): formats 26 (IQ1_S) and 27 (IQ1_M). This covers prompt WMMA, the generic decode path and the tile sizes.ops/kquant_decode_f32.loom): IQ1 lane functions for the SwiGLU gate/up, plain and add-residual decode kernels.tools/generate_iq1_loom.pyproduces both Loom blocks fromiq1s_gridinggml-common.h. It passes flake8 and pyright, and its output matches the committed Loom.kquant_needs_gridbecomeskquant_grid, which returns a codebook id. A gate/up pair is declined only when the two formats need different codebooks; IQ1_S and IQ1_M share one.Measured on strixhalo (Qwen3-4B requantized from Q8_0 with an imatrix, HRX0, balanced power mode, five runs each). The baseline is
cde002d, where these weights run on the CPU.test-backend-ops -b HRX0: 996/996 passed, including 11/11 MUL_MAT cases for each of iq1_s and iq1_m.Known cost: the first decode run in a process is slow (about 5 tok/s for tg128). That run includes roughly 18 s of kernel JIT for the unrolled grid fill. The follow-up is a constant grid buffer bound once through
constant_initializations, in place of the per-workgroup fill. That removes both the JIT time and the fill.🤖 Generated with Claude Code