Skip to content

hrx: IQ1_S and IQ1_M weights on HRX0 (shared dequantizer, K-quant decode kernels) - #53

Merged
bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-iq1-v2
Oct 1, 2026
Merged

bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-iq1-v2

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

IQ1_S and IQ1_M weights now run on HRX0. Until now every such tensor ran on the CPU, which split each layer it was in. Unsloth's smallest UD files (UD-IQ1_S, UD-IQ1_M) and the sub-3-bit UD files all carry these types.

What changed (same pattern as #50 and #51):

  • Shared dequantizer (motifs/dequant.loom): formats 26 (IQ1_S) and 27 (IQ1_M). This covers prompt WMMA, the generic decode path and the tile sizes.
  • K-quant decode kernels (ops/kquant_decode_f32.loom): IQ1 lane functions for the SwiGLU gate/up, plain and add-residual decode kernels.
    • The codebook is staged in workgroup memory like the IQ2/IQ3 grids. The buffer grows from 2 to 4 KiB to hold the 2048 IQ1 entries.
    • Each grid value is -1, 0 or 1, so an entry is a 16-bit code (2 bits per value), packed two to a word.
  • Generator: tools/generate_iq1_loom.py produces both Loom blocks from iq1s_grid in ggml-common.h. It passes flake8 and pyright, and its output matches the committed Loom.
  • Dispatch: the weight-format mapping covers IQ1_S/IQ1_M. kquant_needs_grid becomes kquant_grid, which returns a codebook id. A gate/up pair is declined only when the two formats need different codebooks; IQ1_S and IQ1_M share one.

Measured on strixhalo (Qwen3-4B requantized from Q8_0 with an imatrix, HRX0, balanced power mode, five runs each). The baseline is cde002d, where these weights run on the CPU.

Format pp512 (tok/s) tg128 (tok/s, steady state) KLD vs Q8_0, HRX prefill / decode / CPU
IQ1_S 70.8 → 81.7 14.1 → 18.6 3.046 / 3.043 / 3.033
IQ1_M 70.8 → 80.4 14.1 → 18.1 2.187 / 2.186 / 2.191

test-backend-ops -b HRX0: 996/996 passed, including 11/11 MUL_MAT cases for each of iq1_s and iq1_m.

Known cost: the first decode run in a process is slow (about 5 tok/s for tg128). That run includes roughly 18 s of kernel JIT for the unrolled grid fill. The follow-up is a constant grid buffer bound once through constant_initializations, in place of the per-workgroup fill. That removes both the JIT time and the fill.

🤖 Generated with Claude Code

…ode kernels)

IQ1_S (format 26) and IQ1_M (27) join the shared dequantizer (prompt WMMA, generic decode) and the K-quant
decode kernels, in the IQ2 pattern. The grid tables and lane functions come from
tools/generate_iq1_loom.py (iq1s_grid from ggml-common.h). Each grid value is -1, 0 or 1, so an entry is a
16-bit code with 2 bits per value, two entries per word. The K-quant grid buffer grows to 4 KiB for the
2048 IQ1 entries. A SwiGLU gate/up pair is declined when its two formats need different grids; IQ1_S and
IQ1_M share one.

Qwen3-4B, balanced power mode, HRX0, against cde002d, where these weights run on the CPU:
  IQ1_S  pp512 70.8 -> 81.7 tok/s, tg128 14.1 -> 18.6 (steady state)
  IQ1_M  pp512 70.8 -> 80.4 tok/s, tg128 14.1 -> 18.1
KLD against the Q8_0 model equals the CPU's (IQ1_S 3.04, IQ1_M 2.19, on both the prefill and decode
paths). test-backend-ops -b HRX0: 996/996.

The first decode run in a process includes about 18 s of kernel JIT for the unrolled grid fill. The
follow-up is a constant grid buffer in place of the per-workgroup fill.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the ggml label Oct 1, 2026

@bong-water-water-bong bong-water-water-bong left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review (PR-Agent duty). The approach is right and the numbers are good: IQ1 moves off the CPU, KLD matches the CPU, and test-backend-ops passes 996/996. Before merging I'd like one number, plus two non-blocking notes.

Needed: a decode check on non-IQ formats.
The workgroup grid alloca in kquant_decode_f32.loom goes from 2048 to 4096 bytes for every K-quant decode kernel, not just IQ1. %grid is allocated before scf.if %needs_grid.

  • If Loom drops the unused alloca when the format config makes needs_grid false, this is free.
  • If it doesn't, every Q4_K/Q5_K/Q6_K/Q8_0 decode workgroup reserves 2 KiB more LDS, which can cost occupancy on the showcase models.

Please post one A/B against cde002d, interleaved, balanced mode:

  • tg on Qwen3.8-27B UD-Q4_K_XL (with and without --mtp);
  • tg on ZAYA1-8B Q4_K_M.

Alternatively, size the alloca from the format (4096 only for 26/27, 2048 for the others, 0 when no grid is needed).

Non-blocking.

  1. motifs/dequant.loom is an upstream file with no notice, and it now carries about 2.6k more lines of our code (IQ2 and IQ3_XXS were added before). The follow-up we agreed on #51 (IQ motifs in our own headed file, one-line hooks upstream) is getting more valuable; please do it next.
  2. About 18 s of JIT on the first decode for the unrolled grid fill. The JIT disk cache makes that once per machine, not per process, so it's fine for now. The constant grid buffer follow-up fixes it properly.

Checked: the grid-id switch (kquant_grid) keeps the mixed-pair decline correct and lets IQ1_S + IQ1_M pairs through, and the generator carries the full Apache notice.

@bong-water-water-bong

Copy link
Copy Markdown
Author

Does the 4 KiB grid buffer slow the other formats? No. The K-quant decode grid buffer is now 4 KiB for every format, so I ran an interleaved A/B on strixhalo, cde002d against this branch: balanced power mode, three rounds, llama-bench with 3 reps each, and 1bit serve for the --mtp rows.

cde002d this PR
Qwen3.8-27B UD-Q4_K_XL, tg128 11.8 / 11.9 / 11.9 (avg 11.88) 11.9 / 11.9 / 11.9 (avg 11.88)
ZAYA1-8B Q4_K_M, tg128 72.0* / 91.8 / 93.0 93.2 / 91.9 / 91.8
27B UD-Q4_K_XL --mtp, code (tok/s) 19.1 / 18.6 / 19.3 (avg 19.0) 17.9 / 19.4 / 19.8 (avg 19.0)
27B UD-Q4_K_XL --mtp, prose (tok/s) 14.5 / 14.8 / 14.0 (avg 14.4) 15.3 / 14.7 / 15.0 (avg 15.0)

* The first cde002d ZAYA sample was 28.9 tok/s (JIT warm-up on the first run). The other two were 92.0 and 95.1.

Every difference is inside run-to-run noise, so I kept the 4 KiB buffer for all formats rather than sizing it per format. The script is ~/lb/iq1-ab.sh and the raw output is ~/lb/iq1-ab.out.

🤖 Generated with Claude Code

@bong-water-water-bong

Copy link
Copy Markdown
Author

Merging. The Lint and python type-check jobs need a [self-hosted, fast] runner, and none serves this fork, so they stay queued. #51 merged with the same two jobs pending. I ran the same checks locally on the one Python file this PR adds (ggml/src/ggml-hrx/tools/generate_iq1_loom.py): flake8 with the repository .flake8, and pyright, both clean. The hosted builds pass: ubuntu (x64, arm64), ubuntu-webgpu, windows. The reviewer's 4 KiB grid question is answered by the A/B above.

🤖 Generated with Claude Code

@bong-water-water-bong
bong-water-water-bong merged commit bf5ad7c into 1bit/hrx-vulkan-patched Oct 1, 2026
10 of 27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant