Skip to content

hrx: IQ2_XXS and IQ2_XS weights on HRX0 (shared dequantizer, K-quant decode kernels) - #50

Merged
bong-water-water-bong merged 3 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-iq2
Sep 30, 2026
Merged

bong-water-water-bong merged 3 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-iq2

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

IQ2_XXS and IQ2_XS had no HRX matmul: the load-time 512-token MUL_MAT check
failed, the weights went to CPU_REPACK, and each use split the graph. Unsloth's
UD GGUFs below 4 bits carry both.

  • motifs/dequant.loom: ggml_iq2xxs/iq2xs f16/f32 vector4 decoders (block bytes
    66 / 74, formats 24 / 25 in the tile and row byte tables and both decode
    dispatchers), so prefill WMMA, generic decode and GET_ROWS read them. The
    grids are 16-bit codes (every grid byte is 8, 25 or 43: 2 bits per value);
    signs from ksigns = 7 bits + parity.
  • dispatch-mul-mat-weight-format.h / -common.h: CommonMulMatWeightFormat
    IQ2_XXS (24) and IQ2_XS (25).
  • ops/kquant_decode_f32.loom: lane functions for the 1-token and 2-8 token
    decode kernels, run shape 16; grids staged in workgroup memory once per
    workgroup like IQ3_S (the barrier condition is computed inline from the
    format so it stays workgroup-uniform); dispatch-kquant-decode.cpp accepts
    both, and the SwiGLU matcher only stages a grid when a format needs one.
  • tools/generate_iq2_dequant_loom.py, generate_iq2_kquant_loom.py: the
    generators for the grid tables and decoders; their output is verbatim in the
    two .loom files.

Checked on HRX0 (strixhalo, gfx1151), Qwen3-4B requantized Q8_0 -> IQ2_*:

  • test-backend-ops -b HRX0: 971/971; MUL_MAT iq2_xs 11 OK, iq2_xxs 12 OK
    (64 other iq2_xxs shapes declined, so they run on the CPU as before).
  • KLD vs the Q8_0 model (wikitext-2, 6 x 512), decode path (-b 1):
    IQ2_XXS 1.1306 (CPU 1.141), IQ2_XS 0.7744 (CPU 0.786).
  • llama-bench HRX0, previous (IQ2 on the CPU) -> this:
    IQ2_XXS pp512 59 -> 715, tg128 23.0 -> 39.5;
    IQ2_XS pp512 87 -> 268, tg128 28.0 -> 30.2 (tg: 2 x 6 runs, box locked).

Next gap after this: IQ1_S / IQ1_M.

🤖 Generated with Claude Code

…decode kernels)

IQ2_XXS and IQ2_XS had no HRX matmul: the load-time 512-token MUL_MAT check
failed, the weights went to CPU_REPACK, and each use split the graph. Unsloth's
UD GGUFs below 4 bits carry both.

- motifs/dequant.loom: ggml_iq2xxs/iq2xs f16/f32 vector4 decoders (block bytes
  66 / 74, formats 24 / 25 in the tile and row byte tables and both decode
  dispatchers), so prefill WMMA, generic decode and GET_ROWS read them. The
  grids are 16-bit codes (every grid byte is 8, 25 or 43: 2 bits per value);
  signs from ksigns = 7 bits + parity.
- dispatch-mul-mat-weight-format.h / -common.h: CommonMulMatWeightFormat
  IQ2_XXS (24) and IQ2_XS (25).
- ops/kquant_decode_f32.loom: lane functions for the 1-token and 2-8 token
  decode kernels, run shape 16; grids staged in workgroup memory once per
  workgroup like IQ3_S (the barrier condition is computed inline from the
  format so it stays workgroup-uniform); dispatch-kquant-decode.cpp accepts
  both, and the SwiGLU matcher only stages a grid when a format needs one.
- tools/generate_iq2_dequant_loom.py, generate_iq2_kquant_loom.py: the
  generators for the grid tables and decoders; their output is verbatim in the
  two .loom files.

Checked on HRX0 (strixhalo, gfx1151), Qwen3-4B requantized Q8_0 -> IQ2_*:
- test-backend-ops -b HRX0: 971/971; MUL_MAT iq2_xs 11 OK, iq2_xxs 12 OK
  (64 other iq2_xxs shapes declined, so they run on the CPU as before).
- KLD vs the Q8_0 model (wikitext-2, 6 x 512), decode path (-b 1):
  IQ2_XXS 1.1306 (CPU 1.141), IQ2_XS 0.7744 (CPU 0.786).
- llama-bench HRX0, previous (IQ2 on the CPU) -> this:
  IQ2_XXS pp512 59 -> 715, tg128 23.0 -> 39.5;
  IQ2_XS  pp512 87 -> 268, tg128 28.0 -> 30.2 (tg: 2 x 6 runs, box locked).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the ggml label Sep 30, 2026
@bong-water-water-bong

Copy link
Copy Markdown
Author

Review (standing in for PR-Agent while the local model is paused)

Read the dispatcher changes and the kernel structure; the generated lane bodies were spot-read, not traced line by line.

Dispatcher: correct. Formats 24/25 are consistent across the enum, the ggml type map, the config values and kquant_format.

  • The kquant_needs_grid guard is the right rule. A gate/up pair shares one staged buffer, so the kernel declines only when both sides need a codebook and the codebooks differ. A pair with only one grid format still fuses.

Kernels: correct.

  • Both SwiGLU kernels (1-token and 2-8-token) choose grid_format = gate needs grid ? gate : up, which is right given the guard.
  • The staging condition is computed inline from the two config values, so the barrier's region selector stays workgroup-uniform (STRUCTURE/038). Only the fill dispatch goes through @ggml_kquant_grid_fill_for, which is fine because it sits inside the uniform scf.if.
  • The plain mul_mat kernels do the same with one format.
  • The grid is still one buffer.alloca<workgroup> per kernel, as before. The 16-bit IQ2 codes, two per i32 word, fit in the existing 512-word buffer.

Evidence covers the generated code:

  • test-backend-ops 971/971.
  • Decode-path KLD matches the CPU for both formats (IQ2_XXS 1.131 vs 1.141, IQ2_XS 0.774 vs 0.786).
  • Keeping the generators (tools/generate_iq2_*.py) in-tree, with their output verbatim, makes the ~800-constant fills reviewable and regenerable.

Nit: in kquant_decode_f32.loom, the comment "// Bytes per 256 values of a weight format." now sits above @ggml_kquant_iq2xxs_grid_fill, not above @ggml_kquant_block_bytes. The generated insertion went in between them.

Fine to merge once the runnable checks pass.

bong-water-water-bong and others added 2 commits September 30, 2026 11:12
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit d5048ad into 1bit/hrx-vulkan-patched Sep 30, 2026
17 of 38 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant