Repository navigation
hrx: IQ2_XXS and IQ2_XS weights on HRX0 (shared dequantizer, K-quant decode kernels) - #50
Merged
Merged
Conversation
…decode kernels) IQ2_XXS and IQ2_XS had no HRX matmul: the load-time 512-token MUL_MAT check failed, the weights went to CPU_REPACK, and each use split the graph. Unsloth's UD GGUFs below 4 bits carry both. - motifs/dequant.loom: ggml_iq2xxs/iq2xs f16/f32 vector4 decoders (block bytes 66 / 74, formats 24 / 25 in the tile and row byte tables and both decode dispatchers), so prefill WMMA, generic decode and GET_ROWS read them. The grids are 16-bit codes (every grid byte is 8, 25 or 43: 2 bits per value); signs from ksigns = 7 bits + parity. - dispatch-mul-mat-weight-format.h / -common.h: CommonMulMatWeightFormat IQ2_XXS (24) and IQ2_XS (25). - ops/kquant_decode_f32.loom: lane functions for the 1-token and 2-8 token decode kernels, run shape 16; grids staged in workgroup memory once per workgroup like IQ3_S (the barrier condition is computed inline from the format so it stays workgroup-uniform); dispatch-kquant-decode.cpp accepts both, and the SwiGLU matcher only stages a grid when a format needs one. - tools/generate_iq2_dequant_loom.py, generate_iq2_kquant_loom.py: the generators for the grid tables and decoders; their output is verbatim in the two .loom files. Checked on HRX0 (strixhalo, gfx1151), Qwen3-4B requantized Q8_0 -> IQ2_*: - test-backend-ops -b HRX0: 971/971; MUL_MAT iq2_xs 11 OK, iq2_xxs 12 OK (64 other iq2_xxs shapes declined, so they run on the CPU as before). - KLD vs the Q8_0 model (wikitext-2, 6 x 512), decode path (-b 1): IQ2_XXS 1.1306 (CPU 1.141), IQ2_XS 0.7744 (CPU 0.786). - llama-bench HRX0, previous (IQ2 on the CPU) -> this: IQ2_XXS pp512 59 -> 715, tg128 23.0 -> 39.5; IQ2_XS pp512 87 -> 268, tg128 28.0 -> 30.2 (tg: 2 x 6 runs, box locked). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Author
|
Review (standing in for PR-Agent while the local model is paused) Read the dispatcher changes and the kernel structure; the generated lane bodies were spot-read, not traced line by line. Dispatcher: correct. Formats 24/25 are consistent across the enum, the ggml type map, the config values and
Kernels: correct.
Evidence covers the generated code:
Nit: in Fine to merge once the runnable checks pass. |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
merged commit Sep 30, 2026
d5048ad
into
1bit/hrx-vulkan-patched
17 of 38 checks passed
This was referenced Sep 30, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
IQ2_XXS and IQ2_XS had no HRX matmul: the load-time 512-token MUL_MAT check
failed, the weights went to CPU_REPACK, and each use split the graph. Unsloth's
UD GGUFs below 4 bits carry both.
66 / 74, formats 24 / 25 in the tile and row byte tables and both decode
dispatchers), so prefill WMMA, generic decode and GET_ROWS read them. The
grids are 16-bit codes (every grid byte is 8, 25 or 43: 2 bits per value);
signs from ksigns = 7 bits + parity.
IQ2_XXS (24) and IQ2_XS (25).
decode kernels, run shape 16; grids staged in workgroup memory once per
workgroup like IQ3_S (the barrier condition is computed inline from the
format so it stays workgroup-uniform); dispatch-kquant-decode.cpp accepts
both, and the SwiGLU matcher only stages a grid when a format needs one.
generators for the grid tables and decoders; their output is verbatim in the
two .loom files.
Checked on HRX0 (strixhalo, gfx1151), Qwen3-4B requantized Q8_0 -> IQ2_*:
(64 other iq2_xxs shapes declined, so they run on the CPU as before).
IQ2_XXS 1.1306 (CPU 1.141), IQ2_XS 0.7744 (CPU 0.786).
IQ2_XXS pp512 59 -> 715, tg128 23.0 -> 39.5;
IQ2_XS pp512 87 -> 268, tg128 28.0 -> 30.2 (tg: 2 x 6 runs, box locked).
Next gap after this: IQ1_S / IQ1_M.
🤖 Generated with Claude Code