Repository navigation
ggml-hrx: MUL_MAT_ID at input sizes that are a multiple of 32 (gpt-oss experts on HRX) + decode input row stride - #66
Merged
bong-water-water-bong merged 4 commits intoOct 2, 2026
Conversation
…e of 32 As mul_mat_id_swiglu_f32_f32_wmma already does: whole 256-value tiles with the tail masked (weights past input_size decode as zero, activations load masked), and the row stride from ggml_dequant_weight_row_bytes. input_size mul(32) in the plain and postops kernels, and the MUL_MAT_ID matcher goes back to the per-format dense rule (the decode and gated kernels already took multiples of 32). gpt-oss-20b (2880) and BlackMamba (1152) experts can run on HRX. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ggml_mul_mat_decode_load_f32_block viewed the input as [token_count] x [ceil(input_size / 256)] x 256, so for input sizes that are a multiple of 32 but not of 256 every token / input row after the first was read from the wrong offset (rows 3072 apart instead of 2880 for gpt-oss). Index it as [token_count] x [input_size]; the masks already cover the last partial tile. Only mul_mat_id_decode_f32_wave64 loads rows past the first (the dense decode kernels take one token), and only once MUL_MAT_ID admits such input sizes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s that are not a multiple of 256 MUL_MAT with 2..8 tokens and MUL_MAT_ID with 1 or 4 input rows per token, K = 2880 / 1152 (2048 as the control), MXFP4 / Q8_0 / Q4_0, on HRX0 against the CPU backend. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…multiple of 32 on HRX K = 2880 (gpt-oss), 1152 (BlackMamba) and 2816 (no tail), MXFP4 / Q8_0 / Q4_0, 1 token (decode kernel) and 7 / 40 tokens (WMMA), against the CPU backend; the HRX plan must contain a mul_mat_id kernel. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Author
|
Review (PR-Agent duty): approve.
KLD 0.028 vs CPU is pre-existing and tracked in engine ggml-org#284 (f16 accumulation). Merging when the hosted jobs pass, then pinning. |
bong-water-water-bong
merged commit Oct 2, 2026
1bd9ee6
into
1bit/hrx-vulkan-patched
13 of 31 checks passed
This was referenced Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
gpt-oss-20b's MXFP4 experts now run on HRX. Their input size is 2880, a multiple of 32 but not of 256, as is BlackMamba's 1152. On gpt-oss: pp512 25.8 → 527 tok/s, tg128 12.6 → 24.8 tok/s, with correct text.
Commits
mul_mat_id_swiglu_f32_f32_wmmaalready uses.motifs/mul_mat_id_f32_f32_wmma_core.loom: loop over whole 256-value tiles with the tail masked. Weights pastinput_sizedecode as zero; activations usevector.load.mask. The row stride comes fromggml_dequant_weight_row_bytes.mul(32)in the plain, postops and postops_next_rmsnorm configs. Their file headers are unchanged.input_size % 256there, which also turned away the decode and gated kernels; both already took multiples of 32.t * input_size.@ggml_mul_mat_decode_load_f32_blockin AMD'sops/mul_mat_f32_f32_decode.loomviewed the input as[token_count] x [ceil(K/256)] x 256. For K % 256 != 0, every token or input row after the first was read from the wrong offset: rows 3072 apart instead of 2880.[token_count] x [input_size](4 loads); the existing masks cover the partial tile.common_is_supported_decode_token_countis== 1), and MUL_MAT_ID at K % 256 != 0 isn't admitted there because of ggml-hrx: MXFP4 weights (shared dequantizer, exact E8M0 scale, known-answer test) #65's guard. It becomes reachable with commit 1, wheremul_mat_id_decode_f32_wave64gets input_rows 4 (the down projection) or 2-4 tokens at K = 2880. That's why it ships here rather than on its own.tests/test-hrx-decode-stride.cpp(new):tests/test-hrx-mul-mat-id-k32.cpp(new): MUL_MAT_ID at K = 2880, 1152 and 2816 (no tail); MXFP4, Q8_0, Q4_0; 1 token (decode kernel) and 7/40 tokens (WMMA); against the CPU.Before (commit 1 without commit 2, HRX0):
mul_mat_id_decode_f32_wave64at K = 2880 gave NMSE 20-55 with input_rows 4 or 2 tokens, for MXFP4 and Q8_0 alike. K = 2048 was fine for every format; at 2880, 1-row single-token cases were fine.After, on this branch's own build (HRX0, balanced power mode):
test-hrx-decode-stride72 cases, 0 failures;test-hrx-mul-mat-id-k32all OK;test-hrx-mxfp4bit-exact.test-backend-opsMUL_MAT 322/322 and MUL_MAT_ID 108/108, three runs each; full suite 1051/1051.iree-test-loom, hrx-system 244cd38,--sanitizer=access, MXFP4): K = 1152 exact, plus a random differential against a scalar reference at 2880 and 1152. Expert 1 is the last expert, and the cases cover its last row and the last token, so an unmasked tail would read past the buffers; the sanitizer found nothing.Not changed here
Pinned models with sizes that are a multiple of 32 but not of 256 (scan of 72 local GGUFs): gpt-oss-20b (2880), BlackMamba-1.5B (1152), Zamba-7B (3712), qwen2.5-0.5B (896), and Qwen3.8-Flash-Next expert FFN 640. On the pin, none of them reach the decode stride bug. ZAYA1-8B is all multiples of 256.
🤖 Generated with Claude Code