Repository navigation
ggml-hrx: MXFP4 weights (shared dequantizer, exact E8M0 scale, known-answer test) - #65
Merged
Merged
Conversation
MXFP4 (gpt-oss) as weight format 39: 17-byte blocks of an E8M0 exponent and 16 bytes of E2M1 nibbles, decoded in motifs/dequant_1bit.loom and hooked into the tile, row and value switches. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Build 2^(e - 128) from its f32 bits as GGML_E8M0_TO_FP32_HALF does, instead of expf<afn>, so MXFP4 dequant matches the CPU bit for bit (review). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…WS on HRX E8M0 exponents 120/127/134 and the edges 0/1/2/254, every code in both nibbles, compared with dequantize_row_mxfp4 by memcmp. Runs on HRX0 directly and requires a get_rows kernel in the HRX plan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… subnormal At e = 0 and 1 the scale itself (2^-128, 2^-127) is subnormal and the GPU kernels flush it, so those rows come back as same-sign zeros, also where the CPU product is a normal float (FLT_MIN). Every other value stays bit-exact. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The MUL_MAT_ID kernels declare input_size mul(256), but the matcher reused the dense rule, which admits multiples of 32 for the 32-value block formats. gpt-oss-20b (MXFP4 experts, input 2880) then failed in the JIT (CONFIG/INVALID) instead of running its experts on the CPU. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Author
|
Review (PR-Agent duty): approve.
Not covered yet (stated in the PR): on gpt-oss-20b all 72 MXFP4 tensors are 2880-wide experts, so they run on the CPU. The follow-up (MUL_MAT_ID for multiples of 32) is the next MXFP4 item. Merging when the hosted jobs pass. |
bong-water-water-bong
merged commit Oct 2, 2026
cebcd70
into
1bit/hrx-vulkan-patched
11 of 26 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
MXFP4 (gpt-oss's expert format) as HRX weight format 39 in the shared dequantizer.
Changes
motifs/dequant_1bit.loom(new, ours): MXFP4 decoder. 17-byte blocks (E8M0 exponent, 16 bytes of E2M1 nibbles, IQ4_NL nibble order), scaled by the exactGGML_E8M0_TO_FP32_HALFbuilt from its f32 bits. The file is self-contained: nofunc.declof the inline helpers indequant.loom, because those declarations don't resolve at kernel link time (LINK/MATERIALIZE: unresolved exact declaration).dequant.loom: tile (136 B / 256 values), row (17 B per block) and f16/f32 value hooks.manifest.json: the new file afterdequant_prism.loomin all 87 lists.CONFIG/INVALIDinstead of running its experts on the CPU.tests/test-hrx-mxfp4.cpp(new): bit-exact known-answer check of MXFP4 GET_ROWS on HRX0.memcmpagainstdequantize_row_mxfp4.Tested on HRX0, balanced power mode. Tree:
1bit/hrx-tq(these commits plus TQ1_0/TQ2_0, on d60cc4f).test-hrx-mxfp4: 7 rows x 256 values bit-exact; 448 values with a subnormal scale flushed. The dispatch log showsloom_libs:ggml_get_rows_f32.test-backend-ops: full suite 1051/1051.loom-link --verify=trueandggml-hrx-compile-kernelon the rebased tree (e2b946a) for the kquant and get_rows recipes.gpt-oss-20b (ggml-org
gpt-oss-20b-MXFP4.gguf):🤖 Generated with Claude Code