Skip to content

ggml-hrx: MXFP4 weights (shared dequantizer, exact E8M0 scale, known-answer test) - #65

Merged
bong-water-water-bong merged 5 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-mxfp4-v2
Oct 2, 2026
Merged

bong-water-water-bong merged 5 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-mxfp4-v2

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

MXFP4 (gpt-oss's expert format) as HRX weight format 39 in the shared dequantizer.

Changes

  • motifs/dequant_1bit.loom (new, ours): MXFP4 decoder. 17-byte blocks (E8M0 exponent, 16 bytes of E2M1 nibbles, IQ4_NL nibble order), scaled by the exact GGML_E8M0_TO_FP32_HALF built from its f32 bits. The file is self-contained: no func.decl of the inline helpers in dequant.loom, because those declarations don't resolve at kernel link time (LINK/MATERIALIZE: unresolved exact declaration).
  • dequant.loom: tile (136 B / 256 values), row (17 B per block) and f16/f32 value hooks.
  • manifest.json: the new file after dequant_prism.loom in all 87 lists.
  • Format enum, type mapping, config value 39. Dense input sizes are multiples of 32.
  • MUL_MAT_ID now admits only input sizes that are a multiple of 256, which its kernels declare. It used to reuse the dense rule (multiples of 32 for 32-value block formats). gpt-oss-20b (input 2880) then failed in the JIT with CONFIG/INVALID instead of running its experts on the CPU.
  • tests/test-hrx-mxfp4.cpp (new): bit-exact known-answer check of MXFP4 GET_ROWS on HRX0.
    • Runs with no scheduler, so no CPU fallback.
    • Requires a get_rows kernel in the HRX plan.
    • Exponents 120/127/134 and the edges 0/1/2/254; every code in both nibbles; memcmp against dequantize_row_mxfp4.
    • At e = 0 and 1 the scale itself is subnormal (2^-128, 2^-127). The GPU kernels flush it, so those rows come back as same-sign zeros (values at most 12 * 2^-127). The test accepts that for those exponents only.

Tested on HRX0, balanced power mode. Tree: 1bit/hrx-tq (these commits plus TQ1_0/TQ2_0, on d60cc4f).

  • test-hrx-mxfp4: 7 rows x 256 values bit-exact; 448 values with a subnormal scale flushed. The dispatch log shows loom_libs:ggml_get_rows_f32.
  • test-backend-ops: full suite 1051/1051.
    • MXFP4 MUL_MAT 13 OK / 64 not supported, MUL_MAT_ID 12 OK / 62 not supported, GET_ROWS 1 OK / 3 not supported, 0 FAIL.
    • Same not-supported pattern as Q4_0, Q8_0 and Q4_K (broadcast and batched shapes).
  • loom-link --verify=true and ggml-hrx-compile-kernel on the rebased tree (e2b946a) for the kquant and get_rows recipes.

gpt-oss-20b (ggml-org gpt-oss-20b-MXFP4.gguf):

  • It loads and runs on HRX0. Greedy text is correct ("red, green, and blue ... additive color mixing ...").
  • Its MXFP4 tensors are all expert weights with input size 2880, which isn't a multiple of 256, so MUL_MAT_ID declines them and they run on the CPU. On this model MXFP4 doesn't reach HRX yet. That needs a MUL_MAT_ID kernel that handles input sizes in multiples of 32, which is the next step and is not in this PR.
  • With that split: wikitext KLD vs CPU 0.0309, same top token 88.8% (8 chunks of 512); pp512 25.8, tg128 12.6 ± 1.6 tok/s.
  • The KLD is above what I'd expect and isn't explained yet. It doesn't come from MXFP4 on HRX, which isn't used on this model, but I haven't traced it.

🤖 Generated with Claude Code

bong-water-water-bong and others added 5 commits October 2, 2026 06:40
MXFP4 (gpt-oss) as weight format 39: 17-byte blocks of an E8M0 exponent and 16 bytes of E2M1 nibbles,
decoded in motifs/dequant_1bit.loom and hooked into the tile, row and value switches.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Build 2^(e - 128) from its f32 bits as GGML_E8M0_TO_FP32_HALF does, instead of expf<afn>, so MXFP4 dequant
matches the CPU bit for bit (review).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…WS on HRX

E8M0 exponents 120/127/134 and the edges 0/1/2/254, every code in both nibbles, compared with
dequantize_row_mxfp4 by memcmp. Runs on HRX0 directly and requires a get_rows kernel in the HRX plan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… subnormal

At e = 0 and 1 the scale itself (2^-128, 2^-127) is subnormal and the GPU kernels flush it, so those rows come
back as same-sign zeros, also where the CPU product is a normal float (FLT_MIN). Every other value stays bit-exact.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The MUL_MAT_ID kernels declare input_size mul(256), but the matcher reused the dense rule, which admits
multiples of 32 for the 32-value block formats. gpt-oss-20b (MXFP4 experts, input 2880) then failed in the
JIT (CONFIG/INVALID) instead of running its experts on the CPU.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong

Copy link
Copy Markdown
Author

Review (PR-Agent duty): approve.

  • MXFP4 in the shared dequantizer, in our own file (motifs/dequant_1bit.loom, full notice), with minimal hooks in AMD's dequant.loom, the enum/config/type maps and the manifest.
  • 17 bytes per 32 values, 136-byte tiles, row bytes = blocks x 17, ggml's nibble order.
  • The E8M0 half scale is built from bits exactly as GGML_E8M0_TO_FP32_HALF does. tests/test-hrx-mxfp4.cpp checks it bit for bit on HRX0, and requires an HRX dispatch plan, so a CPU fallback can't pass. The only allowed difference is a same-sign zero for the subnormal scales (e = 0/1, values about 1e-38); the GPU flushes those.
  • The MUL_MAT_ID guard (input_size % 256, which the kernels declare) is a correctness fix: 2880-wide experts used to be accepted and then fail in the JIT.
  • Full suite 1051/1051, MXFP4 ops 0 FAIL.

Not covered yet (stated in the PR): on gpt-oss-20b all 72 MXFP4 tensors are 2880-wide experts, so they run on the CPU. The follow-up (MUL_MAT_ID for multiples of 32) is the next MXFP4 item. Merging when the hosted jobs pass.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant