Skip to content

ggml-hrx: NVFP4 weights (shared dequantizer, decode lanes, bit-exact GET_ROWS test) - #75

Open
bong-water-water-bong wants to merge 4 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-nvfp4
Open

bong-water-water-bong wants to merge 4 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-nvfp4

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

NVFP4 weights (ggml type 40) on HRX0. Before this branch, HRX supported none of the NVFP4 cases in test-backend-ops (MUL_MAT 0/76, MUL_MAT_ID 0/73, GET_ROWS 0/4). About 290 GGUF repos on Hugging Face carry NVFP4, including the most downloaded Qwen3.8-27B NVFP4 files.

Commits

  • 5b57e863: NVFP4 in the shared dequantizer. The decoder lives in our motifs/dequant_1bit.loom and reuses the MXFP4 code table and nibble unpacking. The scale is ggml_ue4m3_to_fp32, built from its f32 bits. It is exact for every byte: 0x7F reads as 0 and bit 7 is ignored, as on the CPU. Both row bytes (18 per 32 values) and tile bytes (144 per 256) are set, per ggml-hrx: PrismML tile bytes for the low-token SwiGLU (PQ2_0 1.7B decode) #74. HRX format code 43, not the ggml type id, because HRX's format 40 is Q4_0. My first build used 40, and every Q4_0 case failed with error 1.0. Input sizes must be multiples of 64.
  • 61e7fc01: NVFP4 on the single-token kquant decode lanes. An NVFP4 scale group is 16 consecutive values, so it fits @ggml_kquant_run16_lane_parts next to packed ternary, TQ1_0, TQ2_0 and MXFP4.
  • 2dce4e29: tests/test-hrx-nvfp4. GET_ROWS on HRX must equal dequantize_row_nvfp4 bit for bit across all 256 UE4M3 scale bytes and every code in both nibbles. Every NVFP4 value is exact in f32 (no subnormals), so nothing is flushed.

Validation on this branch, gfx1151, balanced power mode:

Check Result
test-hrx-nvfp4 (known answer) 16 rows x 256 values bit-exact, all 256 scale bytes
test-hrx-mxfp4 still bit-exact
test-backend-ops NVFP4: MUL_MAT / MUL_MAT_ID / GET_ROWS 12 / 12 / 1 OK, 0 fail; the rest "not supported" (same broadcast/batched shapes as q4_0, q4_K)
test-backend-ops -b HRX0, full suite 1096/1096
Qwen3-0.6B, projections NVFP4 (made with llama-quantize --tensor-type), HRX vs CPU, wikitext-2 8 x 512 KLD 0.00321, same top token 96.5%
same model, projections Q4_0 instead (control) KLD 0.00316, same top token 97.3%
cdiamond/Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF (Apache-2.0, sha256 matches), HRX vs CPU, 8 x 512 KLD 0.00462, same top token 95.8%, PPL 6.657
Routing, both models graph splits = 1, all layers on HRX0, no CPU fallback; decode on kquant.{mul_mat,swiglu,mul_mat_add}.decode_f32
0.6B pp512 / tg128 10,333 / 221 tok/s (Q4_0 control: 12,768 / 282)
27B pp512 / tg128 107 / 8.45 +/- 0.16 tok/s

The NVFP4 and Q4_0 control files have the same KLD, so NVFP4 adds no error of its own on HRX.

Open: 27B prefill is slow (107 tok/s). NVFP4 has no dedicated prefill kernel. Prefill takes the generic f16 mul_mat.f32_f32_wmma / mul_mat_swiglu.f32_f32_wmma, the kernels in engine ggml-org#284. A fast path is a follow-up: NVFP4 codes are small integers with one scale per 16 values, so int8 WMMA can use them.

Not done: an independent scalar reference and kquant check cases for the NVFP4 decode lane, as #69 added for TQ and MXFP4. The lane is covered by test-backend-ops MUL_MAT at n = 1 and the KLD runs above.

Not for the release pin: engine main stays on f5b7f4a until after Sunday.

🤖 Generated with Claude Code

bong-water-water-bong and others added 3 commits October 2, 2026 18:16
NVFP4 (ggml type 40) is 64-value blocks of four UE4M3 scales, one per 16 values, and 32 bytes of E2M1
codes. The shared dequantizer walks 32-value blocks, so a block is half of an NVFP4 block: row bytes are
18 per 32 values and a 256-value tile is 144 bytes (both set, see #74). The decoder in dequant_1bit.loom
reuses the MXFP4 code table and nibble unpacking; the scale is ggml_ue4m3_to_fp32 built from its f32
bits, exact for every byte (0x7F reads as 0, bit 7 is ignored).

The HRX format code is 43, not the ggml type id: HRX format 40 is Q4_0. Input sizes are multiples of 64.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An NVFP4 scale group is 16 consecutive values, so NVFP4 joins packed ternary, TQ1_0, TQ2_0 and MXFP4 on
the 16-value lanes of kquant_decode_f32: lane l reads block l / 4, sub-block l % 4 (scale d[s], codes
qs[8 s..8 s + 7], low nibbles first). 144 bytes per 256 values. Decode projections, SwiGLU pairs and
projection + residual ADD on NVFP4 weights now take the kquant decode kernels instead of the generic path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…WS on HRX

Sixteen rows of 256 values hold all 256 UE4M3 scale bytes and every E2M1 code in both nibbles. GET_ROWS
on HRX must equal ggml's dequantize_row_nvfp4 bit for bit (every NVFP4 value is exact in f32), on the
HRX device with no scheduler, and the dispatch plan must contain the get_rows kernel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong

Copy link
Copy Markdown
Author

Review (PR-Agent duty): looks right; one question before merge, and the merge waits until after Sunday's release.

Checked:

  • Format code 43 (not the ggml id 40, which is HRX's Q4_0), with a comment saying why.
  • row_bytes (K/32 x 18) and tile_bytes (144 per 256 values) are both present, so the low-token SwiGLU won't read row 0 (the ggml-hrx: PrismML tile bytes for the low-token SwiGLU (PQ2_0 1.7B decode) #74 bug).
  • UE4M3 halved scale: e >= 1 gives (8 + m) 2^(e - 11), e = 0 gives m 2^-10, and 0x7F gives 0. This matches ggml, and the all-256-bytes known-answer test pins it down.
  • KLD: Qwen3-0.6B NVFP4 0.00321 vs a Q4_0 control at 0.00316, and the 27B at 0.00462. Splits = 1.
  • Our full Apache header is on test-hrx-nvfp4.cpp.

Question: dispatch-mul-mat-weight-format.h accepts NVFP4 for input_size % 64 == 0, but the kquant decode lane walks 256-value blocks (four 36-byte blocks).

Merge timing: engine main is frozen on f5b7f4a for the first release (Sunday 2026-10-04). I'll merge this after the release run, so no fork-tip change lands during the release weekend.

…2880 and 2048, from one token

The NVFP4 format admits input sizes that are a multiple of 64, the kquant decode lanes only multiples of
256. At 320, 1152 and 2880 one token takes the generic ggml_mul_mat_f32_f32_decode_wave64, 2..8 tokens
ggml_mul_mat_f32_f32_wmma and MUL_MAT_ID ggml_mul_mat_id_decode_f32_wave64; at 2048 the kquant lanes.
All 36 NVFP4 cases match the CPU backend (normalized MSE about 1e-5); 108 cases, 0 failures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong

Copy link
Copy Markdown
Author

Question answered by a43a4f2, approved. At 320, 1152 and 2880 the kquant lanes step aside and the generic decode, WMMA and MUL_MAT_ID decode paths take over. At 2048 the kquant lanes run it. All 36 NVFP4 cases match the CPU (108 cases, 0 failures). I'll merge after Sunday's release run (fork tip stays frozen for the release weekend).

@bong-water-water-bong

Copy link
Copy Markdown
Author

Kept: format support in the shared dequantizer / GET_ROWS is plumbing, not a new kernel (Q2_0 is the 2-bit format the ternary goal uses). Waits for the release hold like everything else; needs a rebase onto f95f2db before merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant