Repository navigation
ggml-hrx: NVFP4 weights (shared dequantizer, decode lanes, bit-exact GET_ROWS test) - #75
bong-water-water-bong wants to merge 4 commits into
Conversation
NVFP4 (ggml type 40) is 64-value blocks of four UE4M3 scales, one per 16 values, and 32 bytes of E2M1 codes. The shared dequantizer walks 32-value blocks, so a block is half of an NVFP4 block: row bytes are 18 per 32 values and a 256-value tile is 144 bytes (both set, see #74). The decoder in dequant_1bit.loom reuses the MXFP4 code table and nibble unpacking; the scale is ggml_ue4m3_to_fp32 built from its f32 bits, exact for every byte (0x7F reads as 0, bit 7 is ignored). The HRX format code is 43, not the ggml type id: HRX format 40 is Q4_0. Input sizes are multiples of 64. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An NVFP4 scale group is 16 consecutive values, so NVFP4 joins packed ternary, TQ1_0, TQ2_0 and MXFP4 on the 16-value lanes of kquant_decode_f32: lane l reads block l / 4, sub-block l % 4 (scale d[s], codes qs[8 s..8 s + 7], low nibbles first). 144 bytes per 256 values. Decode projections, SwiGLU pairs and projection + residual ADD on NVFP4 weights now take the kquant decode kernels instead of the generic path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…WS on HRX Sixteen rows of 256 values hold all 256 UE4M3 scale bytes and every E2M1 code in both nibbles. GET_ROWS on HRX must equal ggml's dequantize_row_nvfp4 bit for bit (every NVFP4 value is exact in f32), on the HRX device with no scheduler, and the dispatch plan must contain the get_rows kernel. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review (PR-Agent duty): looks right; one question before merge, and the merge waits until after Sunday's release. Checked:
Question:
Merge timing: engine main is frozen on f5b7f4a for the first release (Sunday 2026-10-04). I'll merge this after the release run, so no fork-tip change lands during the release weekend. |
…2880 and 2048, from one token The NVFP4 format admits input sizes that are a multiple of 64, the kquant decode lanes only multiples of 256. At 320, 1152 and 2880 one token takes the generic ggml_mul_mat_f32_f32_decode_wave64, 2..8 tokens ggml_mul_mat_f32_f32_wmma and MUL_MAT_ID ggml_mul_mat_id_decode_f32_wave64; at 2048 the kquant lanes. All 36 NVFP4 cases match the CPU backend (normalized MSE about 1e-5); 108 cases, 0 failures. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Question answered by a43a4f2, approved. At 320, 1152 and 2880 the kquant lanes step aside and the generic decode, WMMA and MUL_MAT_ID decode paths take over. At 2048 the kquant lanes run it. All 36 NVFP4 cases match the CPU (108 cases, 0 failures). I'll merge after Sunday's release run (fork tip stays frozen for the release weekend). |
|
Kept: format support in the shared dequantizer / GET_ROWS is plumbing, not a new kernel (Q2_0 is the 2-bit format the ternary goal uses). Waits for the release hold like everything else; needs a rebase onto f95f2db before merge. |
NVFP4 weights (ggml type 40) on HRX0. Before this branch, HRX supported none of the NVFP4 cases in test-backend-ops (MUL_MAT 0/76, MUL_MAT_ID 0/73, GET_ROWS 0/4). About 290 GGUF repos on Hugging Face carry NVFP4, including the most downloaded Qwen3.8-27B NVFP4 files.
Commits
5b57e863: NVFP4 in the shared dequantizer. The decoder lives in ourmotifs/dequant_1bit.loomand reuses the MXFP4 code table and nibble unpacking. The scale isggml_ue4m3_to_fp32, built from its f32 bits. It is exact for every byte:0x7Freads as 0 and bit 7 is ignored, as on the CPU. Both row bytes (18 per 32 values) and tile bytes (144 per 256) are set, per ggml-hrx: PrismML tile bytes for the low-token SwiGLU (PQ2_0 1.7B decode) #74. HRX format code 43, not the ggml type id, because HRX's format 40 is Q4_0. My first build used 40, and every Q4_0 case failed with error 1.0. Input sizes must be multiples of 64.61e7fc01: NVFP4 on the single-token kquant decode lanes. An NVFP4 scale group is 16 consecutive values, so it fits@ggml_kquant_run16_lane_partsnext to packed ternary, TQ1_0, TQ2_0 and MXFP4.2dce4e29:tests/test-hrx-nvfp4. GET_ROWS on HRX must equaldequantize_row_nvfp4bit for bit across all 256 UE4M3 scale bytes and every code in both nibbles. Every NVFP4 value is exact in f32 (no subnormals), so nothing is flushed.Validation on this branch, gfx1151, balanced power mode:
llama-quantize --tensor-type), HRX vs CPU, wikitext-2 8 x 512kquant.{mul_mat,swiglu,mul_mat_add}.decode_f32The NVFP4 and Q4_0 control files have the same KLD, so NVFP4 adds no error of its own on HRX.
Open: 27B prefill is slow (107 tok/s). NVFP4 has no dedicated prefill kernel. Prefill takes the generic f16
mul_mat.f32_f32_wmma/mul_mat_swiglu.f32_f32_wmma, the kernels in engine ggml-org#284. A fast path is a follow-up: NVFP4 codes are small integers with one scale per 16 values, so int8 WMMA can use them.Not done: an independent scalar reference and kquant check cases for the NVFP4 decode lane, as #69 added for TQ and MXFP4. The lane is covered by test-backend-ops MUL_MAT at n = 1 and the KLD runs above.
Not for the release pin: engine main stays on f5b7f4a until after Sunday.
🤖 Generated with Claude Code