Repository navigation
ggml-hrx: TQ1_0 and TQ2_0 weights (decode lanes, select-before-subtract fix, independent reference + check cases) - #69
Conversation
…1_0, TQ2_0 and MXFP4 TQ1_0 (34) and TQ2_0 (35) decoders in motifs/dequant_1bit.loom, hooked into the shared dequantizer. ops/kquant_decode_f32.loom gains 16-consecutive-value lanes for TQ1_0, TQ2_0 and MXFP4 (39), routed with packed ternary (90) through ggml_kquant_run16_lane_parts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The generic TQ1_0 decoder computed g - 5 and p - 4 (and the decode lane l - 10) for every lane and selected the result away; with no-wrap index arithmetic that let the compiler drop the groups 0..4 branch. Select first, then subtract. TQ1_0 GET_ROWS and MUL_MAT n = 9 / 64 now pass (were ERR 2.4 / 130). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… decode ggml_kquant_ref_value had no branch for TQ1_0 (34), TQ2_0 (35) or MXFP4 (39), so the kquant decode kernels had no reference to check them against. ggml_kquant_ref_value_1bit decodes one value of a 256-value group, written from ggml dequantize_row_tq1_0 / dequantize_row_tq2_0 / dequantize_row_mxfp4 and independent of motifs/dequant_1bit.loom (the E8M0 half scale is built from its f32 bits as GGML_E8M0_TO_FP32_HALF). TQ1_0 selects before it subtracts, so no-wrap index arithmetic cannot fold a region away. Six check cases against it: ggml_kquant_mul_mat_decode_f32 for TQ1_0 (with add), TQ2_0 and MXFP4, and ggml_kquant_mul_mat_decode_tokens_f32 for each (3 tokens). The MXFP4 cases fill bytes 124..129 so every scale is normal and between 2^-3 and 2^2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review (PR-Agent duty): approve, conditional on step 4.
|
… positive i32) 0x7C7C7C7C + 5 * 0x01010101 overflowed i32, so iree-test-loom refused the MXFP4 cases (OUT_OF_RANGE in the value materializer). Bytes 121..127 keep every E8M0 scale normal, between 2^-7 and 2^-1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Step 4 done: all six check cases pass, 3 runs each, under The iree-test-loom was built from the engine-pinned hrx-system
Each run reports
Run note: the reference kernel reads |
0b56f47
into
1bit/hrx-vulkan-patched
TQ1_0 and TQ2_0 weights on HRX0. The kernels are bcloud-f3's; the validation and an independent reference are mine. MXFP4 went in with #65; this branch carries only the TQ work, on the fork tip after #66.
Commits
8bd4f890(f3): TQ1_0 / TQ2_0 in the shared dequantizer (motifs/dequant_1bit.loom), and single-token kquant decode lanes for TQ1_0, TQ2_0 and MXFP4 (@ggml_kquant_run16_lane_parts).b9c39393(f3): the TQ1_0 decoders select before they subtract. Index arithmetic is no-wrap, so an unconditionalg - 5/p - 4let the compiler drop the first groups. This broke GET_ROWS (error 2.4) and MUL_MAT at n = 9 and 64 (error 130).b221e1f4(mine): an independent scalar reference for TQ1_0, TQ2_0 and MXFP4 (ggml_kquant_ref_value_1bitinkquant_decode_f32.loom). It is written from ggml'sdequantize_row_tq1_0/tq2_0/mxfp4, not fromdequant_1bit.loom. Before this,ggml_kquant_ref_valuehad no branch for these formats. It adds six check cases:kquant_mul_mat_decode(TQ1_0 with add, TQ2_0, MXFP4) andkquant_mul_mat_decode_tokensfor each.Validation, on this branch (
b221e1f4), gfx1151, balanced power mode:Notes on the table:
llama-quantize --allow-requantize) of the same ternary weights. That is why the CPU perplexity (20.4324) and the KLD match TQ2_0 exactly; the files are 527 vs 590 MiB, at 1.69 vs 2.06 bpw.Pending: step 4. The six check cases under
iree-test-loom --sanitizer=accesshave not run yet. Theiree-test-loomI had is built from a newer Loom (hrx-system244cd38), which rejectsconfig.geton rdna3_5 (TARGET/003). It needs a build from the engine-pinned hrx-system51b1739, scheduled as its own slot.loom-link --verify=truepasses for the decode kernels with the new reference.Model:
ewchampion/Ternary-Bonsai-1.7B-TQ2_0-GGUF(Apache-2.0, built from PrismML's ternary Bonsai).🤖 Generated with Claude Code