Skip to content

ggml-hrx: TQ1_0 and TQ2_0 weights (decode lanes, select-before-subtract fix, independent reference + check cases) - #69

Merged
bong-water-water-bong merged 4 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-tq-pr
Oct 2, 2026
Merged

bong-water-water-bong merged 4 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-tq-pr

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

TQ1_0 and TQ2_0 weights on HRX0. The kernels are bcloud-f3's; the validation and an independent reference are mine. MXFP4 went in with #65; this branch carries only the TQ work, on the fork tip after #66.

Commits

  • 8bd4f890 (f3): TQ1_0 / TQ2_0 in the shared dequantizer (motifs/dequant_1bit.loom), and single-token kquant decode lanes for TQ1_0, TQ2_0 and MXFP4 (@ggml_kquant_run16_lane_parts).
  • b9c39393 (f3): the TQ1_0 decoders select before they subtract. Index arithmetic is no-wrap, so an unconditional g - 5 / p - 4 let the compiler drop the first groups. This broke GET_ROWS (error 2.4) and MUL_MAT at n = 9 and 64 (error 130).
  • b221e1f4 (mine): an independent scalar reference for TQ1_0, TQ2_0 and MXFP4 (ggml_kquant_ref_value_1bit in kquant_decode_f32.loom). It is written from ggml's dequantize_row_tq1_0 / tq2_0 / mxfp4, not from dequant_1bit.loom. Before this, ggml_kquant_ref_value had no branch for these formats. It adds six check cases: kquant_mul_mat_decode (TQ1_0 with add, TQ2_0, MXFP4) and kquant_mul_mat_decode_tokens for each.

Validation, on this branch (b221e1f4), gfx1151, balanced power mode:

Check TQ1_0 TQ2_0
test-backend-ops GET_ROWS, x3 (TQ entries enabled locally; upstream comments them out) 1 OK, 3 not supported 1 OK, 3 not supported
MUL_MAT, x3 11/11 11/11
MUL_MAT_ID 3 not supported 3 not supported
Ternary-Bonsai-1.7B, HRX vs CPU, wikitext-2 10 x 512 KLD 0.000523, same top token 98.67% KLD 0.000523, same top token 98.67%
Greedy, 48 tokens (llama-server, temperature 0), HRX x3 identical all 3 runs identical all 3 runs
pp512 / tg128 on HRX0 3,542 / 113 tok/s 4,100 / 156 tok/s

Notes on the table:

  • "Not supported" is the same broadcast/batched-shape pattern as q4_0 and q4_K on HRX.
  • TQ1_0 is a lossless requant (llama-quantize --allow-requantize) of the same ternary weights. That is why the CPU perplexity (20.4324) and the KLD match TQ2_0 exactly; the files are 527 vs 590 MiB, at 1.69 vs 2.06 bpw.
  • The greedy text matches the CPU's over the first 90 characters ("Berlin. The capital of Japan is Tokyo. The capital of Brazil is Brasília…"). The full 48-token outputs differ later, as GPU and CPU rounding do.
  • tg128 varied by ±16-28 because other jobs shared the box.

Pending: step 4. The six check cases under iree-test-loom --sanitizer=access have not run yet. The iree-test-loom I had is built from a newer Loom (hrx-system 244cd38), which rejects config.get on rdna3_5 (TARGET/003). It needs a build from the engine-pinned hrx-system 51b1739, scheduled as its own slot. loom-link --verify=true passes for the decode kernels with the new reference.

Model: ewchampion/Ternary-Bonsai-1.7B-TQ2_0-GGUF (Apache-2.0, built from PrismML's ternary Bonsai).

🤖 Generated with Claude Code

bong-water-water-bong and others added 3 commits October 2, 2026 09:50
…1_0, TQ2_0 and MXFP4

TQ1_0 (34) and TQ2_0 (35) decoders in motifs/dequant_1bit.loom, hooked into the shared dequantizer.
ops/kquant_decode_f32.loom gains 16-consecutive-value lanes for TQ1_0, TQ2_0 and MXFP4 (39), routed with
packed ternary (90) through ggml_kquant_run16_lane_parts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The generic TQ1_0 decoder computed g - 5 and p - 4 (and the decode lane l - 10) for every lane and selected the
result away; with no-wrap index arithmetic that let the compiler drop the groups 0..4 branch. Select first,
then subtract. TQ1_0 GET_ROWS and MUL_MAT n = 9 / 64 now pass (were ERR 2.4 / 130).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… decode

ggml_kquant_ref_value had no branch for TQ1_0 (34), TQ2_0 (35) or MXFP4
(39), so the kquant decode kernels had no reference to check them against.
ggml_kquant_ref_value_1bit decodes one value of a 256-value group, written
from ggml dequantize_row_tq1_0 / dequantize_row_tq2_0 / dequantize_row_mxfp4
and independent of motifs/dequant_1bit.loom (the E8M0 half scale is built
from its f32 bits as GGML_E8M0_TO_FP32_HALF). TQ1_0 selects before it
subtracts, so no-wrap index arithmetic cannot fold a region away.

Six check cases against it: ggml_kquant_mul_mat_decode_f32 for TQ1_0 (with
add), TQ2_0 and MXFP4, and ggml_kquant_mul_mat_decode_tokens_f32 for each
(3 tokens). The MXFP4 cases fill bytes 124..129 so every scale is normal and
between 2^-3 and 2^2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the ggml label Oct 2, 2026
@bong-water-water-bong

Copy link
Copy Markdown
Author

Review (PR-Agent duty): approve, conditional on step 4.

  • TQ1_0/TQ2_0 in the shared dequantizer (motifs/dequant_1bit.loom, ours) with minimal hooks in AMD's dequant.loom and the weight-format enum/config/type maps. Single-token K-quant decode lanes for TQ1_0, TQ2_0 and MXFP4 are in our kquant_decode_f32.loom, through the run16 lane parts. The TQ1_0 select-before-subtract fix is included.
  • Checks on b221e1f, balanced mode:
    • test-backend-ops MUL_MAT 11/11 for each type, x3.
    • Ternary-Bonsai-1.7B HRX vs CPU: KLD 0.000523, same top token 98.67%, for TQ1_0 and TQ2_0. The two are identical because TQ1_0 is a lossless requant of the same weights; CPU PPL 20.4324 for both.
    • Greedy repeats are identical on HRX, and the first ~90 characters match the CPU before a late divergence, consistent with the small KLD.
    • pp512 / tg128: TQ1_0 3542 / 113, TQ2_0 4100 / 156 tok/s.
  • Pending: the check cases under iree-test-loom --sanitizer=access (step 4, running now). Merge after those pass x3 and the hosted jobs pass.

… positive i32)

0x7C7C7C7C + 5 * 0x01010101 overflowed i32, so iree-test-loom refused the
MXFP4 cases (OUT_OF_RANGE in the value materializer). Bytes 121..127 keep
every E8M0 scale normal, between 2^-7 and 2^-1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong

Copy link
Copy Markdown
Author

Step 4 done: all six check cases pass, 3 runs each, under iree-test-loom --sanitizer=access.

The iree-test-loom was built from the engine-pinned hrx-system 51b1739 (bazel). The module was linked with loom-link --verify=true --mode=merge from this branch (f2c62989).

case config 3 runs
ggml_kquant_mul_mat_decode_tq1_0_add_case weight_format=34 add=1 token_count=1 pass, pass, pass
ggml_kquant_mul_mat_decode_tq2_0_case 35, add=0, 1 pass x3
ggml_kquant_mul_mat_decode_mxfp4_case 39, add=0, 1 pass x3
ggml_kquant_mul_mat_decode_tokens_tq1_0_case 34, add=1, 3 pass x3
ggml_kquant_mul_mat_decode_tokens_tq2_0_case 35, add=0, 3 pass x3
ggml_kquant_mul_mat_decode_tokens_mxfp4_case 39, add=0, 3 pass x3

Each run reports failed_sample_count: 0 with no sanitizer reports. Input sizes are 512 inputs x 6 output rows.

f2c62989 changes test data only: the MXFP4 cases' fill moved from 0x7C7C7C7C to 0x79797979 + k*0x01010101 (bytes 121..127). The old iota overflowed i32, and iree-test-loom refused it (OUT_OF_RANGE). Every E8M0 scale stays normal, between 2^-7 and 2^-1.

Run note: the reference kernel reads ggml.kquant_decode.token_count, so each run passes it as well as the four ggml.kquant_mul_mat_decode.* configs.

@bong-water-water-bong
bong-water-water-bong merged commit 0b56f47 into 1bit/hrx-vulkan-patched Oct 2, 2026
10 of 24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant