Skip to content

ggml-hrx: GET_ROWS of IQ3_S and IQ4_NL rows on HRX (embeddings stay on HRX) - #77

Open
bong-water-water-bong wants to merge 2 commits into
1bit/hrx-q2_0from
1bit/hrx-getrows-iq
Open

bong-water-water-bong wants to merge 2 commits into
1bit/hrx-q2_0from
1bit/hrx-getrows-iq

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

GET_ROWS of IQ3_S and IQ4_NL rows on HRX0, so token embeddings stored in either type stay on HRX. Stacked on the Q2_0 PR (on #75).

Both types were refused twice, in the GET_ROWS matcher and in device_supports_op, even though get_rows_f32 already passes the IQ4_NL table and the IQ3_S grid to the shared dequantizer. IQ4_NL decodes correctly since #41. This removes both blanket refusals. The separate refusal of batched IQ3_S sources ("wrong rows") stays.

Commits

  • 1198e6f0: the two refusals removed (small in-place edits; ggml-hrx.cpp keeps its notice).
  • 97fcffe5: tests/test-hrx-get-rows-iq, 12 x 1024 random rows per type, quantized by ggml_quantize_chunk at scales 2^-6..2^5. HRX must equal dequantize_row_iq4_nl / dequantize_row_iq3_s bit for bit (every value is an f16 scale times small integers).

Validation, gfx1151:

Check IQ4_NL IQ3_S
test-hrx-get-rows-iq bit-exact bit-exact
test-backend-ops GET_ROWS, 3 runs 1 OK, 3 not supported (was 0 OK) 1 OK, 3 not supported (was 0 OK)
Qwen3-0.6B, token_embd in that type (rest Q8_0): graph splits before / after 2 / 1 2 / 1
same, HRX vs CPU KLD (wikitext-2 8 x 512) 0.00309, top 96.9% 0.00312, top 96.4%

The KLD matches HRX's baseline on this model (0.0032 with Q4_0 or NVFP4 projections). The remaining "not supported" cases are the same view/batched shapes q4_0 and q4_K show. Full suite on this branch: 1129/1129.

Not for the release pin.

🤖 Generated with Claude Code

bong-water-water-bong and others added 2 commits October 2, 2026 19:20
The get_rows kernel already hands the IQ4_NL table and the IQ3_S grid to the shared dequantizer, but both
types were refused twice: in the GET_ROWS matcher and in device_supports_op. IQ4_NL decodes correctly
since #41 (no XOR 12 table remap). With both refusals gone, test-backend-ops GET_ROWS passes for the
two-dimensional cases of both types (three runs, error <= 1e-7, the same cases q4_0 and q4_K pass), so
embeddings stored as IQ3_S or IQ4_NL no longer go to the CPU. The separate refusal of batched IQ3_S sources
(wrong rows) stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Twelve rows of 1024 random values per type, quantized by ggml_quantize_chunk at scales 2^-6..2^5. Every
dequantized value is an f16 scale times small integers, so HRX's GET_ROWS must equal dequantize_row_iq4_nl
and dequantize_row_iq3_s bit for bit, on the HRX device with no scheduler, through the get_rows kernel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong

Copy link
Copy Markdown
Author

Review (PR-Agent duty): approved, merging after Sunday's release (stacked on #76). Removing the blanket IQ3_S/IQ4_NL GET_ROWS refusals in the matcher and in device_supports_op is right, because the kernel already decodes both tables; the batched-IQ3_S refusal stays. test-hrx-get-rows-iq is bit-exact for both types; Qwen3-0.6B goes from 2 graph splits to 1 at baseline KLD. The ggml-hrx.cpp change is a 4-line in-place deletion in AMD's file with its notice kept, which is fine under our fork rules. Stack at 97fcffe: test-backend-ops 1129/1129.

@bong-water-water-bong

Copy link
Copy Markdown
Author

Kept: format support in the shared dequantizer / GET_ROWS is plumbing, not a new kernel (Q2_0 is the 2-bit format the ternary goal uses). Waits for the release hold like everything else; needs a rebase onto f95f2db before merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant