Repository navigation
ggml-hrx: kernels to run Laya (ModernBERT encoder) on HRX0 - #31
Merged
Merged
Conversation
Laya's typed-decisions checkpoint (ggmlc's GGUF) now runs entirely on HRX: 95.5% on the engine's 200 routing cases, as on Vulkan and the CPU scorer. New LOOM kernels in our small_rows_f32.loom, matchers in our dispatch-small-rows.cpp, all below the existing kernels (priority -10) so they only take what those refuse: - NORM (LayerNorm without affine), one workgroup per row - ADD/SUB/MUL/DIV with strided or broadcast inputs - CLAMP, standalone and in place (ggml_clamp's output views its input) - CPY F32 (strided) -> F16 - FLASH_ATTN_EXT for one-block-per-head layouts (online softmax per lane) - MUL_MAT of small F16/F32 weights (any K) and Q8_0 weights (K % 32) AMD files: NORM in eager_capability_declared, its epsilon in op-params, and the mul_mat_postops matcher no longer takes K % 256 != 0 (its kernel's config constraint rejected those at prepare time). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ggml_attention_rows_f32_f16: one 64-lane workgroup per (query, head) stages the query row and a score row in workgroup memory (each score computed once, not once per value lane), workgroup max/sum reductions, then lanes over the value dims. Up to 2048 keys, head sizes <= 256; the per-lane kernel stays for the rest (ONEBIT_HRX_ATTN_PER_LANE=1 forces it). - ggml_mul_mat_rows_q8_0_f32: Q8_0 with K >= 2048, one workgroup per output, lanes over the K blocks (ONEBIT_HRX_Q8_PER_OUTPUT=1: the old kernel). Laya on HRX0, steady state at 64 tokens: 40 -> 32 ms a decision; 95.5% on the 200 routing cases unchanged. test-backend-ops HRX0: FLASH_ATTN_EXT 204/204 (with plain K/V tensors), MUL_MAT 209/209, NORM 10/10, ADD 22/22. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ggml_rope_rotate_half_f32: the eight-node RoPE graph compilers emit (x*cos + CONT(CONCAT(NEG(CONT(x[half:])), CONT(x[:half])))*sin) in one kernel, matched from its first node (x*cos) so the fused match covers the rest; the two half views must sit at x's offset and half a row further. - ggml_geglu_strided_f32: CONT(gate view) -> GELU -> MUL(., up view) in one kernel, GELU in ggml's tanh form. With ggmlc giving its graphs a uid (1bit-MONSTER/ggmlc 1bit/main), so HRX's graph-program cache and graph replay hit instead of re-importing and re-recording every call: Laya on HRX0 at 64 tokens 32 -> 15-16 ms a decision (Vulkan 12 ms), 128 tokens 25 ms (Vulkan 24-31). 95.5% on the 200 routing cases, probabilities within ~1e-3 of Vulkan's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
merged commit Sep 28, 2026
22ceda3
into
1bit/hrx-vulkan-patched
7 of 24 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This lets the engine's Laya router scorer run on AMD's own stack (HRX + LOOM) rather than on ggmlc's Vulkan build. ggmlc's
laya, built against this ggml withGGML_HRX, now runs Laya's typed-decisions GGUF entirely on HRX0.New LOOM kernels live in our
small_rows_f32.loom, with matchers in ourdispatch-small-rows.cpp. Every matcher sits below the existing kernels (priority -10), so it only takes what those refuse:NORM(LayerNorm without affine), one workgroup per rowADD/SUB/MUL/DIVwith strided or broadcast inputsCLAMP, standalone and in place (ggml_clamp's output is a view of its input)CPYF32 (strided) -> F16FLASH_ATTN_EXTfor one-block-per-head layouts: one lane per value dim, online softmaxMUL_MATof small F16/F32 weights (any K) and Q8_0 weights (K % 32)AMD files get one-line hookups only:
NORMineager_capability_declared, its epsilon inop-params.cpp, and a guard somul_mat_postopsno longer takes K % 256 != 0. Its kernel's config constraint rejects those at prepare time; this showed up as ModernBERT's K = 2624.Verified on Strix Halo:
test-backend-ops -b HRX0: NORM 10/10, CLAMP 3/3, ADD 22/22, MUL 21/21, CPY 52/52, MUL_MAT 209/209. FLASH_ATTN_EXT is 204/204 when the test's K/V are plain tensors. As shipped, the test makes K/V views of a larger cache, and HRX'ssupports_oprefuses a standalone VIEW, so every FA case (AMD's own included) reports "not supported".GGML_HRX_DISABLE_DISPATCH). pp512 is 2,122 vs 2,050 and tg128 48.5 vs 48.1.🤖 Generated with Claude Code