Skip to content

ggml-hrx: kernels to run Laya (ModernBERT encoder) on HRX0 - #31

Merged
bong-water-water-bong merged 3 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-laya
Sep 28, 2026
Merged

bong-water-water-bong merged 3 commits into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-laya

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

This lets the engine's Laya router scorer run on AMD's own stack (HRX + LOOM) rather than on ggmlc's Vulkan build. ggmlc's laya, built against this ggml with GGML_HRX, now runs Laya's typed-decisions GGUF entirely on HRX0.

New LOOM kernels live in our small_rows_f32.loom, with matchers in our dispatch-small-rows.cpp. Every matcher sits below the existing kernels (priority -10), so it only takes what those refuse:

  • NORM (LayerNorm without affine), one workgroup per row
  • ADD/SUB/MUL/DIV with strided or broadcast inputs
  • CLAMP, standalone and in place (ggml_clamp's output is a view of its input)
  • CPY F32 (strided) -> F16
  • FLASH_ATTN_EXT for one-block-per-head layouts: one lane per value dim, online softmax
  • MUL_MAT of small F16/F32 weights (any K) and Q8_0 weights (K % 32)

AMD files get one-line hookups only: NORM in eager_capability_declared, its epsilon in op-params.cpp, and a guard so mul_mat_postops no longer takes K % 256 != 0. Its kernel's config constraint rejects those at prepare time; this showed up as ModernBERT's K = 2624.

Verified on Strix Halo:

  • Laya end to end: 95.5% on the engine's 200 routing cases, the same as Vulkan and the CPU scorer. Probabilities match Vulkan within about 1e-3. Warm decision takes 44 ms, against 22 ms on Vulkan (the kernels are simple; speed comes next).
  • test-backend-ops -b HRX0: NORM 10/10, CLAMP 3/3, ADD 22/22, MUL 21/21, CPY 52/52, MUL_MAT 209/209. FLASH_ATTN_EXT is 204/204 when the test's K/V are plain tensors. As shipped, the test makes K/V views of a larger cache, and HRX's supports_op refuses a standalone VIEW, so every FA case (AMD's own included) reports "not supported".
  • No change for existing models: ZAYA1-8B Q4_K_M on HRX0 has the same PPL (38.8839, 20x512) with the new dispatches on and off (GGML_HRX_DISABLE_DISPATCH). pp512 is 2,122 vs 2,050 and tg128 48.5 vs 48.1.

🤖 Generated with Claude Code

Laya's typed-decisions checkpoint (ggmlc's GGUF) now runs entirely on HRX:
95.5% on the engine's 200 routing cases, as on Vulkan and the CPU scorer.
New LOOM kernels in our small_rows_f32.loom, matchers in our
dispatch-small-rows.cpp, all below the existing kernels (priority -10) so they
only take what those refuse:

- NORM (LayerNorm without affine), one workgroup per row
- ADD/SUB/MUL/DIV with strided or broadcast inputs
- CLAMP, standalone and in place (ggml_clamp's output views its input)
- CPY F32 (strided) -> F16
- FLASH_ATTN_EXT for one-block-per-head layouts (online softmax per lane)
- MUL_MAT of small F16/F32 weights (any K) and Q8_0 weights (K % 32)

AMD files: NORM in eager_capability_declared, its epsilon in op-params, and
the mul_mat_postops matcher no longer takes K % 256 != 0 (its kernel's
config constraint rejected those at prepare time).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the ggml label Sep 28, 2026
bong-water-water-bong and others added 2 commits September 28, 2026 17:39
- ggml_attention_rows_f32_f16: one 64-lane workgroup per (query, head)
  stages the query row and a score row in workgroup memory (each score
  computed once, not once per value lane), workgroup max/sum reductions, then
  lanes over the value dims. Up to 2048 keys, head sizes <= 256; the per-lane
  kernel stays for the rest (ONEBIT_HRX_ATTN_PER_LANE=1 forces it).
- ggml_mul_mat_rows_q8_0_f32: Q8_0 with K >= 2048, one workgroup per output,
  lanes over the K blocks (ONEBIT_HRX_Q8_PER_OUTPUT=1: the old kernel).

Laya on HRX0, steady state at 64 tokens: 40 -> 32 ms a decision; 95.5% on
the 200 routing cases unchanged. test-backend-ops HRX0: FLASH_ATTN_EXT 204/204
(with plain K/V tensors), MUL_MAT 209/209, NORM 10/10, ADD 22/22.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ggml_rope_rotate_half_f32: the eight-node RoPE graph compilers emit
  (x*cos + CONT(CONCAT(NEG(CONT(x[half:])), CONT(x[:half])))*sin) in one
  kernel, matched from its first node (x*cos) so the fused match covers the
  rest; the two half views must sit at x's offset and half a row further.
- ggml_geglu_strided_f32: CONT(gate view) -> GELU -> MUL(., up view) in one
  kernel, GELU in ggml's tanh form.

With ggmlc giving its graphs a uid (1bit-MONSTER/ggmlc 1bit/main), so HRX's
graph-program cache and graph replay hit instead of re-importing and
re-recording every call: Laya on HRX0 at 64 tokens 32 -> 15-16 ms a decision
(Vulkan 12 ms), 128 tokens 25 ms (Vulkan 24-31). 95.5% on the 200 routing
cases, probabilities within ~1e-3 of Vulkan's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit 22ceda3 into 1bit/hrx-vulkan-patched Sep 28, 2026
7 of 24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant