Skip to content

Bump llama.cpp to c075cc1: HRX0 runs MoE models above 128 experts (Qwen3.6-35B-A3B) - #84

Merged
bong-water-water-bong merged 2 commits into
mainfrom
bump-hrx/35b-prefill
Sep 25, 2026
Merged

bong-water-water-bong merged 2 commits into
mainfrom
bump-hrx/35b-prefill

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves third_party/llama.cpp from 79788e90 to c075cc1, the merge of 1bit-MONSTER/llama.cpp#1 on 1bit/hrx-vulkan-patched. Its tree is identical to the tested 9a7aad6.

The two fixes

  • 96049a2 (256-expert routing): the MoE router wrote its partition table in the 128-expert layout. With more experts:

    • batches of 2 or more tokens faulted;
    • decode ran, but its output was wrong (KLD 15.3 against Vulkan).

    The router now picks the layout by expert count.

  • 9a7aad6 (fused-only nodes): our earlier claim rule sent them to the CPU (the router chain, gated delta net and a few more). They are now claimed on HRX.

Checked on Strix Halo with the engine built at this pin (-DONEBIT_HRX=ON)

  • tests/serve_e2e.sh … hrx:

    • Qwen3-0.6B Q4_K_M: PASS;
    • Qwen3.6-35B-A3B Q8_0: PASS (answers Paris, streams). It faulted on the old pin.
  • llama-bench on HRX0, the 35B Q8_0, fa 1, -r 2, box not idle:

    test tok/s
    pp2 43.7 (faulted before)
    pp512 866
    tg32 30.7

From the fork PR, measured on the fork build

model metric before after
Qwen3.6-35B-A3B pp512 141 909
Qwen3.6-35B-A3B tg128 14.1 35.3
Qwen3.6-35B-A3B KLD vs Vulkan 2.63 0.0046
Qwen3-Coder-30B-A3B tg128 fails 89.4
  • test-backend-ops -b HRX0: 790/790.

docs/hrx.md: a "Fixed" section, and the known issues updated. -fa off and n_seq > 1 are still open.

🤖 Generated with Claude Code

…en3.6-35B-A3B)

1bit-MONSTER/llama.cpp#1 on 1bit/hrx-vulkan-patched (merge c075cc1, same tree
as the tested 9a7aad6): 96049a2 writes the MoE router's partition table in the
layout the mul_mat_id kernels decode once there are more than 128 experts
(before: batches >= 2 faulted, decode was wrong, KLD 15.3 vs Vulkan), and
9a7aad6 claims the fused-only nodes (router chain, gated delta net, ...)
instead of sending them to the CPU. Qwen3.6-35B-A3B Q8_0 on HRX0: pp512
141 -> 909, tg128 14.1 -> 35.3, KLD vs Vulkan 0.0046; Qwen3-Coder-30B-A3B
decodes again (89.4 tok/s); Qwen3-0.6B unchanged; test-backend-ops -b HRX0
790/790. docs/hrx.md: the fix and the remaining known issues.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 05ee06a

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit 05ee06a)

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

1 - Partially compliant

Compliant requirements:

  • GGUF reader and dequant implementation
  • Qwen3 decoder in fp32
  • Loading fails on missing/misshaped tensors
  • Golden test against HF transformers fp32 logits
  • CPU reference implementation for correctness oracle

Non-compliant requirements:

  • Thread pool support (not visible in diff)
  • Unit tests on synthetic data (not visible in diff)

Requires further human verification:

  • Verification that thread pool is correctly implemented
  • Confirmation that unit tests are added/updated
⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Incomplete Documentation of Fixes

The documentation for the MoE fixes is incomplete. It mentions the fixes for 256-expert routing and fused-only nodes, but does not clearly state that these fixes are part of the llama.cpp bump. The documentation should explicitly reference the commits 96049a2 and 9a7aad6 and their impact on the HRX backend.

### Fixed: MoE models with more than 128 experts (Qwen3.6-35B-A3B)

Before this pin, Qwen3.6-35B-A3B (256 experts) failed on `HRX0` in two ways:
- **Prompt batches faulted.** Every batch of 2 or more tokens faulted the GPU
  (`HSA_STATUS_ERROR_MEMORY_FAULT`).
- **Decode was wrong.** Single-token decode ran at about 41 tok/s, but its output was
  wrong: KLD 15.3 against Vulkan, top-1 0%.

There were two causes:

1. **The router assumed 128 experts.** `dispatch-moe-router.cpp` wrote the expert
   partition table in the 128-expert layout (7-bit expert ids) for every expert count.
   The `mul_mat_id` kernels that read the table decode a 9-bit layout. So experts
   128-255 got no partitions, and an expert with 2 or more tokens read past its row.
   Fork commit `96049a2` picks the layout by expert count
   (`dispatch_registration/common/dispatch-moe-routing-layout.h`). Nothing changes up
   to 128 experts.
2. **Fused-only ops went to the CPU.** Our earlier claim rule accepted only nodes that
   run standalone. So the router chain, L2_NORM, SOFTPLUS, GATED_DELTA_NET and the
   per-head RMS_NORM/ROPE, which HRX runs only inside fused patterns, went to the CPU.
   The 35B then split per layer (KLD 2.63), and Qwen3-Coder-30B-A3B could not decode.
   Fork commit `9a7aad6` (`fused-context-claim.h`) claims those nodes when their fused
   pattern's producers are present and their shapes fit.

Measured on Strix Halo (llama-bench, fa on, `-r 3`; KLD against Vulkan):

| model | metric | before | this pin |
|---|---|---|---|
| Qwen3.6-35B-A3B Q8_0 | pp512 | 141 | 909 |
| Qwen3.6-35B-A3B Q8_0 | pp2048 | 156 | 1028 |
| Qwen3.6-35B-A3B Q8_0 | tg128 | 14.1 | 35.3 |
| Qwen3.6-35B-A3B Q8_0 | KLD vs Vulkan | 2.63 | 0.0046 |
| Qwen3-Coder-30B-A3B Q4_K_M | pp512 | 1114 | 2040 |
| Qwen3-Coder-30B-A3B Q4_K_M | tg128 | fails | 89.4 |
| Qwen3-0.6B Q4_K_M | pp512 / tg128 | 21938 / 322.5 | 21789 / 323.8 |

`test-backend-ops -b HRX0`: 790/790.

@bong-water-water-bong
bong-water-water-bong enabled auto-merge (squash) September 25, 2026 14:22
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit 05ee06a

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant