Repository navigation
Bump llama.cpp to c075cc1: HRX0 runs MoE models above 128 experts (Qwen3.6-35B-A3B) - #84
Merged
Merged
Conversation
…en3.6-35B-A3B) 1bit-MONSTER/llama.cpp#1 on 1bit/hrx-vulkan-patched (merge c075cc1, same tree as the tested 9a7aad6): 96049a2 writes the MoE router's partition table in the layout the mul_mat_id kernels decode once there are more than 128 experts (before: batches >= 2 faulted, decode was wrong, KLD 15.3 vs Vulkan), and 9a7aad6 claims the fused-only nodes (router chain, gated delta net, ...) instead of sending them to the CPU. Qwen3.6-35B-A3B Q8_0 on HRX0: pp512 141 -> 909, tg128 14.1 -> 35.3, KLD vs Vulkan 0.0046; Qwen3-Coder-30B-A3B decodes again (89.4 tok/s); Qwen3-0.6B unchanged; test-backend-ops -b HRX0 790/790. docs/hrx.md: the fix and the remaining known issues. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍(Review updated until commit 05ee06a)Here are some key observations to aid the review process:
|
bong-water-water-bong
enabled auto-merge (squash)
September 25, 2026 14:22
|
Persistent review updated to latest commit 05ee06a |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Moves
third_party/llama.cppfrom79788e90toc075cc1, the merge of 1bit-MONSTER/llama.cpp#1 on1bit/hrx-vulkan-patched. Its tree is identical to the tested9a7aad6.The two fixes
96049a2(256-expert routing): the MoE router wrote its partition table in the 128-expert layout. With more experts:The router now picks the layout by expert count.
9a7aad6(fused-only nodes): our earlier claim rule sent them to the CPU (the router chain, gated delta net and a few more). They are now claimed on HRX.Checked on Strix Halo with the engine built at this pin (
-DONEBIT_HRX=ON)tests/serve_e2e.sh … hrx:llama-benchon HRX0, the 35B Q8_0, fa 1,-r 2, box not idle:From the fork PR, measured on the fork build
test-backend-ops -b HRX0: 790/790.docs/hrx.md: a "Fixed" section, and the known issues updated.-fa offand n_seq > 1 are still open.🤖 Generated with Claude Code