Skip to content

docs/npu-lax: the HRX0 35B row was from a wrong-output build - #80

Merged
bong-water-water-bong merged 3 commits into
mainfrom
docs/hrx-35b-numbers
Sep 25, 2026
Merged

bong-water-water-bong merged 3 commits into
mainfrom
docs/hrx-35b-numbers

Conversation

@bong-water-water-bong

@bong-water-water-bong bong-water-water-bong commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

docs/npu-lax.md lists HRX0 decoding Qwen3.6-35B-A3B at 40.9 tok/s. That build's output is wrong.

What is wrong with the pinned build (79788e90) on this model:

  • Wrong decode: KLD 15.3 against Vulkan, top-1 0%. The MoE router hard-codes the 128-expert layout, and this model has 256 experts.
  • Prompt batches fault: any batch of 2 or more tokens faults.

Change: the row now shows fork branch 1bit/hrx-35b-prefill, which fixes both. It was measured at 2631b75c, which has the same tree as the pushed 9a7aad66:

pp512 909 tok/s
pp2048 1028 tok/s
tg128 35.3 tok/s
KLD vs Vulkan 0.0046

Note: that branch is pushed to 1bit-MONSTER/llama.cpp, but it isn't in the engine's pin yet.

🤖 Generated with Claude Code

The pinned HRX build's ~41 tok/s decode of Qwen3.6-35B-A3B computes wrong
logits (KLD 15.3 vs Vulkan: the MoE router assumes 128 experts, the model has
256) and its prefill faults. Replace the row with the fork branch that fixes
both (1bit/hrx-35b-prefill, 2631b75c): pp512 909, pp2048 1028, tg128 35.3,
KLD 0.0046 vs Vulkan. The branch is not in the pin yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit b97c173

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit b97c173)

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Incorrect Performance Claim

The PR updates the documentation to reflect that the previous HRX0 35B performance claim of 40.9 tok/s was based on a faulty build. The new entry correctly states that the fixed fork branch achieves 35.3 tok/s, which is a more accurate performance metric. However, the PR description notes that the fork branch is not yet in the engine's pin, which could lead to confusion for users relying on the documentation alone.

| HRX0 (fork branch `1bit/hrx-35b-prefill`, 9a7aad66; not yet in the pin) | 35.3 tok/s | 909 tok/s (pp2048: 1028) | `llama-bench`, llama.cpp fork off 79788e90 |

@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit 62b4447

@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit b97c173

@bong-water-water-bong
bong-water-water-bong merged commit 572a3b9 into main Sep 25, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the docs/hrx-35b-numbers branch September 25, 2026 13:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant