Skip to content

third_party: bump llama.cpp pin for engine#108 and engine#115 HRX fixes - #119

Merged
bong-water-water-bong merged 1 commit into
mainfrom
fix/bump-hrx-108-115
Sep 26, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
fix/bump-hrx-108-115

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Bumps third_party/llama.cpp to pick up every landed fix for both currently-open HRX issues. The pin was 4 commits behind 1bit/hrx-vulkan-patched.

Closes #108, closes #115.

engine#108 — dense on HRX0, MoE experts on another device (fused router)

Two independent bugs, both now fixed on the upstream fork:

  • Route-ids binding-length crash — llama.cpp#7.
  • Wrong output from the fused router's launch geometry (half the experts' logits never computed on gfx1151) — llama.cpp#12. PR Step 2: HRX + Vulkan in one llama.cpp build, wired into Lemonade #7's description claimed this was included; it wasn't (verified by reading the merged diff) — llama.cpp#12 is the real fix, found and applied this session.

Verified against the exact repro (-dev HRX0,Vulkan0 -ts 1,0 -ot exps=Vulkan0, Qwen3-Coder-30B-A3B-Instruct-Q4_K_M): correct and deterministic across a variety of prompts (arithmetic, factual, multi-turn) with cache_prompt:false, and GGML_HRX_DISABLE_DISPATCH=moe_router reproduces the issue's documented fallback error ("Compute error.") exactly, confirming the fix targets the dispatch actually being exercised.

engine#115 — HRX flash-attention decode-split all_rejected above 2048 KV tokens

Capacity-cap fix — llama.cpp#9. The decode-split matcher now declines above the kernel corpus's real 2048-token capacity and falls through to the general (non-split) flash_attention_f32_f16_wmma dispatch instead of crashing. There is a measured ~32% decode speed cost above that boundary (further declining with depth); a follow-up attempt to recover it with a two-dispatch producer/reducer design was found to have a race/aliasing bug during rigorous re-testing and was retracted (not part of this PR).

Verification

Built onebit (-DONEBIT_HRX=ON) against the bumped pin and re-ran both fixes end to end before opening this PR.

🤖 Generated with Claude Code

Was 4 commits behind 1bit/hrx-vulkan-patched, missing every landed fix for
both open HRX issues:

- engine#108 (dense on HRX0, MoE experts on another device): route-ids
  binding-length fix (llama.cpp#7) plus the fused router launch-geometry fix
  (llama.cpp#12) that PR #7's description claimed was included but wasn't -
  verified separately this session against the exact repro
  (-dev HRX0,Vulkan0 -ts 1,0 -ot exps=Vulkan0, Qwen3-Coder-30B-A3B): correct
  and deterministic across repeated prompts, disabling the fused dispatch
  reproduces the issue's documented fallback error exactly.
- engine#115 (HRX flash-attention decode-split all_rejected above 2048 KV
  tokens): capacity-cap fix (llama.cpp#9) - falls through to the general
  wmma kernel above the cap instead of crashing.

Built onebit (-DONEBIT_HRX=ON) against the new pin and re-verified both
fixes end to end before this bump.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 26, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit a4fc66a

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

108 - Partially compliant

Compliant requirements:

  • Fix for binding length issue in MoE router dispatch
  • Fix for wrong output due to fused router launch geometry
  • Verified with exact repro case
  • Confirmed fallback behavior with GGML_HRX_DISABLE_DISPATCH=moe_router

Non-compliant requirements:

  • None

Requires further human verification:

  • End-to-end verification of the fix with the exact repro case

115 - Partially compliant

Compliant requirements:

  • Fix for flash-attention decode-split template resolution failure
  • Prevents crashes by declining above kernel corpus capacity and falling back to general dispatch

Non-compliant requirements:

  • None

Requires further human verification:

  • Verification that the fix works at the exact token threshold (~3800 tokens)
  • Performance testing to confirm ~32% decode speed cost is acceptable

7 - Partially compliant

Compliant requirements:

  • Bumped llama.cpp to include HRX fixes
  • Build configuration with -DONEBIT_HRX=ON is supported
  • Kernel artifacts are byte-identical
  • Speed parity verified
  • Lemonade wiring with hrx_device option implemented

Non-compliant requirements:

  • None

Requires further human verification:

  • End-to-end testing with Lemonade to confirm device routing works correctly
⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Subproject commit update

The PR updates the third_party/llama.cpp subproject commit to include fixes for HRX issues #108 and #115. While this is a necessary change to incorporate the upstream fixes, it's important to verify that all changes in the new commit are compatible with the current engine's usage and that no regressions are introduced. Specifically, ensure that the MoE router fixes and flash attention fixes are correctly applied and tested.

Subproject commit 1e775cdf6483140ad2d852c27413aafaee32ccf3

@bong-water-water-bong
bong-water-water-bong merged commit a88ab61 into main Sep 26, 2026
3 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the fix/bump-hrx-108-115 branch September 26, 2026 09:07
bong-water-water-bong added a commit that referenced this pull request Sep 26, 2026
…bit-MONSTER/llama.cpp #11) (#127)

The pin adds sliding-window attention and a second rope base for ZAYA1-74B-preview on top of
1e775cd (#119). ZAYA1-8B is unchanged: Q4_K_M perplexity 21.5731 on Vulkan before and after,
and test-llama-archs -a zaya passes. ZAYA1-74B-preview Q4_K_M passes serve_e2e through
1bit serve --device vulkan.

registry/architectures.json, regenerated: it records the new pin (it still named 5556bf2
after #119 moved the pin to 1e775cd without regenerating it).

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

1 participant