Was 4 commits behind 1bit/hrx-vulkan-patched, missing every landed fix for
both open HRX issues:
- engine#108 (dense on HRX0, MoE experts on another device): route-ids
binding-length fix (llama.cpp#7) plus the fused router launch-geometry fix
(llama.cpp#12) that PR #7's description claimed was included but wasn't -
verified separately this session against the exact repro
(-dev HRX0,Vulkan0 -ts 1,0 -ot exps=Vulkan0, Qwen3-Coder-30B-A3B): correct
and deterministic across repeated prompts, disabling the fused dispatch
reproduces the issue's documented fallback error exactly.
- engine#115 (HRX flash-attention decode-split all_rejected above 2048 KV
tokens): capacity-cap fix (llama.cpp#9) - falls through to the general
wmma kernel above the cap instead of crashing.
Built onebit (-DONEBIT_HRX=ON) against the new pin and re-verified both
fixes end to end before this bump.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Bumps
third_party/llama.cppto pick up every landed fix for both currently-open HRX issues. The pin was 4 commits behind1bit/hrx-vulkan-patched.Closes #108, closes #115.
engine#108 — dense on HRX0, MoE experts on another device (fused router)
Two independent bugs, both now fixed on the upstream fork:
Verified against the exact repro (
-dev HRX0,Vulkan0 -ts 1,0 -ot exps=Vulkan0, Qwen3-Coder-30B-A3B-Instruct-Q4_K_M): correct and deterministic across a variety of prompts (arithmetic, factual, multi-turn) withcache_prompt:false, andGGML_HRX_DISABLE_DISPATCH=moe_routerreproduces the issue's documented fallback error ("Compute error.") exactly, confirming the fix targets the dispatch actually being exercised.engine#115 — HRX flash-attention decode-split
all_rejectedabove 2048 KV tokensCapacity-cap fix — llama.cpp#9. The decode-split matcher now declines above the kernel corpus's real 2048-token capacity and falls through to the general (non-split)
flash_attention_f32_f16_wmmadispatch instead of crashing. There is a measured ~32% decode speed cost above that boundary (further declining with depth); a follow-up attempt to recover it with a two-dispatch producer/reducer design was found to have a race/aliasing bug during rigorous re-testing and was retracted (not part of this PR).Verification
Built
onebit(-DONEBIT_HRX=ON) against the bumped pin and re-ran both fixes end to end before opening this PR.🤖 Generated with Claude Code