Repository navigation
serve: --device hrx decodes without decode-split flash attention by default (#140) - #148
Merged
Merged
Conversation
…efault (#140) flash_attention_decode_split_next_q8 gives nondeterministic, sometimes wrong attention on HRX0 (up to 3.66 nats between identical requests, Qwen3-0.6B) and faulted the GPU in 2 of 3 decode runs of Qwen3-Coder-30B-A3B, at the current pin and at 5556bf2 alike. Without it decode uses the flash-attention fallback: deterministic after warm-up, 5 of 5 Qwen3-Coder runs clean, and faster on Qwen3-0.6B (322 vs 169 tok/s) and ZAYA1-8B (25.5 vs 23.5); Qwen3-Coder gives up ~20% of its surviving runs' speed (66-71 vs 80-88 tok/s). 1bit serve --device hrx now starts llama-server with GGML_HRX_DISABLE_DISPATCH=decode_split. A value the user sets wins; ONEBIT_HRX_DECODE_SPLIT=1 turns the kernel back on for testing a fix. The --prefill-device hrx split is unchanged (it decodes on Vulkan). docs/hrx.md, Known issues. Checked: serve_e2e --device hrx on Qwen3-Coder-30B-A3B passes 3 of 3; the child's environment carries the variable by default and not with ONEBIT_HRX_DECODE_SPLIT=1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
This was referenced Sep 26, 2026
bong-water-water-bong
added a commit
that referenced
this pull request
Sep 27, 2026
…verted) (#178) #148 turned HRX0's decode-split flash attention off because it gave nondeterministic attention (#140) and faulted MoE models (#123). The cause was the q8 pack barrier, fixed in llama.cpp 00adc2b (#176). Measured again after the fix (llama-bench, interleaved runs), decode-split is as fast or faster: Qwen3-0.6B 140-158 vs 118-122 tok/s at ctx 2100, Qwen3-Coder-30B-A3B 42-62 vs 47-49, ZAYA1-8B 38-42 vs 34-36; within noise at ctx 0. ONEBIT_HRX_DECODE_SPLIT=0 now turns it off (it used to be =1 to turn it on), and a GGML_HRX_DISABLE_DISPATCH the user sets still wins. docs/hrx.md carries the new table. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Mitigates #140 until the kernel is fixed.
flash_attention_decode_split_next_q8gives nondeterministic, sometimes wrong attention on HRX0. It also faults the GPU intermittently on Qwen3-Coder-30B-A3B, at the current pin and at5556bf2alike, so it isn't a recent regression.Behaviour:
1bit serve --device hrxstarts llama-server withGGML_HRX_DISABLE_DISPATCH=decode_split.ONEBIT_HRX_DECODE_SPLIT=1turns the kernel back on, for testing a fix.--prefill-device hrxis unchanged, since it decodes on Vulkan.Checked on strixhalo:
serve_e2e --device hrxpasses 3 of 3 on Qwen3-Coder-30B-A3B.ONEBIT_HRX_DECODE_SPLIT=1.🤖 Generated with Claude Code