Skip to content

serve: --device hrx decodes without decode-split flash attention by default (#140) - #148

Merged
bong-water-water-bong merged 1 commit into
mainfrom
hrx-no-decode-split
Sep 26, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
hrx-no-decode-split

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Mitigates #140 until the kernel is fixed.

flash_attention_decode_split_next_q8 gives nondeterministic, sometimes wrong attention on HRX0. It also faults the GPU intermittently on Qwen3-Coder-30B-A3B, at the current pin and at 5556bf2 alike, so it isn't a recent regression.

model (Q4_K_M, decode) with decode-split without (new default)
Qwen3-0.6B 169 ± 34 tok/s; identical requests differ by up to 3.66 nats 322 tok/s, deterministic after warm-up
Qwen3-Coder-30B-A3B GPU fault in 2 of 3 runs (80–88 tok/s when it survives) 66–71 tok/s, 5 of 5 runs
ZAYA1-8B 23.5 tok/s 25.5 tok/s

Behaviour:

  • 1bit serve --device hrx starts llama-server with GGML_HRX_DISABLE_DISPATCH=decode_split.
  • A value the user sets for that variable wins.
  • ONEBIT_HRX_DECODE_SPLIT=1 turns the kernel back on, for testing a fix.
  • --prefill-device hrx is unchanged, since it decodes on Vulkan.

Checked on strixhalo:

  • serve_e2e --device hrx passes 3 of 3 on Qwen3-Coder-30B-A3B.
  • The child's environment carries the variable by default, and not with ONEBIT_HRX_DECODE_SPLIT=1.

🤖 Generated with Claude Code

…efault (#140)

flash_attention_decode_split_next_q8 gives nondeterministic, sometimes wrong attention on HRX0
(up to 3.66 nats between identical requests, Qwen3-0.6B) and faulted the GPU in 2 of 3 decode
runs of Qwen3-Coder-30B-A3B, at the current pin and at 5556bf2 alike. Without it decode uses the
flash-attention fallback: deterministic after warm-up, 5 of 5 Qwen3-Coder runs clean, and faster
on Qwen3-0.6B (322 vs 169 tok/s) and ZAYA1-8B (25.5 vs 23.5); Qwen3-Coder gives up ~20% of its
surviving runs' speed (66-71 vs 80-88 tok/s).

1bit serve --device hrx now starts llama-server with GGML_HRX_DISABLE_DISPATCH=decode_split. A
value the user sets wins; ONEBIT_HRX_DECODE_SPLIT=1 turns the kernel back on for testing a fix.
The --prefill-device hrx split is unchanged (it decodes on Vulkan). docs/hrx.md, Known issues.

Checked: serve_e2e --device hrx on Qwen3-Coder-30B-A3B passes 3 of 3; the child's environment
carries the variable by default and not with ONEBIT_HRX_DECODE_SPLIT=1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 26, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit c64f7f3

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

140 - Partially compliant

Compliant requirements:

  • Mitigates nondeterministic and wrong attention output from flash_attention_decode_split_next_q8 on HRX0
  • Makes decode deterministic by default on HRX0
  • Allows user to override the default behavior via GGML_HRX_DISABLE_DISPATCH environment variable
  • Allows testing of the decode-split kernel via ONEBIT_HRX_DECODE_SPLIT=1
  • Documents the change in docs/hrx.md

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Environment variable handling

The code checks for ONEBIT_HRX_DECODE_SPLIT to determine whether to disable the decode-split kernel. However, it only checks if the variable is set to "1", which might lead to unexpected behavior if the variable is set to any other value (including empty string or non-numeric values). This could cause the intended behavior to be bypassed unintentionally.

const char* keep_split = std::getenv("ONEBIT_HRX_DECODE_SPLIT");
if (device == "hrx" && !(keep_split && std::string(keep_split) == "1"))
    env.push_back("GGML_HRX_DISABLE_DISPATCH=decode_split");

@bong-water-water-bong
bong-water-water-bong merged commit 7161c67 into main Sep 26, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the hrx-no-decode-split branch September 26, 2026 16:29
bong-water-water-bong added a commit that referenced this pull request Sep 27, 2026
…verted) (#178)

#148 turned HRX0's decode-split flash attention off because it gave
nondeterministic attention (#140) and faulted MoE models (#123). The cause was
the q8 pack barrier, fixed in llama.cpp 00adc2b (#176). Measured again after
the fix (llama-bench, interleaved runs), decode-split is as fast or faster:
Qwen3-0.6B 140-158 vs 118-122 tok/s at ctx 2100, Qwen3-Coder-30B-A3B 42-62 vs
47-49, ZAYA1-8B 38-42 vs 34-36; within noise at ctx 0.

ONEBIT_HRX_DECODE_SPLIT=0 now turns it off (it used to be =1 to turn it on), and
a GGML_HRX_DISABLE_DISPATCH the user sets still wins. docs/hrx.md carries the
new table.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant