Skip to content

serve --parallel on HRX: shared KV cache, no qwen.attention (Qwen3-4B 4 slots 39 -> 153 tok/s) - #240

Merged
bong-water-water-bong merged 1 commit into
mainfrom
hrx-parallel
Sep 30, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
hrx-parallel

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Makes 1bit serve --device hrx --parallel N fast and correct, and pins llama.cpp acf9c74 (fork #47, #48).

The problem

llama-server gives each slot its own KV stream, so K/V are 4-D. HRX's flash-attention matchers all require 3-D tensors, so attention fell back to the CPU in every layer. Qwen3-4B with 4 slots managed 39 tok/s in total, less than a single stream (76).

The fix (app/serve.cpp, HRX only)

  • -kvu: one shared KV cache across all slots, so attention stays on the GPU.
  • GGML_HRX_DISABLE_DISPATCH=qwen.attention, merged into any value already in the list. AMD's Qwen attention path builds its mask from token positions and ignores the other sequences. Under -kvu that made 3 of 4 parallel answers carry another prompt's content (TCP text in the palindrome, Mars and German answers).
  • Refuses --parallel for qwen35 / qwen35moe / qwen3next (Qwen3.5/3.8), with a message pointing to Vulkan. Multi-sequence batches of those models stop at GATED_DELTA_NET, which has no multi-sequence kernel on HRX yet. Before this they failed mid-request.

Pinned llama.cpp (acf9c74)

Measured on strixhalo

4 slots, greedy, llama-server with the flags serve now passes:

Model Aggregate tok/s before After Single stream Vulkan (4 slots)
Qwen3-4B Q4_K_M 39 153 76 216
Qwen3-Coder-30B-A3B 34 74 86 —
ZAYA1-8B 47 71 79 —
  • Parallel answers stay on topic. The number that exactly match their single-stream run (0–3 of 4) is in the same range as Vulkan's; batching changes the arithmetic on both backends.
  • Single-stream decode is unchanged: Coder-30B 90.4, Qwen3.8-27B 12.17 tok/s.
  • test-backend-ops MUL_MAT and SOFTPLUS pass on HRX0.

docs/hrx.md gets a section on this.

I did not run the rebuilt 1bit binary itself end-to-end, only llama-server with identical flags. serve.cpp passes a syntax-only compile, and CI builds it.

🤖 Generated with Claude Code

…llama.cpp acf9c74

llama-server's per-slot KV streams make K/V 4-D; HRX flash attention then
runs on the CPU in every layer (Qwen3-4B, 4 slots: 39 tok/s aggregate,
below one stream). 1bit serve --device hrx --parallel N now passes -kvu and
turns off AMD's Qwen attention path, which builds its mask from positions and
let the slots' answers bleed into each other under -kvu (measured: 3 of 4
parallel answers taken over by another prompt). Qwen3-4B 4 slots 39 -> 153
tok/s (Vulkan 216), Qwen3-Coder-30B-A3B 34 -> 74, ZAYA1-8B 47 -> 71, answers
on topic. Gated delta-net models (qwen35, qwen35moe, qwen3next) are refused
with --parallel on HRX: no multi-sequence GATED_DELTA_NET kernel yet.

Pins llama.cpp acf9c74 (#47, #48): token kernels accumulate vectors;
kquant matchers skip q8-only inputs (Qwen3-Coder-30B crashed with
qwen.attention off) and a softplus kernel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 30, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 2561ea6

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

47 - Partially compliant

Compliant requirements:

  • HRX decoding past 256 tokens: The PR fixes the cooperative reducer issue by implementing -kvu for shared KV cache and disabling qwen.attention on HRX, which resolves the attention fallback to CPU issue
  • Determining Unsloth files: The PR includes documentation updates that reference the correct Unsloth quantization files (UD-Q4_K_XL and UD-Q5_K_XL) for various models
  • Blog post rendering: The PR includes changes to docs/hrx.md and the necessary build process to render the blog post

Non-compliant requirements:

  • Measurements (1-3% logit difference from Vulkan at 248-2,040 tokens; chat right at 404-2,844): The PR does not include actual measured data to confirm the claimed logit differences

Requires further human verification:

  • The claimed performance improvements (e.g., Qwen3-4B 4 slots: 153 tok/s aggregate) need to be verified on actual hardware

48 - Partially compliant

Compliant requirements:

  • Added Context7 widget to every page via tools/site.py when CONTEXT7_LIBRARY is set
  • Configured context7.json to index docs/ and blog/ and skip specified directories
  • Implemented logic in tools/site.py to handle the CONTEXT7_LIBRARY variable for widget configuration
  • Added error handling for malformed IDs in the build process

Non-compliant requirements:

  • The actual activation of the Context7 widget after merge (steps 1-4 in the ticket description) requires manual intervention and cannot be verified through code review

Requires further human verification:

  • Verification that the Context7 widget is correctly displayed on all pages after setting CONTEXT7_LIBRARY
  • Confirmation that the indexing configuration in context7.json works as expected
⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Model Compatibility Check

The code introduces a check to prevent using --parallel with Qwen3.5, Qwen3.5Moe, and Qwen3Next models on HRX, but this check only considers the GGUF architecture string. It's possible that other model variants or future models with similar architectures could be missed by this check, potentially leading to runtime failures or incorrect behavior.

if (arch == "qwen35" || arch == "qwen35moe" || arch == "qwen3next")
    throw std::runtime_error("--parallel on --device hrx does not run " + arch +
                             " yet (no multi-sequence gated delta-net on HRX); use --device vulkan");
Environment Variable Handling

The code modifies the GGML_HRX_DISABLE_DISPATCH environment variable by appending ,qwen.attention to it if it already exists. This approach could lead to issues if the environment variable contains other values that are not comma-separated, or if the format is not exactly as expected, potentially causing the dispatch disabling to not work correctly.

for (std::string& e : env)
    if (e.rfind("GGML_HRX_DISABLE_DISPATCH=", 0) == 0) { e += ",qwen.attention"; merged = true; }
if (!merged) env.push_back("GGML_HRX_DISABLE_DISPATCH=qwen.attention");

@bong-water-water-bong
bong-water-water-bong merged commit a346ffb into main Sep 30, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the hrx-parallel branch September 30, 2026 02:36
bong-water-water-bong added a commit that referenced this pull request Sep 30, 2026
…244)

Without --parallel, serve passes no -np and llama-server opens its default
several slots. On HRX0 a second concurrent request then shares a batch with
the first, and Qwen3.5/3.8 (qwen35, qwen35moe, qwen3next; Ternary Bonsai too)
fail with HTTP 500 at the multi-sequence softplus / GATED_DELTA_NET (found by
the leaderboard session: Bonsai GSM8K died on its second request). serve now
passes -np 1 for those architectures on HRX; requests queue instead.
Complements #243 (-np 1 for an explicit --parallel 1) and #240 (--parallel >1
refused for them on HRX).

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant