Repository navigation
serve --parallel on HRX: shared KV cache, no qwen.attention (Qwen3-4B 4 slots 39 -> 153 tok/s) - #240
Conversation
…llama.cpp acf9c74 llama-server's per-slot KV streams make K/V 4-D; HRX flash attention then runs on the CPU in every layer (Qwen3-4B, 4 slots: 39 tok/s aggregate, below one stream). 1bit serve --device hrx --parallel N now passes -kvu and turns off AMD's Qwen attention path, which builds its mask from positions and let the slots' answers bleed into each other under -kvu (measured: 3 of 4 parallel answers taken over by another prompt). Qwen3-4B 4 slots 39 -> 153 tok/s (Vulkan 216), Qwen3-Coder-30B-A3B 34 -> 74, ZAYA1-8B 47 -> 71, answers on topic. Gated delta-net models (qwen35, qwen35moe, qwen3next) are refused with --parallel on HRX: no multi-sequence GATED_DELTA_NET kernel yet. Pins llama.cpp acf9c74 (#47, #48): token kernels accumulate vectors; kquant matchers skip q8-only inputs (Qwen3-Coder-30B crashed with qwen.attention off) and a softplus kernel. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
…244) Without --parallel, serve passes no -np and llama-server opens its default several slots. On HRX0 a second concurrent request then shares a batch with the first, and Qwen3.5/3.8 (qwen35, qwen35moe, qwen3next; Ternary Bonsai too) fail with HTTP 500 at the multi-sequence softplus / GATED_DELTA_NET (found by the leaderboard session: Bonsai GSM8K died on its second request). serve now passes -np 1 for those architectures on HRX; requests queue instead. Complements #243 (-np 1 for an explicit --parallel 1) and #240 (--parallel >1 refused for them on HRX). Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Makes
1bit serve --device hrx --parallel Nfast and correct, and pins llama.cppacf9c74(fork #47, #48).The problem
llama-server gives each slot its own KV stream, so K/V are 4-D. HRX's flash-attention matchers all require 3-D tensors, so attention fell back to the CPU in every layer. Qwen3-4B with 4 slots managed 39 tok/s in total, less than a single stream (76).
The fix (
app/serve.cpp, HRX only)-kvu: one shared KV cache across all slots, so attention stays on the GPU.GGML_HRX_DISABLE_DISPATCH=qwen.attention, merged into any value already in the list. AMD's Qwen attention path builds its mask from token positions and ignores the other sequences. Under-kvuthat made 3 of 4 parallel answers carry another prompt's content (TCP text in the palindrome, Mars and German answers).--parallelfor qwen35 / qwen35moe / qwen3next (Qwen3.5/3.8), with a message pointing to Vulkan. Multi-sequence batches of those models stop at GATED_DELTA_NET, which has no multi-sequence kernel on HRX yet. Before this they failed mid-request.Pinned llama.cpp (
acf9c74)qwen.attentionoff, Qwen3-Coder-30B crashed on its first request ("reads transient before write"). Docs chat on 1bit.gg: Context7 widget + context7.json #48 also adds a softplus kernel.Measured on strixhalo
4 slots, greedy, llama-server with the flags serve now passes:
test-backend-opsMUL_MAT and SOFTPLUS pass on HRX0.docs/hrx.mdgets a section on this.I did not run the rebuilt
1bitbinary itself end-to-end, only llama-server with identical flags.serve.cpppasses a syntax-only compile, and CI builds it.🤖 Generated with Claude Code