Repository navigation
1bit serve: --parallel, --adaptive, RAG (--embed/--rerank), --mtp-p-min; MMQ-only ROCm build - #42
Merged
Merged
Conversation
…ode at --mtp-max 6 --mtp-p-min 0.5) and the MoE counter-case Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e llama.cpp devices Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…kan and ROCm) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e --adaptive-at requests are in flight Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, none past that) A request's own speculative.n_max still wins. MTP spends the GPU on verification, which is free for one stream and costs throughput once requests batch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… each on its own Vulkan llama-server Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…Q-only measured more accurate) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… request, ROCm batches the rest) Measured: a server with MTP loaded batches at about two thirds of the throughput whatever each request's draft length, so the per-request draft scaling is gone and --adaptive-at defaults to 1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e rerankers did not help here) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A burst of simultaneous requests all read the first backend as empty and piled onto its single slot (measured: Vulkan took 35 of 67 requests at adaptive-at 1). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-mtp alone for one) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
enabled auto-merge (squash)
September 24, 2026 18:50
bong-water-water-bong
added a commit
that referenced
this pull request
Sep 29, 2026
… 72% -> 93% of Vulkan) (#235) llama.cpp #42: one-token projections on Q4_K, Q5_K, Q6_K, IQ4_NL, IQ4_XS and Q8_0 read the GGUF blocks directly (SwiGLU gate/up pair fused, residual add folded into projections). Decode vs Vulkan on gfx1151: Qwen3.8-27B UD-Q4_K_XL 11.60 / 12.49, ZAYA1-8B 92.7 / 93, Qwen3-0.6B 326.6 / 357, Qwen3-Coder-30B-A3B 90.6 / 93.7. docs/hrx.md updated. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The rest of the serve work: #41 was merged after its first two commits, so these twelve are rebased onto main.
--parallel N(llama-server-np, continuous batching). Total tok/s at 16 requests: Qwen3.8-27B 68.0 on ROCm (5.8x one stream), Qwen3-Coder-30B 318.9 on ROCm. Vulkan peaks at 8 (50.9 / 228.2).--adaptive [--adaptive-at N]: the engine grows with the load. A Vulkan backend with MTP (N slots, default 1) takes the lone request; a ROCm backend with 16 slots and no MTP batches the rest. They are separate because a server with MTP loaded batches at about two thirds of the throughput whatever each request's draft length (measured; the per-request draft scaling that tried otherwise is gone). Choosing the backend and reserving its slot is one locked step: without that a burst of requests all read "empty" and piled onto one slot (measured, fixed).--mtp-p-min P, with the full MTP sweep on Qwen3.8-27B in docs/serve.md: default draft 3 = 35.9 / 28.2 / 28.8 tok/s;--mtp-max 6 --mtp-p-min 0.5= 42.0 on code (3.4x). On Qwen3-Coder-30B-A3B no draft model beats no drafting.--embed MODEL→/v1/embeddings,--rerank MODEL→/v1/rerank, each on its own Vulkan llama-server. Measured end to end over the engine's docs: embeddings with a top-8 context answered 3/4 (the fourth gave the same fact in other units); the rerankers tried did not help (Qwen3-Reranker scores everything ~1.0 through llama.cpp).GGML_CUDA_FORCE_MMQ. hipBLAS returns wrong GEMMs on gfx1151 (sgemm silently returns wrong results on gfx1151 (Radeon 8060S): max rel err 1.7e4, no error raised ROCm/rocm-libraries#11530); MMQ-only measured more accurate at the same speed (Qwen3.8-27B UD-Q4_K_XL KLD 0.0064 vs 0.0094).Built and run on Strix Halo;
--adaptiveround 4 (after the race fix) is measuring now and its numbers go into docs/serve.md when done.🤖 Generated with Claude Code