Skip to content

1bit serve: --parallel, --adaptive, RAG (--embed/--rerank), --mtp-p-min; MMQ-only ROCm build - #42

Merged
bong-water-water-bong merged 14 commits into
mainfrom
serve-scale
Sep 24, 2026
Merged

bong-water-water-bong merged 14 commits into
mainfrom
serve-scale

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

The rest of the serve work: #41 was merged after its first two commits, so these twelve are rebased onto main.

  • --parallel N (llama-server -np, continuous batching). Total tok/s at 16 requests: Qwen3.8-27B 68.0 on ROCm (5.8x one stream), Qwen3-Coder-30B 318.9 on ROCm. Vulkan peaks at 8 (50.9 / 228.2).
  • --adaptive [--adaptive-at N]: the engine grows with the load. A Vulkan backend with MTP (N slots, default 1) takes the lone request; a ROCm backend with 16 slots and no MTP batches the rest. They are separate because a server with MTP loaded batches at about two thirds of the throughput whatever each request's draft length (measured; the per-request draft scaling that tried otherwise is gone). Choosing the backend and reserving its slot is one locked step: without that a burst of requests all read "empty" and piled onto one slot (measured, fixed).
  • --mtp-p-min P, with the full MTP sweep on Qwen3.8-27B in docs/serve.md: default draft 3 = 35.9 / 28.2 / 28.8 tok/s; --mtp-max 6 --mtp-p-min 0.5 = 42.0 on code (3.4x). On Qwen3-Coder-30B-A3B no draft model beats no drafting.
  • RAG: --embed MODEL → /v1/embeddings, --rerank MODEL → /v1/rerank, each on its own Vulkan llama-server. Measured end to end over the engine's docs: embeddings with a top-8 context answered 3/4 (the fourth gave the same fact in other units); the rerankers tried did not help (Qwen3-Reranker scores everything ~1.0 through llama.cpp).
  • Lean ROCm build: GGML_CUDA_FORCE_MMQ. hipBLAS returns wrong GEMMs on gfx1151 (sgemm silently returns wrong results on gfx1151 (Radeon 8060S): max rel err 1.7e4, no error raised ROCm/rocm-libraries#11530); MMQ-only measured more accurate at the same speed (Qwen3.8-27B UD-Q4_K_XL KLD 0.0064 vs 0.0094).

Built and run on Strix Halo; --adaptive round 4 (after the race fix) is measuring now and its numbers go into docs/serve.md when done.

🤖 Generated with Claude Code

bong-water-water-bong and others added 12 commits September 24, 2026 13:19
…ode at --mtp-max 6 --mtp-p-min 0.5) and the MoE counter-case

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e llama.cpp devices

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…kan and ROCm)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e --adaptive-at requests are in flight

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, none past that)

A request's own speculative.n_max still wins. MTP spends the GPU on verification, which is
free for one stream and costs throughput once requests batch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… each on its own Vulkan llama-server

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…Q-only measured more accurate)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… request, ROCm batches the rest)

Measured: a server with MTP loaded batches at about two thirds of the throughput whatever
each request's draft length, so the per-request draft scaling is gone and --adaptive-at
defaults to 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e rerankers did not help here)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A burst of simultaneous requests all read the first backend as empty and piled onto its
single slot (measured: Vulkan took 35 of 67 requests at adaptive-at 1).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-mtp alone for one)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong enabled auto-merge (squash) September 24, 2026 18:50
@bong-water-water-bong
bong-water-water-bong merged commit c877ffb into main Sep 24, 2026
3 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the serve-scale branch September 24, 2026 18:51
bong-water-water-bong added a commit that referenced this pull request Sep 29, 2026
… 72% -> 93% of Vulkan) (#235)

llama.cpp #42: one-token projections on Q4_K, Q5_K, Q6_K, IQ4_NL, IQ4_XS
and Q8_0 read the GGUF blocks directly (SwiGLU gate/up pair fused, residual
add folded into projections). Decode vs Vulkan on gfx1151: Qwen3.8-27B
UD-Q4_K_XL 11.60 / 12.49, ZAYA1-8B 92.7 / 93, Qwen3-0.6B 326.6 / 357,
Qwen3-Coder-30B-A3B 90.6 / 93.7. docs/hrx.md updated.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant