Skip to content

1bit serve: --mtp (up to 3.4x), --parallel (up to 5.8x), --adaptive (Vulkan + ROCm overflow), --device rocm - #41

Merged
bong-water-water-bong merged 3 commits into
mainfrom
serve-mtp-rocm
Sep 24, 2026
Merged

bong-water-water-bong merged 3 commits into
mainfrom
serve-mtp-rocm

Conversation

@bong-water-water-bong

@bong-water-water-bong bong-water-water-bong commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Serve options, measured on Strix Halo, and the lean docs updated with the round-2 results.

--mtp <head.gguf> [--mtp-max N]: llama-server's draft-mtp speculative decoding on the llama.cpp devices (vulkan, hrx, rocm). The model's own MTP head drafts; the model checks the drafts in one batch, so output quality is the model's.

--mtp-p-min P: drafts a token only when the head is at least that sure. The full sweep on Qwen3.8-27B: default draft 3 = 35.9 / 28.2 / 28.8 (all-round); --mtp-max 6 --mtp-p-min 0.5 = 42.0 on code (3.4x). On Qwen3-Coder-30B-A3B (3B active) no draft model beats no drafting.

--parallel N: continuous batching (llama-server -np). Total tok/s at 16 requests: Qwen3.8-27B 68.0 on ROCm (5.8x), Qwen3-Coder-30B 318.9 on ROCm (3.7x); Vulkan peaks at 8 requests (50.9 / 228.2).

--adaptive [--adaptive-at N]: the engine grows with the load. It runs a Vulkan backend (N slots, default 8, with --mtp if given) and a ROCm overflow backend (16 slots) side by side, and routes each request to Vulkan until N are in flight there, then to ROCm. Verified: 12 simultaneous requests, 8 on Vulkan and 4 on ROCm.

--device rocm: any GGUF on the ROCm build (ONEBIT_LEAN_ROCM, ROCmFPX's tree), not only ROCmI4.

Qwen3.8-27B UD-Q4_K_XL through 1bit serve, decode tok/s on code / prose / short prompts:

Route Without --mtp With --mtp
--device vulkan (upstream pin) 12.2 / 12.0 / 12.0 35.0 / 28.3 / 28.9
--device rocm 11.8 / 11.8 / 12.2 38.8 / 20.6 / 16.9

Docs:

  • docs/serve.md: the MTP section and table; rocm in the device table.
  • docs/lean.md:
    • Round 2: ROCmFPX rebuilt with Unsloth's imatrix still trails Unsloth's own UD-Q3_K_XL, which is smaller, faster and more accurate. So the smaller/faster pick is UD-Q3_K_XL on the default route. --lean keeps ROCmI4 for prefill (455 tok/s) and ROCmFP2 for memory.
    • Two backends at once: Vulkan + ROCm decoding at the same time give +18% total throughput (14.3 vs 12.2 tok/s), with each stream slower, so it pays for multi-request serving.

Built and run on Strix Halo (~/1bit-engine-lean, branch head); the numbers above are from that build.

🤖 Generated with Claude Code

…nd --device rocm for any GGUF

--mtp <head.gguf> [--mtp-max N] passes llama-server's draft-mtp speculative decoding:
the model's MTP head drafts tokens on the same device and the model checks them in one
batch. --device rocm now runs any GGUF on the ROCm build (ONEBIT_LEAN_ROCM, ROCmFPX's
tree), not only ROCmI4; ROCmI4 files still take its W4A4 path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…x) and two backends at once

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong enabled auto-merge (squash) September 24, 2026 13:51
@bong-water-water-bong
bong-water-water-bong merged commit c8c679c into main Sep 24, 2026
3 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the serve-mtp-rocm branch September 24, 2026 13:54
@bong-water-water-bong bong-water-water-bong changed the title 1bit serve --mtp (2.4-2.9x on Qwen3.8-27B) and --device rocm for any GGUF 1bit serve --mtp (up to 3.4x), --parallel (continuous batching, up to 5.8x total) and --device rocm Sep 24, 2026
@bong-water-water-bong bong-water-water-bong changed the title 1bit serve --mtp (up to 3.4x), --parallel (continuous batching, up to 5.8x total) and --device rocm 1bit serve: --mtp (up to 3.4x), --parallel (up to 5.8x), --adaptive (Vulkan + ROCm overflow), --device rocm Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant