Repository navigation
1bit serve: --mtp (up to 3.4x), --parallel (up to 5.8x), --adaptive (Vulkan + ROCm overflow), --device rocm - #41
Merged
Conversation
…nd --device rocm for any GGUF --mtp <head.gguf> [--mtp-max N] passes llama-server's draft-mtp speculative decoding: the model's MTP head drafts tokens on the same device and the model checks them in one batch. --device rocm now runs any GGUF on the ROCm build (ONEBIT_LEAN_ROCM, ROCmFPX's tree), not only ROCmI4; ROCmI4 files still take its W4A4 path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…x) and two backends at once Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
enabled auto-merge (squash)
September 24, 2026 13:51
This was referenced Sep 24, 2026
Merged
This was referenced Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Serve options, measured on Strix Halo, and the lean docs updated with the round-2 results.
--mtp <head.gguf> [--mtp-max N]: llama-server'sdraft-mtpspeculative decoding on the llama.cpp devices (vulkan, hrx, rocm). The model's own MTP head drafts; the model checks the drafts in one batch, so output quality is the model's.--mtp-p-min P: drafts a token only when the head is at least that sure. The full sweep on Qwen3.8-27B: default draft 3 = 35.9 / 28.2 / 28.8 (all-round);--mtp-max 6 --mtp-p-min 0.5= 42.0 on code (3.4x). On Qwen3-Coder-30B-A3B (3B active) no draft model beats no drafting.--parallel N: continuous batching (llama-server-np). Total tok/s at 16 requests: Qwen3.8-27B 68.0 on ROCm (5.8x), Qwen3-Coder-30B 318.9 on ROCm (3.7x); Vulkan peaks at 8 requests (50.9 / 228.2).--adaptive [--adaptive-at N]: the engine grows with the load. It runs a Vulkan backend (N slots, default 8, with--mtpif given) and a ROCm overflow backend (16 slots) side by side, and routes each request to Vulkan until N are in flight there, then to ROCm. Verified: 12 simultaneous requests, 8 on Vulkan and 4 on ROCm.--device rocm: any GGUF on the ROCm build (ONEBIT_LEAN_ROCM, ROCmFPX's tree), not only ROCmI4.Qwen3.8-27B UD-Q4_K_XL through
1bit serve, decode tok/s on code / prose / short prompts:--mtp--mtp--device vulkan(upstream pin)--device rocmDocs:
docs/serve.md: the MTP section and table;rocmin the device table.docs/lean.md:--leankeeps ROCmI4 for prefill (455 tok/s) and ROCmFP2 for memory.Built and run on Strix Halo (
~/1bit-engine-lean, branch head); the numbers above are from that build.🤖 Generated with Claude Code