Skip to content

blog: speculative decoding, smaller files, one GPU many users - #43

Merged
bong-water-water-bong merged 4 commits into
mainfrom
blog-sep24
Sep 24, 2026
Merged

bong-water-water-bong merged 4 commits into
mainfrom
blog-sep24

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Two posts from today's measurements on Strix Halo; every number is from the runs behind #37, #41 and #42.

  • Speculative decoding: where it pays, and where it does not. Qwen3.8-27B with its MTP head: 2.3-2.9x by default, 3.4x on code at --mtp-max 6 --mtp-p-min 0.5. Qwen3-Coder-30B-A3B (3B active): no draft model beats no drafting across 36 combinations.
  • Smaller files: Unsloth, ROCmFPX and a ternary 27B. Unsloth's UD-Q3_K_XL beats both 4-bit ROCmFPX formats even with the imatrix. Ternary-Bonsai-2-27B is effectively lossless (PQ2_0 on CPU: KL 0.00014) and decodes at 27 tok/s on ROCm (PTQ1_0); PrismML's ROCm kernels are the gap. The hipBLAS trap (sgemm silently returns wrong results on gfx1151 (Radeon 8060S): max rel err 1.7e4, no error raised ROCm/rocm-libraries#11530).

Please read them before merging: they are written in the project's voice. A third post, on serving many users at once (batching and --adaptive), follows when its last load test is in.

🤖 Generated with Claude Code

…h, ROCmFPX, a ternary 27B)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong bong-water-water-bong changed the title blog: speculative decoding, and smaller files (Unsloth vs ROCmFPX vs ternary) blog: speculative decoding, smaller files, one GPU many users Sep 24, 2026
@bong-water-water-bong
bong-water-water-bong enabled auto-merge (squash) September 24, 2026 18:51
@bong-water-water-bong
bong-water-water-bong merged commit 5eec8a6 into main Sep 24, 2026
3 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the blog-sep24 branch September 24, 2026 19:16
bong-water-water-bong added a commit that referenced this pull request Sep 29, 2026
…B 97% of Vulkan, 74 C) (#236)

llama.cpp #43: Q3_K and IQ3_S in the K-quant decode kernels (Qwen3.8-27B
UD-Q4_K_XL 11.60 -> 12.15 tok/s, Vulkan 12.48). llama.cpp #44: long stream
waits sleep instead of spinning a CPU core in ROCr's signal wait (27B decode
71-76 C / ~125 W instead of 83-93 C / 133-137 W, -0.4%). docs/hrx.md updated.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant