Skip to content

sycl: add IQ3_S multi-column MMVQ - #29500

Merged
arthw merged 1 commit into
ggml-org:masterfrom
clemenswasser:sycl-iq3s-multicol-mmvq
Oct 7, 2026
Merged

arthw merged 1 commit into
ggml-org:masterfrom
clemenswasser:sycl-iq3s-multicol-mmvq

Conversation

@clemenswasser

Copy link
Copy Markdown
Contributor

Overview

Port the IQ4_XS switch_ncols path to IQ3_S. Noticed this bottleneck on Qwen3.8-27B IQ3_S-heavy GGUF.

bench before after speedup
17408×5120 N=4 1 978µs (39 GB/s) 361µs (106 GB/s) 2.71x
17408×5120 N=1 1 250µs 249µs 1.00x
inference (Qwen3.8-27B) before after
short-prompt TG 15 tok/s 22 tok/s
22k-ctx TG 9.4 tok/s 10.1 tok/s

All benchmarked on my local Arc A770

Additional information

Extends #21845 to IQ3_S

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: yes, for discovering this slowdown with Qwen3.8-27B inference and initial mirroring/porting the identical code from IQ4_XS paths, I manually reviewed and checked everything

Footnotes

  1. test-backend-ops test -b SYCL0 -o MUL_MAT -p 'iq3_s' 3-run median ↩ ↩2

@clemenswasser
clemenswasser requested a review from a team as a code owner September 26, 2026 21:50
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language labels Sep 26, 2026

@arthw arthw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On B60:

GGUF Test fa Base t/s Primary t/s Increase Rate (Primary vs Base)
Qwen3-30B-A3B-UD-IQ3_XXS.gguf pp8192 0 467.21 478.42 2.40%
Qwen3-30B-A3B-UD-IQ3_XXS.gguf pp8192 1 616.53 636.08 3.17%
Qwen3-30B-A3B-UD-IQ3_XXS.gguf tg128 0 46.83 46.79 -0.09%
Qwen3-30B-A3B-UD-IQ3_XXS.gguf tg128 1 48.95 48.95 0.00%

It's good job!

Thank you!

@arthw
arthw merged commit 78651c4 into ggml-org:master Oct 7, 2026
16 checks passed
anantshri added a commit to anantshri/llama.cpp that referenced this pull request Oct 7, 2026
Adopt ggml-org#29500's plain IQ3_S multi-column MMVQ verbatim; keep the
persistent reorder paths (reordered multicol + single-col decode +
reorder-aware prefill dequant) and IQ-quant grouped MoE MMVQ.
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants