Repository navigation
vulkan: small M matrix optimizations for qwen - #28457
Merged
Merged
Conversation
Allow split_k with small M. Make small vs med tile selection (for coopmat2) depend on M, not just N.
jeffbolznv
force-pushed
the
small_mat_qwen
branch
from
September 9, 2026 19:30
5b0b52f to
3bee02b
Compare
0cc4m
approved these changes
Sep 10, 2026
Contributor
|
A bunch of ARGSORT tests failed completely on the T4 CM2 CI, please check. |
Contributor
Author
This seems to be affecting multiple PRs. Maybe related to the driver update? I haven't been able to reproduce it locally so far. |
Contributor
|
Very likely the driver, yes. CM1 passed on 580 GB10, the 615 T4 failed. |
1 task done
pl752
pushed a commit
to pl752/llama.cpp
that referenced
this pull request
Sep 15, 2026
* vulkan: optimize m=1 mul_mat by swapping A/B * vulkan: Improve small M perf Allow split_k with small M. Make small vs med tile selection (for coopmat2) depend on M, not just N.
zsogitbe
pushed a commit
to zsogitbe/llama.cpp
that referenced
this pull request
Sep 17, 2026
* vulkan: optimize m=1 mul_mat by swapping A/B * vulkan: Improve small M perf Allow split_k with small M. Make small vs med tile selection (for coopmat2) depend on M, not just N.
gaetan-puleo
pushed a commit
to halo-box/strix-llama.cpp
that referenced
this pull request
Oct 3, 2026
The m == 1 operand swap (ggml-org#28457) runs B^T*A as a mat-vec when dst has one row. Extend it to dst->ne[0] <= mul_mat_vec_max_cols: the src0 rows become the mat-vec columns (NUM_COLS = m), the mat-vec writes the result transposed into the split_k scratch buffer, and a strided f32 copy of m*n floats puts it in place. Without this, a [k, m] f32 weight with tiny m goes through mul_mm and is padded to a full tile. The motivating shape is the Flash-Next (qwen4exp) hyper-connection inject projection, MUL_MAT f32 m=4 n=2048 k=10240, 95 calls per 2048-token ubatch. Decode (n <= 8) and batched / broadcast cases keep their existing paths. Tests: small-m cases (m = 2..5, 8, 9; n = 9..2048; the k = 10240 inject shape at n = 1, 4, 8, 9, 512, 2048; f16 src0; a batched case that must not take the route) added to test-backend-ops, plus two perf cases. Skipping the transposing copy fails all 25 routed cases, so the tests see the route. Tested on gfx1151 (RADV) only. Assisted-by: Claude Opus 5.5
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
* vulkan: optimize m=1 mul_mat by swapping A/B * vulkan: Improve small M perf Allow split_k with small M. Make small vs med tile selection (for coopmat2) depend on M, not just N.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Optimize m=1 by swapping A and B matrices. codex rediscovered this then told me ggml-cuda already did this in #26171.
Optimize small m (e.g. m=32) by changing tile size selection heuristic and allowing split_k.
These changes target these buckets which appear in recent qwen models:
Requirements