Skip to content

metal : skip the empty half of the mul_mm_id token tile, load iq2/iq3 codebooks as uint32 - #28301

Merged
ggerganov merged 1 commit into
ggml-org:masterfrom
masterFoad:metal-mmid-halftile-pr
Sep 11, 2026
Merged

ggerganov merged 1 commit into
ggml-org:masterfrom
masterFoad:metal-mmid-halftile-pr

Conversation

@masterFoad

@masterFoad masterFoad commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Overview

kernel_mul_mm_id works on 32 routed token rows at a time. An expert often receives fewer than 32 rows, but the unused rows are still multiplied and discarded.

This splits the token tile into two NR1H = 16 halves and skips the upper half when nr1 <= NR1H.

On the tensor-ops path, each half gets its own matmul2d call. On the simdgroup path, the two simdgroups responsible for rows 16..31 are disabled through sg_active when the upper half is not needed.

nr1 is the number of rows remaining in the current tile. An expert with 33 routed rows therefore gets one full tile and a 1-row tail, and the tail skips its upper half.

The tB extents are also corrected for the split tile. They were (NR1, NK) for a [NR1][NK] row-major tile, which was harmless while both dimensions were 32.

For reference, mean rows per expert are:

ubatch * n_expert_used / n_expert

A 256-expert top-8 model therefore averages 8 rows per expert at -ub 256 and 16 at -ub 512.

mul_mm_id is used from 32 tokens up. Below that, MoE uses mul_mv_id, so this kernel is not reached.

The IQ codebook and padding-staging changes from the first revision have been removed. This PR now contains only the tile split, the tB extent fix, and the related tests/benchmark hook.

Performance

Apple M5, against master 5a4d0feca. Four paired runs per case. Negative is faster.

Tensor ops

./build/bin/test-backend-ops perf -o MUL_MAT_ID -b MTL0 \
  -p "type_a=(q4_0|q4_K|iq2_xs|iq2_s|iq3_xxs),.*,n=(64|128|256|512|1024|2048),"
n q4_0 q4_K iq2_xs iq2_s iq3_xxs
64 -4.5% -11.4% -18.1% -16.5% -15.2%
128 -5.4% -8.7% -13.7% -13.5% -11.3%
256 -4.0% -3.8% -7.6% -6.9% -5.8%
512 -3.0% -0.8% -4.4% -4.6% -3.4%
1024 -3.4% -2.0% -3.9% -4.0% -1.8%
2048 +0.2% +1.2% +0.3% -0.7% +0.6%

The largest regression is +1.2% for q4_K at n = 2048. At smaller batch sizes, where experts receive fewer rows, the gain is larger.

iq2_s, iq3_xxs, and n > 512 are not in the upstream perf set. I added them locally for these measurements; the test change is not part of this PR.

Measurement-only test-backend-ops change
@@ static std::vector<std::unique_ptr<test_case>> make_test_cases_perf() {
     // qwen3-30b-a3b
-    for (int bs : {1, 4, 8, 32, 64, 128, 256, 512}) {
-        for (ggml_type type_a : {GGML_TYPE_F32, GGML_TYPE_F16, GGML_TYPE_Q4_0, GGML_TYPE_Q8_0, GGML_TYPE_Q4_K, GGML_TYPE_Q6_K, GGML_TYPE_IQ2_XS}) {
+    for (int bs : {1, 4, 8, 32, 64, 128, 256, 512, 1024, 2048}) {
+        for (ggml_type type_a : {GGML_TYPE_F32, GGML_TYPE_F16, GGML_TYPE_Q4_0, GGML_TYPE_Q8_0, GGML_TYPE_Q4_K, GGML_TYPE_Q6_K, GGML_TYPE_IQ2_XS, GGML_TYPE_IQ2_S, GGML_TYPE_IQ3_XXS}) {
             for (ggml_type type_b : {GGML_TYPE_F32}) {
                 test_cases.emplace_back(new test_mul_mat_id(type_a, type_b, 128, 8, false, 768, bs, 2048));

The same change was made to the 32, 4, false, 1792 loop below it.

Simdgroup path

GGML_METAL_TENSOR_DISABLE=1 ./build/bin/test-backend-ops perf -o MUL_MAT_ID -b MTL0 \
  -p "type_a=(q4_0|q4_K|iq2_xs),.*,n=(32|64|128|256|512),"
mean rows per expert 2 4 8 16 32 64
change -41.2% -40.8% -40.6% -22.6% -12.0% -8.8%

Median over q4_0, q4_K, and iq2_xs.

All 30 case medians improve. 119 of 120 paired runs are faster.

The gain is larger on this path because all four simdgroups otherwise run simdgroup_multiply_accumulate over the full tile regardless of nr1.

End to end

Tiel-Coder-35B-A3B IQ3_XXS, 256 experts top-8, Apple M5, against the same master.

./build/bin/llama-bench -m <model>.gguf \
  -p 512 -n 0 -fa 1 -r 5 -ub 256,512

./build/bin/llama-batched-bench -m <model>.gguf \
  -c 16384 -b 512 -ub 256 -fa on \
  -npp 128 -ntg 32 -npl 1,8,32,64
test throughput change runs faster
pp512, -ub 256 +5.1% 8/8
pp512, -ub 512 +2.9% 8/8
tg64, batch 1 +0.00% 3/6
batched decode, npl = 8 -0.3% 1/4
batched decode, npl = 32 +3.6% 3/4
batched decode, npl = 64 +2.6% 4/4

Positive is faster.

Batch-1 decode is unchanged, as expected, because it uses mul_mv_id.

The batched cases start using mul_mm_id at 32 tokens, which is also where the gain appears.

Testing

./build/bin/test-backend-ops -o MUL_MAT_ID -b MTL0
./build/bin/test-backend-ops -o MUL_MAT    -b MTL0

GGML_METAL_TENSOR_DISABLE=1 ./build/bin/test-backend-ops -o MUL_MAT_ID -b MTL0
GGML_METAL_TENSOR_DISABLE=1 ./build/bin/test-backend-ops -o MUL_MAT    -b MTL0

On 5bda51bf:

  • MUL_MAT_ID: 843/843 passed, 78 not supported
  • MUL_MAT: 1265/1265 passed, 417 not supported

Both pass with tensor ops enabled and disabled.

This PR adds 33 test cases across q4_K, iq2_xs, and f16.

Thirty use n_used == n_mats == 4, so every token routes to every expert and each expert receives exactly n rows.

The new mul_mm_id cases cover the split boundary and tail sizes directly:

n final tile
32 32 rows
33 1 row
47 15 rows
48 16 rows
49 17 rows

The 16- and 17-row cases cover the point where the upper half starts running again.

n = 1, 15, 16, 17, 31 cover the mul_mv_id path. Three additional cases use eight experts with one selected to cover empty experts.

test_mul_mat_id also redraws expert IDs between timed perf iterations through reinit_perf_iter, so perf runs do not repeatedly use the same routing assignment.

Related

#25377, #27370 and #26223 touch nearby code. #26223 changes the same B-tile staging code, so whichever lands second will need a small merge.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, AI was used during implementation and profiling, guided and reviewed by me.

@masterFoad
masterFoad requested review from a team and ggerganov as code owners September 3, 2026 07:09
@masterFoad
masterFoad marked this pull request as draft September 3, 2026 07:10
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 3, 2026

Copy link
Copy Markdown

Hi @masterFoad, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 3 open PRs.

  • AI-generated content: While code is allowed to be generated by AI, please write the PR description and commit messages on your own without the help of AI.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) labels Sep 3, 2026
@ggerganov ggerganov self-assigned this Sep 6, 2026
@masterFoad
masterFoad marked this pull request as ready for review September 8, 2026 16:43
@ggerganov

Copy link
Copy Markdown
Member

The dequantize_iq2_xxs/iq2_xs/iq3_xxs/iq3_s/iq2_s functions now read each codebook entry using one or two uint32 loads instead of 4 to 8 uint8 loads. The ksigns_iq2xs lookup is also replaced by the identity encoded by that table, k | (parity(k) << 7). The result is bit-exact. This relies on little-endian layout, which covers Metal targets. It is the same general idea as #27370 for q8_0.

This part is independent of the tile change and I can split it into a separate PR if that is easier to review.

Let's split the loading changes in a separate PR.

@forforever73

Copy link
Copy Markdown
Contributor

Looks good to me overall. I think we're in good shape once the loading changes are split out.

@masterFoad
masterFoad force-pushed the metal-mmid-halftile-pr branch from 1f20843 to bd351fc Compare September 11, 2026 05:54
Comment on lines 640 to 666
// restage every row unconditionally, as upstream does (out-of-range rows read
// clamped-safe duplicate addresses and are discarded by the final store loop)
{
if (FC_mul_mm_bc_inp) {
for (short i = 0; i < 8; ++i) {
const short sx = (tiitg%NL1);
const short sy = (tiitg/NL1)/8;

const short lx = i;
const short ly = (tiitg/NL1)%8;
//const short lx = (tiitg/NL1)%8;
//const short ly = i;

*(sb + NK*(8*sy + ly) + 8*sx + lx) = loop_k + iy + i < args.ne00 ? (S1) *((device T1 *) y + i) : 0;
}
} else {
const short sx = (tiitg%NL1);
const short sy = (tiitg/NL1)/8;

const short lx = i;
//const short lx = i;
const short ly = (tiitg/NL1)%8;
//const short lx = (tiitg/NL1)%8;
//const short ly = i;

*(sb + NK*(8*sy + ly) + 8*sx + lx) = loop_k + iy + i < args.ne00 ? (S1) *((device T1 *) y + i) : 0;
*(threadgroup S1_2x4 *)(sb + NK*(8*sy + ly) + 8*sx) = (S1_2x4)(*((device T1_2x4 *) y));
}
} else {
const short sx = (tiitg%NL1);
const short sy = (tiitg/NL1)/8;

//const short lx = i;
const short ly = (tiitg/NL1)%8;
//const short lx = (tiitg/NL1)%8;
//const short ly = i;

*(threadgroup S1_2x4 *)(sb + NK*(8*sy + ly) + 8*sx) = (S1_2x4)(*((device T1_2x4 *) y));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This comment + indentation is not really needed - let's keep the code block as it is on master

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After addressing the comment, we can merge

kernel_mul_mm_id splits its NR1 = 32 token tile into two 16-row halves and skips
the upper half when the expert did not fill it, on both the tensor and simdgroup
paths. The tB extents are corrected to (NK, NR1H) for the [NR1][NK] row-major tile.

The B tile is staged unconditionally, as on master: rows past nr1 restage a clamped
duplicate of a valid row, lie in the output-row dimension so they never contribute
to a valid row, and are dropped by the final store loop.

test-backend-ops: re-draw the expert ids between perf iterations of test_mul_mat_id
so MoE perf numbers are not warm-cache, and add token-tile boundary coverage using
n_used == n_mats, which routes every token to every expert so each expert receives
exactly n rows; n = 32, 33, 47, 48, 49 reach mul_mm_id and leave a last tile of 32,
1, 15, 16 and 17 rows.
@masterFoad
masterFoad force-pushed the metal-mmid-halftile-pr branch from bd351fc to e80a322 Compare September 11, 2026 10:53
@masterFoad

Copy link
Copy Markdown
Contributor Author

@ggerganov Done, Thanks!

@ggerganov
ggerganov merged commit 5bda51b into ggml-org:master Sep 11, 2026
28 of 35 checks passed
fairydreaming pushed a commit to fairydreaming/llama.cpp that referenced this pull request Sep 13, 2026
@BrewTestBot BrewTestBot mentioned this pull request Sep 14, 2026
1 task done
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
kernel_mul_mm_id splits its NR1 = 32 token tile into two 16-row halves and skips
the upper half when the expert did not fill it, on both the tensor and simdgroup
paths. The tB extents are corrected to (NK, NR1H) for the [NR1][NK] row-major tile.

The B tile is staged unconditionally, as on master: rows past nr1 restage a clamped
duplicate of a valid row, lie in the output-row dimension so they never contribute
to a valid row, and are dropped by the final store loop.

test-backend-ops: re-draw the expert ids between perf iterations of test_mul_mat_id
so MoE perf numbers are not warm-cache, and add token-tile boundary coverage using
n_used == n_mats, which routes every token to every expert so each expert receives
exactly n rows; n = 32, 33, 47, 48, 49 reach mul_mm_id and leave a last tile of 32,
1, 15, 16 and 17 rows.
quimmedes pushed a commit to quimmedes/cafe-llama.cpp that referenced this pull request Sep 16, 2026
kernel_mul_mm_id splits its NR1 = 32 token tile into two 16-row halves and skips
the upper half when the expert did not fill it, on both the tensor and simdgroup
paths. The tB extents are corrected to (NK, NR1H) for the [NR1][NK] row-major tile.

The B tile is staged unconditionally, as on master: rows past nr1 restage a clamped
duplicate of a valid row, lie in the output-row dimension so they never contribute
to a valid row, and are dropped by the final store loop.

test-backend-ops: re-draw the expert ids between perf iterations of test_mul_mat_id
so MoE perf numbers are not warm-cache, and add token-tile boundary coverage using
n_used == n_mats, which routes every token to every expert so each expert receives
exactly n rows; n = 32, 33, 47, 48, 49 reach mul_mm_id and leave a last tile of 32,
1, 15, 16 and 17 rows.
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
kernel_mul_mm_id splits its NR1 = 32 token tile into two 16-row halves and skips
the upper half when the expert did not fill it, on both the tensor and simdgroup
paths. The tB extents are corrected to (NK, NR1H) for the [NR1][NK] row-major tile.

The B tile is staged unconditionally, as on master: rows past nr1 restage a clamped
duplicate of a valid row, lie in the output-row dimension so they never contribute
to a valid row, and are dropped by the final store loop.

test-backend-ops: re-draw the expert ids between perf iterations of test_mul_mat_id
so MoE perf numbers are not warm-cache, and add token-tile boundary coverage using
n_used == n_mats, which routes every token to every expert so each expert receives
exactly n rows; n = 32, 33, 47, 48, 49 reach mul_mm_id and leave a last tile of 32,
1, 15, 16 and 17 rows.
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
kernel_mul_mm_id splits its NR1 = 32 token tile into two 16-row halves and skips
the upper half when the expert did not fill it, on both the tensor and simdgroup
paths. The tB extents are corrected to (NK, NR1H) for the [NR1][NK] row-major tile.

The B tile is staged unconditionally, as on master: rows past nr1 restage a clamped
duplicate of a valid row, lie in the output-row dimension so they never contribute
to a valid row, and are dropped by the final store loop.

test-backend-ops: re-draw the expert ids between perf iterations of test_mul_mat_id
so MoE perf numbers are not warm-cache, and add token-tile boundary coverage using
n_used == n_mats, which routes every token to every expert so each expert receives
exactly n rows; n = 32, 33, 47, 48, 49 reach mul_mm_id and leave a last tile of 32,
1, 15, 16 and 17 rows.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants