Skip to content

metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024 - #28692

Merged
ggerganov merged 3 commits into
ggml-org:masterfrom
masterFoad:metal-iq-rowsplit
Sep 11, 2026
Merged

ggerganov merged 3 commits into
ggml-org:masterfrom
masterFoad:metal-iq-rowsplit

Conversation

@masterFoad

Copy link
Copy Markdown
Contributor

Overview

Follow up to #28086, applying the same mul_mv row split from iq3_xxs to the six other IQ kernels with the same thread mapping:

  • iq1_s
  • iq1_m
  • iq2_xxs
  • iq2_xs
  • iq2_s
  • iq3_s

These kernels assign one 32-element chunk to each simdgroup lane with ix = tiisg. When nb32 < 32, some lanes have no chunk to process. At ne00 = 512, 16 of 32 lanes are idle; at ne00 = 256, 24 are idle.

As in #28086, when nb32 < 32 and nb32 divides 32, the idle lanes share the existing chunks and split the nr0 output rows between them.

The split path uses N_R0_*_SPLIT = 8. The non-split path keeps the existing N_R0_* = 4 and the original thread mapping.

K-quants are left out because they use different thread mappings as mentioned before.

Testing

./build/bin/test-backend-ops -o MUL_MAT    -b MTL0
./build/bin/test-backend-ops -o MUL_MAT_ID -b MTL0
  • MUL_MAT: 1256/1256 against CPU
  • MUL_MAT_ID: 810/810 against CPU
  • Covers k = 256 on the split path and k = 768 on the non-split path

Performance

All measurements were run on an Apple M5 24gb (macOS 26.6.2), on AC power, against master f3f1a8f27.

Kernels

./build/bin/test-backend-ops perf -b MTL0 --test-file shapes.txt

shapes.txt contains MUL_MAT cases with src0 = [k, 4096] and src1 = [k, 1], for k in {256, 512, 768, 1024, 4096} across the seven IQ types.

The default perf set does not include these narrow shapes. Its quantized MUL_MAT cases use k >= 2048, so they do not take the split path, I used --test-file shapes.txt to run k = 256, 512, 768, 1024, 4096 across the seven IQ types; the file is included below so the results can be reproduced.

shapes.txt
29 0 4096 1 1 1 0 2 16 256 4096 1 1 0 0 0 0 0 256 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 16 512 4096 1 1 0 0 0 0 0 512 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 16 768 4096 1 1 0 0 0 0 0 768 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 16 1024 4096 1 1 0 0 0 0 0 1024 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 16 4096 4096 1 1 0 0 0 0 0 4096 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 17 256 4096 1 1 0 0 0 0 0 256 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 17 512 4096 1 1 0 0 0 0 0 512 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 17 768 4096 1 1 0 0 0 0 0 768 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 17 1024 4096 1 1 0 0 0 0 0 1024 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 17 4096 4096 1 1 0 0 0 0 0 4096 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 18 256 4096 1 1 0 0 0 0 0 256 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 18 512 4096 1 1 0 0 0 0 0 512 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 18 768 4096 1 1 0 0 0 0 0 768 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 18 1024 4096 1 1 0 0 0 0 0 1024 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 18 4096 4096 1 1 0 0 0 0 0 4096 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 19 256 4096 1 1 0 0 0 0 0 256 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 19 512 4096 1 1 0 0 0 0 0 512 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 19 768 4096 1 1 0 0 0 0 0 768 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 19 1024 4096 1 1 0 0 0 0 0 1024 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 19 4096 4096 1 1 0 0 0 0 0 4096 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 21 256 4096 1 1 0 0 0 0 0 256 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 21 512 4096 1 1 0 0 0 0 0 512 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 21 768 4096 1 1 0 0 0 0 0 768 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 21 1024 4096 1 1 0 0 0 0 0 1024 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 21 4096 4096 1 1 0 0 0 0 0 4096 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 22 256 4096 1 1 0 0 0 0 0 256 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 22 512 4096 1 1 0 0 0 0 0 512 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 22 768 4096 1 1 0 0 0 0 0 768 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 22 1024 4096 1 1 0 0 0 0 0 1024 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 22 4096 4096 1 1 0 0 0 0 0 4096 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 29 256 4096 1 1 0 0 0 0 0 256 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 29 512 4096 1 1 0 0 0 0 0 512 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 29 768 4096 1 1 0 0 0 0 0 768 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 29 1024 4096 1 1 0 0 0 0 0 1024 1 1 1 0 0 0 0 -
29 0 4096 1 1 1 0 2 29 4096 4096 1 1 0 0 0 0 0 4096 1 1 1 0 0 0 0 -

Time per run vs master, negative is faster:

type k=256 k=512
iq1_s -44.2% -20.0%
iq1_m -48.4% -23.6%
iq2_xxs -50.2% -29.9%
iq2_xs -48.2% -28.7%
iq2_s -50.1% -27.2%
iq3_s -53.1% -30.1%

Median across these 12 shapes is -37.2%. Keeping the original N_R0 = 4 and applying only the thread remapping gives -30.9%, so most of the gain comes from the remapping.

iq3_xxs, already fixed by #28086, was measured in the same runs and stayed within 1.5% of master.

For the non-split shapes (k = 768, 1024, 4096, 18 measurements), the median change is +1.0%, with a range of -2.3% to +2.7%. Spread between repeated runs is of similar size, so I cannot distinguish a consistent effect on these shapes. The non-split path keeps the existing N_R0 = 4 and original thread mapping.

End to end

bartowski/Ornith-1.5-35B-A3B-GGUF has 512-wide expert tensors, which use the split path.

./build/bin/llama-bench -m Ornith-1.5-35B-A3B-IQ2_XXS.gguf \
  -fa 1 -ub 512 -p 512 -d 0,4096,16384 -r 3 -n 64 -t 1

The same command was run with the IQ2_XS and IQ2_S model files.

Token generation, master -> patch -> master (ABA), with the patch compared against the mean of the two master runs:

quant tg64 tg64 @ d4096 tg64 @ d16384
IQ2_XXS +3.6% +5.2% +1.7%
IQ2_XS +4.0% +0.8% +3.3%
IQ2_S +2.8% +3.3% +4.5%

All nine token-generation measurements improved, with a median of +3.3%.

Prefill over the same nine cases ranged from -7.8% to +3.7% with no consistent direction. The two master runs differed by as much as 10% on prefill. These batches use mul_mm / mul_mm_id, which this patch does not change.

Batch width

For these expert tensors, Metal switches from mul_mv_id to mul_mm_id when ne21 >= 32, where ne21 is the number of tokens in the batch.

I swept batch widths B = 1, 2, 4, 8, 16, 24, 32, 48, 64 with llama-batched-bench. Every tested width below 32 improved, while results at 32 and above were within run-to-run variation. The same boundary is visible in the master results, where throughput increases from about 90 to 109 tok/s between B = 24 and B = 32.

This also applies to speculative decoding: verify batches below 32 tokens can use the patched mul_mv_id path, while batches of 32 or more use mul_mm_id.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, AI was used during implementation and profiling, guided and reviewed by me.

…< 1024

Generalize the row split from ggml-org#28086 to the six other kernels that use the
same lane-to-block mapping: iq1_s, iq1_m, iq2_xxs, iq2_xs, iq2_s and iq3_s.

Each of them assigns one 32-element chunk per thread, so when a row has
fewer than 32 chunks the rest of the simdgroup is idle. When nb32 < 32 and
nb32 divides 32, 32/nb32 threads now share each chunk and each takes a
slice of the rows, reusing the FC_mul_mv_split function constant and the
dispatch wrapper introduced for iq3_xxs.

The plain path is untouched: wide matrices keep one thread per chunk and
N_R0_<TYPE> = 4. Only the split path uses N_R0_<TYPE>_SPLIT = 8. The
K-quants have the same idle-thread issue but a different lane mapping, so
they are left for a separate change.
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) labels Sep 10, 2026
@ggerganov ggerganov self-assigned this Sep 10, 2026
Comment on lines +1934 to +1936
device const block_iq2_xxs * xr = x + ibl;
device const uint16_t * q2 = xr->qs + 4 * ib;
device const half * dh = &xr->d;
device const uint16_t * q2 = xr->qs + 4 * ib + (uint64_t) row0*args.nb01/2;
device const half * dh = &xr->d + (uint64_t) row0*args.nb01/2;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Wonder, instead of offsetting each individual pointer (q2, dh, etc.) with row0*nb01, can we offset just the xr pointer alone?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, xr is now offset once before the loop instead. Tests still pass.

q2, dh, sc, qh and signs are all derived from xr, so the row slice
offset only has to be applied to xr.
const int ib = ib32 % (QK_K / 32);

device const block_iq2_xxs * xr = x + ibl;
device const block_iq2_xxs * xr = (device const block_iq2_xxs *) ((device const char *) x + (uint64_t) row0*args.nb01) + ibl;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you can actually move the row0 and row1 initializations earlier in the functions - before initializing the x pointer. And directly do the offset into offset0. Should be cleaner and aligned with the existing pattern.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. I moved it; Correctness still passes - I also reran perf after the change: split shapes improved slightly further, from a 37.8% median reduction to 39.1% median, while non-split shapes stayed the same at +0.1%.

Compute row0 and row1 before initializing the source pointers and apply
the row slice directly to offset0.

This keeps x and its derived pointers on the existing path while applying
the split row offset once.
@ggerganov
ggerganov marked this pull request as ready for review September 11, 2026 09:29
@ggerganov
ggerganov requested a review from a team as a code owner September 11, 2026 09:29
@ggerganov
ggerganov merged commit aac8102 into ggml-org:master Sep 11, 2026
27 of 31 checks passed
@BrewTestBot BrewTestBot mentioned this pull request Sep 14, 2026
1 task done
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…< 1024 (ggml-org#28692)

* metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024

Generalize the row split from ggml-org#28086 to the six other kernels that use the
same lane-to-block mapping: iq1_s, iq1_m, iq2_xxs, iq2_xs, iq2_s and iq3_s.

Each of them assigns one 32-element chunk per thread, so when a row has
fewer than 32 chunks the rest of the simdgroup is idle. When nb32 < 32 and
nb32 divides 32, 32/nb32 threads now share each chunk and each takes a
slice of the rows, reusing the FC_mul_mv_split function constant and the
dispatch wrapper introduced for iq3_xxs.

The plain path is untouched: wide matrices keep one thread per chunk and
N_R0_<TYPE> = 4. Only the split path uses N_R0_<TYPE>_SPLIT = 8. The
K-quants have the same idle-thread issue but a different lane mapping, so
they are left for a separate change.

* metal : offset the src0 row pointer once in the iq mul_mv kernels

q2, dh, sc, qh and signs are all derived from xr, so the row slice
offset only has to be applied to xr.

* metal : fold iq mul_mv row split into offset0

Compute row0 and row1 before initializing the source pointers and apply
the row slice directly to offset0.

This keeps x and its derived pointers on the existing path while applying
the split row offset once.
quimmedes pushed a commit to quimmedes/cafe-llama.cpp that referenced this pull request Sep 16, 2026
…< 1024 (ggml-org#28692)

* metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024

Generalize the row split from ggml-org#28086 to the six other kernels that use the
same lane-to-block mapping: iq1_s, iq1_m, iq2_xxs, iq2_xs, iq2_s and iq3_s.

Each of them assigns one 32-element chunk per thread, so when a row has
fewer than 32 chunks the rest of the simdgroup is idle. When nb32 < 32 and
nb32 divides 32, 32/nb32 threads now share each chunk and each takes a
slice of the rows, reusing the FC_mul_mv_split function constant and the
dispatch wrapper introduced for iq3_xxs.

The plain path is untouched: wide matrices keep one thread per chunk and
N_R0_<TYPE> = 4. Only the split path uses N_R0_<TYPE>_SPLIT = 8. The
K-quants have the same idle-thread issue but a different lane mapping, so
they are left for a separate change.

* metal : offset the src0 row pointer once in the iq mul_mv kernels

q2, dh, sc, qh and signs are all derived from xr, so the row slice
offset only has to be applied to xr.

* metal : fold iq mul_mv row split into offset0

Compute row0 and row1 before initializing the source pointers and apply
the row slice directly to offset0.

This keeps x and its derived pointers on the existing path while applying
the split row offset once.
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
…< 1024 (ggml-org#28692)

* metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024

Generalize the row split from ggml-org#28086 to the six other kernels that use the
same lane-to-block mapping: iq1_s, iq1_m, iq2_xxs, iq2_xs, iq2_s and iq3_s.

Each of them assigns one 32-element chunk per thread, so when a row has
fewer than 32 chunks the rest of the simdgroup is idle. When nb32 < 32 and
nb32 divides 32, 32/nb32 threads now share each chunk and each takes a
slice of the rows, reusing the FC_mul_mv_split function constant and the
dispatch wrapper introduced for iq3_xxs.

The plain path is untouched: wide matrices keep one thread per chunk and
N_R0_<TYPE> = 4. Only the split path uses N_R0_<TYPE>_SPLIT = 8. The
K-quants have the same idle-thread issue but a different lane mapping, so
they are left for a separate change.

* metal : offset the src0 row pointer once in the iq mul_mv kernels

q2, dh, sc, qh and signs are all derived from xr, so the row slice
offset only has to be applied to xr.

* metal : fold iq mul_mv row split into offset0

Compute row0 and row1 before initializing the source pointers and apply
the row slice directly to offset0.

This keeps x and its derived pointers on the existing path while applying
the split row offset once.
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…< 1024 (ggml-org#28692)

* metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024

Generalize the row split from ggml-org#28086 to the six other kernels that use the
same lane-to-block mapping: iq1_s, iq1_m, iq2_xxs, iq2_xs, iq2_s and iq3_s.

Each of them assigns one 32-element chunk per thread, so when a row has
fewer than 32 chunks the rest of the simdgroup is idle. When nb32 < 32 and
nb32 divides 32, 32/nb32 threads now share each chunk and each takes a
slice of the rows, reusing the FC_mul_mv_split function constant and the
dispatch wrapper introduced for iq3_xxs.

The plain path is untouched: wide matrices keep one thread per chunk and
N_R0_<TYPE> = 4. Only the split path uses N_R0_<TYPE>_SPLIT = 8. The
K-quants have the same idle-thread issue but a different lane mapping, so
they are left for a separate change.

* metal : offset the src0 row pointer once in the iq mul_mv kernels

q2, dh, sc, qh and signs are all derived from xr, so the row slice
offset only has to be applied to xr.

* metal : fold iq mul_mv row split into offset0

Compute row0 and row1 before initializing the source pointers and apply
the row slice directly to offset0.

This keeps x and its derived pointers on the existing path while applying
the split row offset once.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants