Repository navigation
metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024 - #28692
Conversation
…< 1024 Generalize the row split from ggml-org#28086 to the six other kernels that use the same lane-to-block mapping: iq1_s, iq1_m, iq2_xxs, iq2_xs, iq2_s and iq3_s. Each of them assigns one 32-element chunk per thread, so when a row has fewer than 32 chunks the rest of the simdgroup is idle. When nb32 < 32 and nb32 divides 32, 32/nb32 threads now share each chunk and each takes a slice of the rows, reusing the FC_mul_mv_split function constant and the dispatch wrapper introduced for iq3_xxs. The plain path is untouched: wide matrices keep one thread per chunk and N_R0_<TYPE> = 4. Only the split path uses N_R0_<TYPE>_SPLIT = 8. The K-quants have the same idle-thread issue but a different lane mapping, so they are left for a separate change.
| device const block_iq2_xxs * xr = x + ibl; | ||
| device const uint16_t * q2 = xr->qs + 4 * ib; | ||
| device const half * dh = &xr->d; | ||
| device const uint16_t * q2 = xr->qs + 4 * ib + (uint64_t) row0*args.nb01/2; | ||
| device const half * dh = &xr->d + (uint64_t) row0*args.nb01/2; |
There was a problem hiding this comment.
Wonder, instead of offsetting each individual pointer (q2, dh, etc.) with row0*nb01, can we offset just the xr pointer alone?
There was a problem hiding this comment.
Done, xr is now offset once before the loop instead. Tests still pass.
q2, dh, sc, qh and signs are all derived from xr, so the row slice offset only has to be applied to xr.
| const int ib = ib32 % (QK_K / 32); | ||
|
|
||
| device const block_iq2_xxs * xr = x + ibl; | ||
| device const block_iq2_xxs * xr = (device const block_iq2_xxs *) ((device const char *) x + (uint64_t) row0*args.nb01) + ibl; |
There was a problem hiding this comment.
I think you can actually move the row0 and row1 initializations earlier in the functions - before initializing the x pointer. And directly do the offset into offset0. Should be cleaner and aligned with the existing pattern.
There was a problem hiding this comment.
Done. I moved it; Correctness still passes - I also reran perf after the change: split shapes improved slightly further, from a 37.8% median reduction to 39.1% median, while non-split shapes stayed the same at +0.1%.
Compute row0 and row1 before initializing the source pointers and apply the row slice directly to offset0. This keeps x and its derived pointers on the existing path while applying the split row offset once.
…< 1024 (ggml-org#28692) * metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024 Generalize the row split from ggml-org#28086 to the six other kernels that use the same lane-to-block mapping: iq1_s, iq1_m, iq2_xxs, iq2_xs, iq2_s and iq3_s. Each of them assigns one 32-element chunk per thread, so when a row has fewer than 32 chunks the rest of the simdgroup is idle. When nb32 < 32 and nb32 divides 32, 32/nb32 threads now share each chunk and each takes a slice of the rows, reusing the FC_mul_mv_split function constant and the dispatch wrapper introduced for iq3_xxs. The plain path is untouched: wide matrices keep one thread per chunk and N_R0_<TYPE> = 4. Only the split path uses N_R0_<TYPE>_SPLIT = 8. The K-quants have the same idle-thread issue but a different lane mapping, so they are left for a separate change. * metal : offset the src0 row pointer once in the iq mul_mv kernels q2, dh, sc, qh and signs are all derived from xr, so the row slice offset only has to be applied to xr. * metal : fold iq mul_mv row split into offset0 Compute row0 and row1 before initializing the source pointers and apply the row slice directly to offset0. This keeps x and its derived pointers on the existing path while applying the split row offset once.
…< 1024 (ggml-org#28692) * metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024 Generalize the row split from ggml-org#28086 to the six other kernels that use the same lane-to-block mapping: iq1_s, iq1_m, iq2_xxs, iq2_xs, iq2_s and iq3_s. Each of them assigns one 32-element chunk per thread, so when a row has fewer than 32 chunks the rest of the simdgroup is idle. When nb32 < 32 and nb32 divides 32, 32/nb32 threads now share each chunk and each takes a slice of the rows, reusing the FC_mul_mv_split function constant and the dispatch wrapper introduced for iq3_xxs. The plain path is untouched: wide matrices keep one thread per chunk and N_R0_<TYPE> = 4. Only the split path uses N_R0_<TYPE>_SPLIT = 8. The K-quants have the same idle-thread issue but a different lane mapping, so they are left for a separate change. * metal : offset the src0 row pointer once in the iq mul_mv kernels q2, dh, sc, qh and signs are all derived from xr, so the row slice offset only has to be applied to xr. * metal : fold iq mul_mv row split into offset0 Compute row0 and row1 before initializing the source pointers and apply the row slice directly to offset0. This keeps x and its derived pointers on the existing path while applying the split row offset once.
…< 1024 (ggml-org#28692) * metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024 Generalize the row split from ggml-org#28086 to the six other kernels that use the same lane-to-block mapping: iq1_s, iq1_m, iq2_xxs, iq2_xs, iq2_s and iq3_s. Each of them assigns one 32-element chunk per thread, so when a row has fewer than 32 chunks the rest of the simdgroup is idle. When nb32 < 32 and nb32 divides 32, 32/nb32 threads now share each chunk and each takes a slice of the rows, reusing the FC_mul_mv_split function constant and the dispatch wrapper introduced for iq3_xxs. The plain path is untouched: wide matrices keep one thread per chunk and N_R0_<TYPE> = 4. Only the split path uses N_R0_<TYPE>_SPLIT = 8. The K-quants have the same idle-thread issue but a different lane mapping, so they are left for a separate change. * metal : offset the src0 row pointer once in the iq mul_mv kernels q2, dh, sc, qh and signs are all derived from xr, so the row slice offset only has to be applied to xr. * metal : fold iq mul_mv row split into offset0 Compute row0 and row1 before initializing the source pointers and apply the row slice directly to offset0. This keeps x and its derived pointers on the existing path while applying the split row offset once.
…< 1024 (ggml-org#28692) * metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024 Generalize the row split from ggml-org#28086 to the six other kernels that use the same lane-to-block mapping: iq1_s, iq1_m, iq2_xxs, iq2_xs, iq2_s and iq3_s. Each of them assigns one 32-element chunk per thread, so when a row has fewer than 32 chunks the rest of the simdgroup is idle. When nb32 < 32 and nb32 divides 32, 32/nb32 threads now share each chunk and each takes a slice of the rows, reusing the FC_mul_mv_split function constant and the dispatch wrapper introduced for iq3_xxs. The plain path is untouched: wide matrices keep one thread per chunk and N_R0_<TYPE> = 4. Only the split path uses N_R0_<TYPE>_SPLIT = 8. The K-quants have the same idle-thread issue but a different lane mapping, so they are left for a separate change. * metal : offset the src0 row pointer once in the iq mul_mv kernels q2, dh, sc, qh and signs are all derived from xr, so the row slice offset only has to be applied to xr. * metal : fold iq mul_mv row split into offset0 Compute row0 and row1 before initializing the source pointers and apply the row slice directly to offset0. This keeps x and its derived pointers on the existing path while applying the split row offset once.
Overview
Follow up to #28086, applying the same
mul_mvrow split fromiq3_xxsto the six other IQ kernels with the same thread mapping:iq1_siq1_miq2_xxsiq2_xsiq2_siq3_sThese kernels assign one 32-element chunk to each simdgroup lane with
ix = tiisg. Whennb32 < 32, some lanes have no chunk to process. Atne00 = 512, 16 of 32 lanes are idle; atne00 = 256, 24 are idle.As in #28086, when
nb32 < 32andnb32divides 32, the idle lanes share the existing chunks and split thenr0output rows between them.The split path uses
N_R0_*_SPLIT = 8. The non-split path keeps the existingN_R0_* = 4and the original thread mapping.K-quants are left out because they use different thread mappings as mentioned before.
Testing
MUL_MAT: 1256/1256 against CPUMUL_MAT_ID: 810/810 against CPUk = 256on the split path andk = 768on the non-split pathPerformance
All measurements were run on an Apple M5 24gb (macOS 26.6.2), on AC power, against master
f3f1a8f27.Kernels
shapes.txtcontainsMUL_MATcases withsrc0 = [k, 4096]andsrc1 = [k, 1], forkin {256, 512, 768, 1024, 4096} across the seven IQ types.The default
perfset does not include these narrow shapes. Its quantizedMUL_MATcases usek >= 2048, so they do not take the split path, I used --test-file shapes.txt to run k = 256, 512, 768, 1024, 4096 across the seven IQ types; the file is included below so the results can be reproduced.shapes.txt
Time per run vs master, negative is faster:
iq1_siq1_miq2_xxsiq2_xsiq2_siq3_sMedian across these 12 shapes is -37.2%. Keeping the original
N_R0 = 4and applying only the thread remapping gives -30.9%, so most of the gain comes from the remapping.iq3_xxs, already fixed by #28086, was measured in the same runs and stayed within 1.5% of master.For the non-split shapes (
k = 768, 1024, 4096, 18 measurements), the median change is +1.0%, with a range of -2.3% to +2.7%. Spread between repeated runs is of similar size, so I cannot distinguish a consistent effect on these shapes. The non-split path keeps the existingN_R0 = 4and original thread mapping.End to end
bartowski/Ornith-1.5-35B-A3B-GGUFhas 512-wide expert tensors, which use the split path.The same command was run with the IQ2_XS and IQ2_S model files.
Token generation, master -> patch -> master (ABA), with the patch compared against the mean of the two master runs:
IQ2_XXSIQ2_XSIQ2_SAll nine token-generation measurements improved, with a median of +3.3%.
Prefill over the same nine cases ranged from -7.8% to +3.7% with no consistent direction. The two master runs differed by as much as 10% on prefill. These batches use
mul_mm/mul_mm_id, which this patch does not change.Batch width
For these expert tensors, Metal switches from
mul_mv_idtomul_mm_idwhenne21 >= 32, wherene21is the number of tokens in the batch.I swept batch widths
B = 1, 2, 4, 8, 16, 24, 32, 48, 64withllama-batched-bench. Every tested width below 32 improved, while results at 32 and above were within run-to-run variation. The same boundary is visible in the master results, where throughput increases from about 90 to 109 tok/s betweenB = 24andB = 32.This also applies to speculative decoding: verify batches below 32 tokens can use the patched
mul_mv_idpath, while batches of 32 or more usemul_mm_id.Requirements