Skip to content

metal : optimize sparse FA + clean-up - #29377

Merged
ggerganov merged 4 commits into
masterfrom
gg/metal-fa-sparse-opt
Sep 24, 2026
Merged

ggerganov merged 4 commits into
masterfrom
gg/metal-fa-sparse-opt

Conversation

@ggerganov

Copy link
Copy Markdown
Member

Overview

Optimize the FA sparse path by moving the indices to shared mem.

Additional information

make -j && ./bin/llama-batched-bench -hf ggml-org/DeepSeek-V4-Flash-Vision-Exp-GGUF:Q2_K -c 70000 -b 2048 -ub 2048 -npp 2048,4096,8192,16384,32768,65536 -ntg 32 -npl 1 -lv 4

# master
|    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
|-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
|  2048 |     32 |    1 |   2080 |    4.886 |   419.19 |    1.212 |    26.40 |    6.098 |   341.11 |
|  4096 |     32 |    1 |   4128 |   10.214 |   401.03 |    1.187 |    26.95 |   11.401 |   362.07 |
|  8192 |     32 |    1 |   8224 |   21.065 |   388.88 |    1.196 |    26.76 |   22.261 |   369.43 |
| 16384 |     32 |    1 |  16416 |   43.488 |   376.74 |    1.210 |    26.45 |   44.698 |   367.26 |
| 32768 |     32 |    1 |  32800 |   91.006 |   360.07 |    1.253 |    25.54 |   92.259 |   355.52 |
| 65536 |     32 |    1 |  65568 |  215.277 |   304.43 |    1.349 |    23.72 |  216.626 |   302.68 |

# PR
|    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
|-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
|  2048 |     32 |    1 |   2080 |    4.759 |   430.32 |    1.173 |    27.28 |    5.932 |   350.62 |
|  4096 |     32 |    1 |   4128 |    9.805 |   417.76 |    1.186 |    26.98 |   10.991 |   375.59 |
|  8192 |     32 |    1 |   8224 |   20.070 |   408.18 |    1.195 |    26.79 |   21.264 |   386.75 |
| 16384 |     32 |    1 |  16416 |   41.335 |   396.37 |    1.213 |    26.38 |   42.548 |   385.82 |
| 32768 |     32 |    1 |  32800 |   86.611 |   378.33 |    1.250 |    25.60 |   87.861 |   373.32 |
| 65536 |     32 |    1 |  65568 |  188.258 |   348.12 |    1.339 |    23.90 |  189.597 |   345.83 |

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

@ggerganov
ggerganov requested a review from a team as a code owner September 24, 2026 13:43
@github-actions github-actions Bot added documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) labels Sep 24, 2026
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
@ggerganov
ggerganov force-pushed the gg/metal-fa-sparse-opt branch from faf2fb8 to 4c9f0cb Compare September 24, 2026 18:14
@ggerganov
ggerganov merged commit cdc0642 into master Sep 24, 2026
22 checks passed
@nikwen

nikwen commented Sep 26, 2026

Copy link
Copy Markdown
Member

It brings me so much joy to see all these performance improvements. Thanks, Georgi!

@ggerganov
ggerganov deleted the gg/metal-fa-sparse-opt branch September 27, 2026 18:14
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* metal : cache sparse FA indices in shared memory

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : simplify shared memory size calculation

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* pi : update general

* metal : unroll sparse index load

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
* metal : cache sparse FA indices in shared memory

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : simplify shared memory size calculation

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* pi : update general

* metal : unroll sparse index load

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
(cherry picked from commit cdc0642)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants