Repository navigation
CUDA: tune FA for Gemma 4 on Ampere or newer - #29152
Merged
Merged
Conversation
am17an
approved these changes
Sep 20, 2026
Contributor
|
Deepseek also has a 512 head size. |
Contributor
Author
|
Deepseek has |
Contributor
|
@JohannesGaessler for the deepseek 4 family |
ynankani
approved these changes
Sep 20, 2026
Contributor
There was a problem hiding this comment.
Performance DGX-SPARK
| GPU | Model | Microbatch size | Test | t/s master | t/s a3cec6a | Speedup |
|---|---|---|---|---|---|---|
| DGX-SPARK | gemma4 26B.A4B Q4_0 | 1 | pp256 | 104.74 | 104.39 | 1 |
| DGX-SPARK | gemma4 26B.A4B Q4_0 | 1 | pp256@d32768 | 75.67 | 75.37 | 1 |
| DGX-SPARK | gemma4 26B.A4B Q4_0 | 2 | pp256 | 158.76 | 158.82 | 1 |
| DGX-SPARK | gemma4 26B.A4B Q4_0 | 2 | pp256@d32768 | 122.27 | 122.5 | 1 |
| DGX-SPARK | gemma4 26B.A4B Q4_0 | 4 | pp256 | 230.22 | 230.03 | 1 |
| DGX-SPARK | gemma4 26B.A4B Q4_0 | 4 | pp256@d32768 | 190.11 | 190.09 | 1 |
| DGX-SPARK | gemma4 31B Q4_0 | 1 | pp256 | 13.2 | 13.21 | 1 |
| DGX-SPARK | gemma4 31B Q4_0 | 1 | pp256@d32768 | 11.19 | 11.12 | 0.99 |
| DGX-SPARK | gemma4 31B Q4_0 | 2 | pp256 | 26.19 | 26.08 | 1 |
| DGX-SPARK | gemma4 31B Q4_0 | 2 | pp256@d32768 | 22.12 | 22.05 | 1 |
| DGX-SPARK | gemma4 31B Q4_0 | 4 | pp256 | 51.59 | 51.2 | 0.99 |
| DGX-SPARK | gemma4 31B Q4_0 | 4 | pp256@d32768 | 43.49 | 43.34 | 1 |
| DGX-SPARK | gemma4 12B Q4_0 | 1 | pp256 | 36.79 | 36.66 | 1 |
| DGX-SPARK | gemma4 12B Q4_0 | 1 | pp256@d32768 | 32.42 | 32.36 | 1 |
| DGX-SPARK | gemma4 12B Q4_0 | 2 | pp256 | 70.21 | 70.08 | 1 |
| DGX-SPARK | gemma4 12B Q4_0 | 2 | pp256@d32768 | 61.87 | 62.01 | 1 |
| DGX-SPARK | gemma4 12B Q4_0 | 4 | pp256 | 139.79 | 139.44 | 1 |
| DGX-SPARK | gemma4 12B Q4_0 | 4 | pp256@d32768 | 122.65 | 122.75 | 1 |
| DGX-SPARK | gemma4 E2B Q4_0 | 1 | pp256 | 162.3 | 162.41 | 1 |
| DGX-SPARK | gemma4 E2B Q4_0 | 1 | pp256@d32768 | 120.35 | 120.35 | 1 |
| DGX-SPARK | gemma4 E2B Q4_0 | 2 | pp256 | 315.53 | 319.75 | 1.01 |
| DGX-SPARK | gemma4 E2B Q4_0 | 2 | pp256@d32768 | 233.18 | 238.19 | 1.02 |
| DGX-SPARK | gemma4 E2B Q4_0 | 4 | pp256 | 616.71 | 615.92 | 1 |
| DGX-SPARK | gemma4 E2B Q4_0 | 4 | pp256@d32768 | 450.88 | 456.45 | 1.01 |
IMbackK
approved these changes
Sep 20, 2026
Contributor
|
@JohannesGaessler please rebase |
JohannesGaessler
force-pushed
the
cuda-fa-tune-2
branch
from
September 20, 2026 18:51
a3cec6a to
9b916dc
Compare
Contributor
|
your rebase erroneously removed the |
JohannesGaessler
force-pushed
the
cuda-fa-tune-2
branch
from
September 20, 2026 19:01
9b916dc to
2216cae
Compare
IMbackK
approved these changes
Sep 20, 2026
hariag
added a commit
to hariag/llama.cpp
that referenced
this pull request
Sep 23, 2026
Merges the danielhanchen qwen38-mtp branch (004b547) and the three gemma4-related commits on top of it into upstream master (e6ab7c1): - 6f29477 CUDA: tune FA for Gemma 4 on Ampere or newer (backport ggml-org#29152) - 6cee122 model: fix gemma4-assistant SWA pattern array length (backport ggml-org#28183) - 372d950 ggml: port INT8 ConvRot support (stable-diffusion vendor build) Conflicts resolved in favor of upstream where it already contains an equivalent fix (gemma4-assistant via ggml-org#28868, fattn.cu >=256 condition), keeping convrot additions merged alongside upstream Hadamard logic. # Conflicts: # ggml/src/ggml-cpu/ggml-cpu.cpp # ggml/src/ggml-cuda/fattn.cu # ggml/src/ggml-vulkan/ggml-vulkan.cpp # src/models/gemma4-assistant.cpp # src/models/qwen4exp.cpp
1 task done
turbo-tan
pushed a commit
to turbo-tan/llama.cpp-tq3
that referenced
this pull request
Sep 23, 2026
ggml-org#28102 replay dropped fork's vec_dot_fattn_vec_KQ_{tq3_0,turbo3_0,turbo4_0}, dequantize_V_{tq3_0,turbo3_0,turbo4_0}, turbo4_decode_element + getter arms (static_assert 'bad type' on GB10 builds). Re-landed from main onto the post-ggml-org#28102/ggml-org#29152 upstream shape (185 lines, no other deltas).
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 1, 2026
# Conflicts: # ggml/src/ggml-cuda/fattn.cu
LadislavSopko
pushed a commit
to 0ics-srls/llama.cpp
that referenced
this pull request
Oct 5, 2026
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
This PR tunes the CUDA FlashAttention code for Ampere or newer for head sizes 256 and 512 (relevant for Gemma 4) and batch sizes 1-4 to favor larger CUDA blocks and to preferably use the mma kernel for batch size 1. This yields a bit of performance, primarily for the small models.
Performance
Table generated with
scripts/compare-llama-bench.pyThe RTX 3090 has cooling issues with the thermal pads which interfered with the measurements for 31b.
Additional information
Requirements