Skip to content

CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 - #26289

Merged
ggerganov merged 2 commits into
ggml-org:masterfrom
animeshsri14:perf/pascal-fattn-tile-fp16-tune
Sep 27, 2026
Merged

ggerganov merged 2 commits into
ggml-org:masterfrom
animeshsri14:perf/pascal-fattn-tile-fp16-tune

Conversation

@animeshsri14

Copy link
Copy Markdown
Contributor

Overview

This PR optimizes the FlashAttention FP16 tile configurations for NVIDIA GPUs, addressing the header TODO in ggml/src/ggml-cuda/fattn-tile.cuh to "optimize kernel parameters for FP16 NVIDIA (P100)" for head sizes 40, 64, 72, 80, 96, and 112.

By retuning exactly 13 rows in ggml_cuda_fattn_tile_get_config_nvidia_fp16, this PR achieves a mean kernel-level speedup of +14.5% on P100 (up to +31.5%), with zero regressions across Pascal (P100), Volta (V100), and Blackwell architectures.

End-to-end, this yields a solid +2% decode speedup for phi-2 at short contexts, scaling to +9.4% at 8K context, while remaining neutral for other models and prefill workloads.

Additional information

1. Hardware & Methodology

GPU cc driver toolchain
Tesla P100-PCIE/SXM 16GB 6.0 580.x CUDA 12.9, g++-14
Tesla V100-SXM2-16GB 7.0 580.159.03 CUDA 12.9, g++-14
RTX Pro 6000 Blackwell 12.0 - CUDA 12.9

Methodology:

  • Tests were performed via paired alternating A/B runs.
  • Execution used LD_LIBRARY_PATH (because RUNPATH would otherwise override).
  • GPU persistence mode was ON and clocks were locked.

2. Microbenchmark Gains (Isolated Kernel)

Rows were timed via test-backend-ops perf -o FLASH_ATTN_EXT (4 heads, KV 4096, F16 K/V, F32 accum). ncols maps to the batch size / GQA ratio.

Noise Floor: In-run noise floors were established by the 13 unmodified guard rows at 2.6% for P100 and 1.8% for V100. We define "neutral" as any delta falling within this noise floor. Under this definition, every modified row is win-or-neutral across all architectures it runs on.

hs ncols default tuned P100 (cc6.0) V100 (cc7.0) Blackwell (cc12.0)
40 2 (64,2,64,40) (128,3,128,40) +26.6% +14.0% +7.0%
40 4 (128,2,64,40) (128,2,128,40) +21.3% +20.9% +15.2%
40 8 (256,2,64,40) (128,3,128,40) +9.4% +24.4% +5.5%
64 8 (256,2,64,64) (256,3,128,64) +6.0% +9.0% n/a*
72 2 (64,2,64,72) (128,2,64,72) +4.6% +2.7% +13.3%
72 4 (128,2,64,72) (128,3,128,72) +24.3% +14.8% +21.0%
80 2 (64,2,64,40) (128,2,64,40) +4.9% -0.7% n/a*
80 32 (256,2,64,40) (256,2,128,40) +10.9% +0.0% n/a*
96 32 (256,2,64,48) (256,2,128,24) +6.8% +0.0% n/a*
112 2 (64,2,64,56) (128,3,64,56) +31.5% +1.5% n/a*
112 8 (256,2,64,56) (256,3,64,112) +18.5% +53.9% n/a*
112 16 (256,2,64,56) (256,2,32,112) +17.8% +41.0% n/a*
112 32 (256,2,64,56) (256,2,128,56) +5.9% +0.0% n/a*

(Note: Every tensor-core branch is explicitly guarded by ne[0] != 40 && ne[0] != 72, meaning head sizes 40 and 72 fall through to the TILE kernel on all fast-fp16 NVIDIA GPUs through Blackwell.)
(*row does not execute on the arch)

3. Tuning Mechanisms

  • nbatch_fa 64 -> 128: Amortizes softmax/rescale overhead by halving the KQ-tile loop iteration count. Main driver for small-ncols rows.
  • occupancy 2 -> 3: Hides more latency with a third resident block per SM, taking advantage of the shared-memory headroom left by small heads.
  • nthreads 64 -> 128: Doubles per-block parallelism at small batch sizes (ncols=2), fixing severe SM underutilization.
  • nbatch_K full-DV: Halves K-loop overhead for several rows by loading the full DV per iteration instead of DV/2.

4. End-to-End Scaling & Cross-Model Safety (P100)

Tested via llama-bench.

Cross-model Safety:
Models hitting untuned batch-1 rows showed flat performance within the run-to-run noise margin for both prefill and decode, indicating zero regressions in production scenarios.

model params head size decode Δ prefill Δ
TinyLlama-1.1B-Chat 1.1B 64 -1.15% (noise) +0.22%
phi-2 2.7B 80 +2.13% -0.14%
Phi-3-mini-4k 3.8B 96 -0.02% -0.02%
Mistral-7B-Instruct 7B 128 -0.02% +0.24%

Phi-2 (hs 80) Decode Scaling:
Single-batch decode hits the tuned 80/2 row. The speedup scales effectively with context length as FlashAttention consumes a larger share of the decode step relative to weight-streaming:

KV depth base t/s tuned t/s decode Δ sd
0 82.80 83.73 +1.10% 2.10
1024 78.45 80.36 +2.44% 0.75
2048 74.38 74.70 +0.42% 2.48
4096 67.50 70.14 +3.91% 0.07
8192 54.91 60.08 +9.41% 0.14

Caveat: Depth > 2048 exceeds phi-2's native context; these are throughput measurements, not a generation-quality claim.

5. Correctness

Passed test-backend-ops test -o FLASH_ATTN_EXT:

  • P100 (cc6.0): 2880/2880 OK, 0 FAIL.
  • V100 (cc7.0): 2656/2656 OK, byte-identical against the untuned baseline for affected head sizes. (Note: The full V100 suite throws a cudaFuncSetAttribute invalid argument abort on the hs 320/256 MLA test at fattn-mma-f16.cuh:1945 because it exceeds Volta's 96 KB smem limit. Verified on the baseline(untuned branch) to be identical error, unrelated to this pr.
  • Blackwell (cc12.0): 2880/2880 OK.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: yes, codex used for automated benchmarking (verified after each run by me, tuned, iterated and methodology done by me)

@animeshsri14
animeshsri14 requested a review from a team as a code owner July 29, 2026 18:26
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Jul 29, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Jul 29, 2026

Copy link
Copy Markdown

Hi @animeshsri14, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@animeshsri14

Copy link
Copy Markdown
Contributor Author

Previous pr has been converted to draft for now, as I would like to allot my one-pr slot to this pr.

@JohannesGaessler JohannesGaessler added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 26, 2026
@ggerganov
ggerganov merged commit 36d7b08 into ggml-org:master Sep 27, 2026
2 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 1, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 2, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants