Skip to content

CUDA: skip compilation of superfluous FA kernels - #21768

Merged
IMbackK merged 1 commit into
ggml-org:masterfrom
JohannesGaessler:cuda-faster-compile
Apr 11, 2026
Merged

IMbackK merged 1 commit into
ggml-org:masterfrom
JohannesGaessler:cuda-faster-compile

Conversation

@JohannesGaessler

Copy link
Copy Markdown
Contributor

Fixup to #20998 .

The compilation of FA kernels with head size 512 is supposed to be skipped for GQA ratios of 1 and 2 because those are never used. However, because the invocation of the corresponding template specializations is not guarded with an if constexpr they are being compiled regardless; this PR adds them. On my server with a 64 core EPYC CPU the total compilation time of the full project without CCache goes down from 330s to 300s.

Requirements

@JohannesGaessler
JohannesGaessler requested a review from a team as a code owner April 11, 2026 13:26
@github-actions github-actions Bot added Nvidia GPU Issues specific to Nvidia GPUs ggml changes relating to the ggml tensor library for machine learning labels Apr 11, 2026
@IMbackK
IMbackK merged commit ff5ef82 into ggml-org:master Apr 11, 2026
48 checks passed
oussamaahmia pushed a commit to oussamaahmia/llama-cpp-turboquant-gemma4 that referenced this pull request Apr 13, 2026
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
MrLordCat referenced this pull request in MrLordCat/llama.cpp-rdna-lab Jul 16, 2026
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Nvidia GPU Issues specific to Nvidia GPUs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants