Skip to content

CUDA: fix FP16 overflow in tile FA kernel - #17875

Merged
JohannesGaessler merged 1 commit into
ggml-org:masterfrom
JohannesGaessler:cuda-fa-tile-fix-overflow
Dec 9, 2025
Merged

JohannesGaessler merged 1 commit into
ggml-org:masterfrom
JohannesGaessler:cuda-fa-tile-fix-overflow

Conversation

@JohannesGaessler

Copy link
Copy Markdown
Contributor

Follow-up to #17558 .

I forgot that the tile kernel would also be suffering the same problem with overflows. Unfortunately unlike the vector kernel the tile kernel is compute bound rather than I/O bound so simply switching the arithmetic to FP32 is not a good option. I think the least bad option is to scale down Q and to then scale up KQ afterwards again. The tradeoff is potentially flushing some small values < $2^{-12}$ to zero. Affected GPUs are P100s, V100s, and RDNA1 or older AMD GPUs.

@github-actions github-actions Bot added Nvidia GPU Issues specific to Nvidia GPUs ggml changes relating to the ggml tensor library for machine learning labels Dec 9, 2025
@JohannesGaessler
JohannesGaessler merged commit 0cdce38 into ggml-org:master Dec 9, 2025
65 of 67 checks passed
Seunghhon pushed a commit to Seunghhon/llama.cpp that referenced this pull request Apr 26, 2026
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
MrLordCat referenced this pull request in MrLordCat/llama.cpp-rdna-lab Jul 16, 2026
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Nvidia GPU Issues specific to Nvidia GPUs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants