Skip to content

ggml-hip: enable -ffast-math for HIP builds - #23862

Merged
am17an merged 1 commit into
ggml-org:masterfrom
a-huk:hip-fast-math
Jul 6, 2026
Merged

am17an merged 1 commit into
ggml-org:masterfrom
a-huk:hip-fast-math

Conversation

@a-huk

@a-huk a-huk commented May 29, 2026

Copy link
Copy Markdown
Contributor

Overview

This add -ffast-math to HIP builds, mirroring the -use_fast_math flag already applied to CUDA builds in ggml-cuda/CMakeLists.txt.

The flag was previously missing from HIP because it had no effect in older toolchains. It was changed by ROCm/LLVM: https://reviews.llvm.org/D154790

I tried disuccing this in #23339, @IMbackK confirmed the historical reason and @am17an suggested adding the flag rather than using __expf directly.

To prove the changes I ran some benchmarks.

Benchmarks

Benchmarked on gfx1151 (RDNA3.5/Strix Halo, 40 CUs) against latest master (2f6c815). These two models cover both the FA paths.

Qwen3.5-27B Q4_K_M — head_dim=256 (TILE path), FA=1, prompt t/s:

pp upstream + fast-math delta
512 307 321 +4.6%
2048 302 323 +7.0%
4096 300 312 +4.0%
8192 288 307 +6.6%
16384 263 280 +6.5%
32768 233 246 +5.6%

Qwen3-0.6B BF16 — head_dim=64 (MMA path), FA=1, prompt t/s:

pp upstream + fast-math delta
512 9325 9642 +3.4%
2048 8609 8892 +3.3%
4096 7485 7661 +2.4%
8192 5770 5846 +1.3%
16384 3581 3636 +1.5%
32768 2037 2049 +0.6%

FA=0 results were essentially the same across both builds.


Requirements

Mirrors the -use_fast_math flag already applied to CUDA builds
(ggml-cuda/CMakeLists.txt). Previously this was missing from HIP
builds because fast-math had no effect in older HIP/ROCm toolchains,
but this has since been resolved upstream in ROCm/LLVM.

Benchmarked on gfx1151 (RDNA3.5, Strix Halo) against latest master:
+5-10% prompt processing throughput on FA=1 across all context lengths
up to 32K with no measurable quality impact.
@a-huk
a-huk requested a review from IMbackK as a code owner May 29, 2026 09:59
@github-actions github-actions Bot added the ggml changes relating to the ggml tensor library for machine learning label May 29, 2026
@a-huk a-huk closed this May 29, 2026
@am17an

am17an commented May 29, 2026

Copy link
Copy Markdown
Contributor

Looks like you closed the wrong PR.

@am17an am17an reopened this May 29, 2026
@ggml-gh-bot

This comment was marked as resolved.

@a-huk a-huk mentioned this pull request May 29, 2026
@a-huk

a-huk commented May 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Looks like you closed the wrong PR.

Indeed, my bad, I was wondering where it went :/, thanks

@IMbackK IMbackK self-assigned this Jun 1, 2026
@ssakar

ssakar commented Jun 1, 2026

Copy link
Copy Markdown

I think there needs to be an exception -fno-finite-math-only otherwise it can lead to undefined behavior: warning: use of infinity via a macro is undefined behavior due to the currently enabled floating-point options [-Wnan-infinity-disabled]

@am17an
am17an merged commit d06ddd3 into ggml-org:master Jul 6, 2026
2 checks passed
iacopPBK pushed a commit to DENEB1312/mx-llama.cpp that referenced this pull request Jul 7, 2026
@Beinsezii Beinsezii mentioned this pull request Jul 9, 2026
1 task done
iacopPBK pushed a commit to DENEB1312/mx-llama.cpp that referenced this pull request Jul 13, 2026
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants