Skip to content

sycl: fuse RMS_NORM + MUL - #26015

Merged
ggerganov merged 1 commit into
ggml-org:masterfrom
Titaniumtown:pr/sycl-rms-norm-mul-fusion
Jul 31, 2026
Merged

ggerganov merged 1 commit into
ggml-org:masterfrom
Titaniumtown:pr/sycl-rms-norm-mul-fusion

Conversation

@Titaniumtown

Copy link
Copy Markdown
Contributor

Overview

Port of cuda path ggml_cuda_op_rms_norm_fused to SYCL.

Also adds a function ggml_sycl_can_fuse that can be used in future fusion PRs that I am am working on. This is analogous to the cuda ggml_cuda_can_fuse.

Additional information

Benchmarked on an Arc B70:

master @ 0278d83:

| model                          |       size |     params | backend    | ngl |  fa |            test |                  t/s |
|------------------------------ | ---------: | ---------: | ---------- | --: | --: | --------------: | -------------------: |
| qwen35 27B Q4_K - Medium       |  16.67 GiB |    27.32 B | SYCL       | 999 |   1 |           pp512 |        849.40 ± 1.34 |
| qwen35 27B Q4_K - Medium       |  16.67 GiB |    27.32 B | SYCL       | 999 |   1 |          pp2048 |        845.89 ± 1.58 |
| qwen35 27B Q4_K - Medium       |  16.67 GiB |    27.32 B | SYCL       | 999 |   1 |           tg128 |         23.37 ± 0.02 |

Patched — master + rms_norm+mul fusion:

| model                          |       size |     params | backend    | ngl |  fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | --------------: | -------------------: |
| qwen35 27B Q4_K - Medium       |  16.67 GiB |    27.32 B | SYCL       | 999 |   1 |           pp512 |        854.02 ± 1.54 |
| qwen35 27B Q4_K - Medium       |  16.67 GiB |    27.32 B | SYCL       | 999 |   1 |          pp2048 |        851.84 ± 0.72 |
| qwen35 27B Q4_K - Medium       |  16.67 GiB |    27.32 B | SYCL       | 999 |   1 |           tg128 |         23.48 ± 0.02 |

Uplift: pp512 +0.54%, pp2048 +0.70%, tg128 +0.47%.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - Claude Opus 4.8 was used to help understand the codebase and testing.

@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language labels Jul 22, 2026

@arthw arthw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With Qwen3.5-4B-Q4_K_M.gguf on B60:

Test fa Base t/s Primary t/s Increase Rate (Primary vs Base)
pp512 0 1048.44 1056.17 0.74%
pp512 1 1048.79 1055.80 0.67%
tg128 0 80.09 81.02 1.16%
tg128 1 80.60 81.65 1.30%

Comment thread ggml/src/ggml-sycl/ggml-sycl.cpp Outdated
@Titaniumtown
Titaniumtown force-pushed the pr/sycl-rms-norm-mul-fusion branch from f585fd2 to 75450ee Compare July 28, 2026 22:09
@Titaniumtown
Titaniumtown marked this pull request as ready for review July 28, 2026 22:10
@Titaniumtown
Titaniumtown requested a review from a team as a code owner July 28, 2026 22:10
@Titaniumtown

Copy link
Copy Markdown
Contributor Author

Ready for review!

Comment thread ggml/src/ggml-sycl/norm.cpp

@arthw arthw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's good job!

Thank you!

@arthw arthw added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Jul 30, 2026
@ggerganov
ggerganov merged commit 1553725 into ggml-org:master Jul 31, 2026
23 of 29 checks passed
@malsbat

malsbat commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

@Titaniumtown, I was wondering what other fusions you have in the works? I was looking at bringing some work we had done earlier up to date and found this PR that already did the RMS_NORM+MUL.

I don't want to duplicate effort if you're already working on these ones:

@Titaniumtown

Copy link
Copy Markdown
Contributor Author

@malsbat I have both of those already locally. I can submit PRs for those.

huaxel pushed a commit to huaxel/CachyLLama that referenced this pull request Aug 2, 2026
@malsbat

malsbat commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

@malsbat I have both of those already locally. I can submit PRs for those.

Great, looking forward to those PRs!

satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants