Skip to content

sycl: FWHT optimizations - #29605

Merged
Titaniumtown merged 2 commits into
ggml-org:masterfrom
PrismML-Eng:sycl-fwht-f16
Oct 8, 2026
Merged

Titaniumtown merged 2 commits into
ggml-org:masterfrom
PrismML-Eng:sycl-fwht-f16

Conversation

@bri-prism

Copy link
Copy Markdown
Contributor

Overview

sycl: FWHT optimizations

Additional information

$ build/bin/test-backend-ops test -b SYCL0 -o MUL_MAT_HADAMARD
Backend 1/2: SYCL0
  Device description: Intel(R) Arc(TM) B390 GPU
  32/32 tests passed
  Backend SYCL0: OK
$ master/build/bin/test-backend-ops perf -b SYCL0 -o MUL_MAT_HADAMARD
  MUL_MAT_HADAMARD(type_a=f32,type_b=f16,m=512,n=2048,k=512,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):               4888 runs -   208.44 us/run -   1.07 GFLOP/run -   5.15 TFLOPS
  MUL_MAT_HADAMARD(type_a=f32,type_b=f16,m=1024,n=2048,k=1024,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):                     1512 runs -   663.34 us/run -   4.29 GFLOP/run -   6.47 TFLOPS
  MUL_MAT_HADAMARD(type_a=f32,type_b=f16,m=4096,n=2048,k=4096,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):                      104 runs -  9637.58 us/run -  68.72 GFLOP/run -   7.13 TFLOPS
  MUL_MAT_HADAMARD(type_a=f32,type_b=f16,m=8192,n=2048,k=8192,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):                       27 runs - 38292.48 us/run - 274.88 GFLOP/run -   7.18 TFLOPS
$ pr/build/bin/test-backend-ops perf -b SYCL0 -o MUL_MAT_HADAMARD
  MUL_MAT_HADAMARD(type_a=f32,type_b=f16,m=512,n=2048,k=512,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):              13912 runs -    72.20 us/run -   1.07 GFLOP/run -  14.87 TFLOPS
  MUL_MAT_HADAMARD(type_a=f32,type_b=f16,m=1024,n=2048,k=1024,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):                     8616 runs -   116.31 us/run -   4.29 GFLOP/run -  36.93 TFLOPS
  MUL_MAT_HADAMARD(type_a=f32,type_b=f16,m=4096,n=2048,k=4096,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):                     2094 runs -   477.78 us/run -  68.72 GFLOP/run - 143.83 TFLOPS
  MUL_MAT_HADAMARD(type_a=f32,type_b=f16,m=8192,n=2048,k=8192,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):                     1055 runs -   948.51 us/run - 274.88 GFLOP/run - 289.80 TFLOPS

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. Claude Code was used to help develop and test the fix and to format this description to the PR template. I reviewed every line and take full responsibility for the changes.

Assisted-by: Claude Code
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language labels Sep 28, 2026

@arthw arthw left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On Arc770, there are perf increase on the type_b=fp16 cases. No impact to type_b=fp32.

cmd: ./build/bin/test-backend-ops perf -b SYCL0 -o MUL_MAT_HADAMARD

Config Base PR Change
f32/f32 m=128 n=1 k=128 8.00 GFLOPS 7.96 GFLOPS -0.5%
f32/f32 m=64 n=1 k=64 2.33 GFLOPS 2.32 GFLOPS -0.4%
f32/f32 m=256 n=1 k=256 29.39 GFLOPS 29.37 GFLOPS -0.1%
f32/f32 m=128 n=32 k=128 252.54 GFLOPS 251.12 GFLOPS -0.6%
f32/f32 m=64 n=2048 k=64 2.48 TFLOPS 2.47 TFLOPS -0.4%
f32/f32 m=128 n=2048 k=128 6.73 TFLOPS 6.70 TFLOPS -0.4%
f32/f32 m=256 n=2048 k=256 15.68 TFLOPS 15.64 TFLOPS -0.3%
f32/f32 m=512 n=2048 k=512 35.32 TFLOPS 35.25 TFLOPS -0.2%
f32/f16 m=128 n=1 k=128 3.97 GFLOPS 8.04 GFLOPS +102.5%
f32/f16 m=64 n=1 k=64 1.10 GFLOPS 2.32 GFLOPS +110.9%
f32/f16 m=256 n=1 k=256 15.36 GFLOPS 29.02 GFLOPS +88.9%
f32/f16 m=128 n=32 k=128 107.38 GFLOPS 248.37 GFLOPS +131.3%
f32/f16 m=64 n=2048 k=64 1.13 TFLOPS 2.51 TFLOPS +122.1%
f32/f16 m=128 n=2048 k=128 2.85 TFLOPS 7.15 TFLOPS +150.9%
f32/f16 m=256 n=2048 k=256 6.02 TFLOPS 16.53 TFLOPS +174.6%
f32/f16 m=512 n=2048 k=512 10.08 TFLOPS 41.67 TFLOPS +313.4%

@arthw

arthw commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

@bri-prism
Could you add the perf test case for type_b=fp16 in tests/test-backend-ops.cpp?
So that user can reproduce the test easily.

Thank you!

Assisted-by: Claude Code
@github-actions github-actions Bot added the testing Everything test related label Sep 30, 2026
@bri-prism

Copy link
Copy Markdown
Contributor Author

Thanks for testing on Arc. Added the type_b=fp16 perf cases to test-backend-ops in 4537cdf, same shapes as your table, so the numbers should reproduce with the same command.

@arthw arthw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's good job!

Thank you!

@bri-prism
bri-prism marked this pull request as ready for review October 7, 2026 17:43
@bri-prism
bri-prism requested review from a team and ggerganov as code owners October 7, 2026 17:43
@Titaniumtown
Titaniumtown merged commit 46baf1f into ggml-org:master Oct 8, 2026
21 checks passed
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants