sycl: FWHT kernels for block widths above 512 - #29243
Conversation
arthw
left a comment
There was a problem hiding this comment.
It's good job!
I test the related UT cases and all are passed on Arc770.
Thank you!
The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus 384/640/768/1280 via the Kronecker/Paley construction added separately in Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through to the default case and run as a dense GEMM against the materialized rotation tensor, correct but O(n^2) instead of O(n log n). fwht_kernel_wide runs one row per work-group instead of per sub-group, so each work-item keeps N/NT values rather than N/WARP_SIZE. Butterflies below the sub-group width still shuffle; those up to the work-group width go through work-group local memory; the rest stay in registers. Same butterfly and sign convention as the existing narrow kernel. ggml's SYCL backend registration (dpct::dev_mgr) unconditionally requires a GPU-labeled platform to exist and throws before any op-level test can run, so test-backend-ops could not be exercised on this box (a GPU-less pod) even via the CPU device. Verified instead with a standalone harness: the same kernel body run through a real SYCL CPU device (Intel oneAPI DPC++ 2026.1, OpenCL CPU backend), checked against an independent recursive-doubling Hadamard reference, cross-validated by first running the existing unmodified narrow kernel through the identical harness and confirming it passes (rules out a reference-convention bug before trusting a pass on the new code). Random-input results for all four widths, single- and multi-row: N=1024 NT=256 rows=1 max_abs_err=1.7e-07 max_rel_err=4.9e-04 PASS N=2048 NT=256 rows=1 max_abs_err=1.9e-07 max_rel_err=2.0e-04 PASS N=4096 NT=256 rows=1 max_abs_err=2.0e-07 max_rel_err=1.4e-04 PASS N=8192 NT=256 rows=1 max_abs_err=2.5e-07 max_rel_err=3.8e-03 PASS N=1024 NT=256 rows=7 max_abs_err=2.4e-07 max_rel_err=1.0e-03 PASS N=2048 NT=256 rows=5 max_abs_err=3.0e-07 max_rel_err=9.4e-04 PASS N=4096 NT=256 rows=3 max_abs_err=2.7e-07 max_rel_err=1.7e-03 PASS N=8192 NT=256 rows=2 max_abs_err=2.5e-07 max_rel_err=1.9e-03 PASS This covers the kernel algorithm itself; it does not exercise the ggml dispatch/supports_op integration end to end, which needs a real GPU (or a SYCL GPU plugin) to get past backend registration. test-backend-ops build is verified: fwht.cpp recompiles with zero warnings as part of ggml-sycl.
e041aaf to
5427536
Compare
|
Rebased onto current master now that #29095 is in, and marked ready. The test cases this PR added for 1024 to 8192 are already on master from #29095, so the diff is back to the single SYCL file, unchanged from the version @arthw approved. The F16-input cases that came with them take the regular matmul path on SYCL, since the FWHT only handles F32. |
|
Tested on an Intel Arc B390 (Xe3 iGPU, Windows, oneAPI 2025.3, Level Zero) at 5427536: |
The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus 384/640/768/1280 via the Kronecker/Paley construction added separately in Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through to the default case and run as a dense GEMM against the materialized rotation tensor, correct but O(n^2) instead of O(n log n). fwht_kernel_wide runs one row per work-group instead of per sub-group, so each work-item keeps N/NT values rather than N/WARP_SIZE. Butterflies below the sub-group width still shuffle; those up to the work-group width go through work-group local memory; the rest stay in registers. Same butterfly and sign convention as the existing narrow kernel. ggml's SYCL backend registration (dpct::dev_mgr) unconditionally requires a GPU-labeled platform to exist and throws before any op-level test can run, so test-backend-ops could not be exercised on this box (a GPU-less pod) even via the CPU device. Verified instead with a standalone harness: the same kernel body run through a real SYCL CPU device (Intel oneAPI DPC++ 2026.1, OpenCL CPU backend), checked against an independent recursive-doubling Hadamard reference, cross-validated by first running the existing unmodified narrow kernel through the identical harness and confirming it passes (rules out a reference-convention bug before trusting a pass on the new code). Random-input results for all four widths, single- and multi-row: N=1024 NT=256 rows=1 max_abs_err=1.7e-07 max_rel_err=4.9e-04 PASS N=2048 NT=256 rows=1 max_abs_err=1.9e-07 max_rel_err=2.0e-04 PASS N=4096 NT=256 rows=1 max_abs_err=2.0e-07 max_rel_err=1.4e-04 PASS N=8192 NT=256 rows=1 max_abs_err=2.5e-07 max_rel_err=3.8e-03 PASS N=1024 NT=256 rows=7 max_abs_err=2.4e-07 max_rel_err=1.0e-03 PASS N=2048 NT=256 rows=5 max_abs_err=3.0e-07 max_rel_err=9.4e-04 PASS N=4096 NT=256 rows=3 max_abs_err=2.7e-07 max_rel_err=1.7e-03 PASS N=8192 NT=256 rows=2 max_abs_err=2.5e-07 max_rel_err=1.9e-03 PASS This covers the kernel algorithm itself; it does not exercise the ggml dispatch/supports_op integration end to end, which needs a real GPU (or a SYCL GPU plugin) to get past backend registration. test-backend-ops build is verified: fwht.cpp recompiles with zero warnings as part of ggml-sycl.
…gml-org#28254, ggml-org#29243) (#302) * sycl : fix test-backend-ops CI break && restore Kronecker product FWHT support (ggml-org#28016) (ggml-org#28254) * Reapply "sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 12…" (ggml-org#28184) This reverts commit c845263. * tests : fix unused variable M in test-backend-ops * tests: fix trailing space error and isolate kronecker tests for sycl backend only (cherry picked from commit 4d91760) * sycl: FWHT kernels for block widths above 512 (ggml-org#29243) The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus 384/640/768/1280 via the Kronecker/Paley construction added separately in Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through to the default case and run as a dense GEMM against the materialized rotation tensor, correct but O(n^2) instead of O(n log n). fwht_kernel_wide runs one row per work-group instead of per sub-group, so each work-item keeps N/NT values rather than N/WARP_SIZE. Butterflies below the sub-group width still shuffle; those up to the work-group width go through work-group local memory; the rest stay in registers. Same butterfly and sign convention as the existing narrow kernel. ggml's SYCL backend registration (dpct::dev_mgr) unconditionally requires a GPU-labeled platform to exist and throws before any op-level test can run, so test-backend-ops could not be exercised on this box (a GPU-less pod) even via the CPU device. Verified instead with a standalone harness: the same kernel body run through a real SYCL CPU device (Intel oneAPI DPC++ 2026.1, OpenCL CPU backend), checked against an independent recursive-doubling Hadamard reference, cross-validated by first running the existing unmodified narrow kernel through the identical harness and confirming it passes (rules out a reference-convention bug before trusting a pass on the new code). Random-input results for all four widths, single- and multi-row: N=1024 NT=256 rows=1 max_abs_err=1.7e-07 max_rel_err=4.9e-04 PASS N=2048 NT=256 rows=1 max_abs_err=1.9e-07 max_rel_err=2.0e-04 PASS N=4096 NT=256 rows=1 max_abs_err=2.0e-07 max_rel_err=1.4e-04 PASS N=8192 NT=256 rows=1 max_abs_err=2.5e-07 max_rel_err=3.8e-03 PASS N=1024 NT=256 rows=7 max_abs_err=2.4e-07 max_rel_err=1.0e-03 PASS N=2048 NT=256 rows=5 max_abs_err=3.0e-07 max_rel_err=9.4e-04 PASS N=4096 NT=256 rows=3 max_abs_err=2.7e-07 max_rel_err=1.7e-03 PASS N=8192 NT=256 rows=2 max_abs_err=2.5e-07 max_rel_err=1.9e-03 PASS This covers the kernel algorithm itself; it does not exercise the ggml dispatch/supports_op integration end to end, which needs a real GPU (or a SYCL GPU plugin) to get past backend registration. test-backend-ops build is verified: fwht.cpp recompiles with zero warnings as part of ggml-sycl. (cherry picked from commit c829670) --------- Co-authored-by: Jingxin (Philip) Li <philipaslee@gmail.com>
Overview
The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus 384/640/768/1280 via the Kronecker/Paley construction added separately in #28016. Neither path covers the pure power-of-2 widths above 512 that the Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through to the default case and run as a dense GEMM against the materialized rotation tensor, correct but O(n^2) instead of O(n log n).
fwht_kernel_wide runs one row per work-group instead of per sub-group, so each work-item keeps N/NT values rather than N/WARP_SIZE. Butterflies below the sub-group width still shuffle; those up to the work-group width go through work-group local memory; the rest stay in registers. Same butterfly and sign convention as the existing narrow kernel.
Additional information
Verification:
test-backend-ops test -b SYCL0 -o MUL_MAT_HADAMARDpasses 26/26 on an Intel Arc B390 (Xe3 iGPU, Windows, oneAPI 2025.3, Level Zero) at the current head, covering every width from 1024 to 16384 with both F32 and F16 src1. Before that I checked the kernel algorithm with a standalone harness on a SYCL CPU device against an independent Hadamard reference; the numbers are in the commit message. The test cases for these widths are already on master from #29095, so this PR only touches the SYCL FWHT.Requirements
structure of the existing narrow sub-group kernel, and to format this description to the PR template.
I reviewed every line and take full responsibility for the changes.