Repository navigation
Conversation
|
Hi @Lonny154, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
76dcdf3 to
791a265
Compare
|
/bot review |
|
❌ Code review failed. |
|
where is this operation required? |
|
@am17an I did some further research and I have not found a specific runtime path where these dimensions are required. When tracing the current KV-cache rotation path, the rotation size gets reduced to a power of two dimension. I originally saw the SYCL implementation adding support for non power of two dimensions and saw an opportunity to do something similar for CUDA support. |
Overview
I added CUDA support for
MUL_MAT_HADAMARD/ FWHT dimensions (384, 640, 768, and 1280). They are not powers of two.This was done using Kronecker products with the fixed
H12/H20transforms.The existing power-of-two CUDA path is untouched.
Additional information
16/16CUDA tests passedGPU used RTX 4060 Ti (Compute capability 8.9, CUDA 13.3):
./build/bin/test-backend-ops test -b CUDA0 -o MUL_MAT_HADAMARDRequirements