Skip to content

cuda: add Kronecker FWHT support - #29051

Open
Lonny154 wants to merge 1 commit into
ggml-org:masterfrom
Lonny154:cuda-kronecker-fwht
Open

Lonny154 wants to merge 1 commit into
ggml-org:masterfrom
Lonny154:cuda-kronecker-fwht

Conversation

@Lonny154

@Lonny154 Lonny154 commented Sep 17, 2026 •

Copy link
Copy Markdown

Overview

I added CUDA support for MUL_MAT_HADAMARD / FWHT dimensions (384, 640, 768, and 1280). They are not powers of two.

This was done using Kronecker products with the fixed H12/H20 transforms.

The existing power-of-two CUDA path is untouched.

Additional information

16/16 CUDA tests passed
GPU used RTX 4060 Ti (Compute capability 8.9, CUDA 13.3):

./build/bin/test-backend-ops test -b CUDA0 -o MUL_MAT_HADAMARD

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES — I used ChatGPT to help with learning the existing FWHT code and Kronecker algorithm, and ChatGPT helped to generate the CUDA code. I reviewed the code and tested the resulting code.

@Lonny154
Lonny154 requested review from a team and ggerganov as code owners September 17, 2026 21:14
@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Sep 17, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 17, 2026

Copy link
Copy Markdown

Hi @Lonny154, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

  • AI-generated content: While code is allowed to be generated by AI, please write the PR description and commit messages on your own without the help of AI.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@ggml-gh-bot ggml-gh-bot Bot added the draft PR will be changed to draft by github-actions bot label Sep 17, 2026
@github-actions
github-actions Bot marked this pull request as draft September 17, 2026 21:20
@github-actions github-actions Bot removed the draft PR will be changed to draft by github-actions bot label Sep 17, 2026
@Lonny154
Lonny154 marked this pull request as ready for review September 17, 2026 22:38
@ggml-org ggml-org deleted a comment from ggml-gh-bot Bot Sep 18, 2026
@pwilkin

pwilkin commented Sep 18, 2026

Copy link
Copy Markdown
Member

/bot review

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

❌ Code review failed.

Error: failed to fetch diff (503) from https://github.com/ggml-org/llama.cpp/pull/29051.diff

@am17an

am17an commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

where is this operation required?

@Lonny154

Copy link
Copy Markdown
Author

@am17an I did some further research and I have not found a specific runtime path where these dimensions are required. When tracing the current KV-cache rotation path, the rotation size gets reduced to a power of two dimension.

I originally saw the SYCL implementation adding support for non power of two dimensions and saw an opportunity to do something similar for CUDA support.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants