Skip to content

cuda: FWHT kernels for block widths above 512 - #29100

Open
bri-prism wants to merge 1 commit into
ggml-org:masterfrom
PrismML-Eng:up/fwht-cuda-wide
Open

bri-prism wants to merge 1 commit into
ggml-org:masterfrom
PrismML-Eng:up/fwht-cuda-wide

Conversation

@bri-prism

@bri-prism bri-prism commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Overview

The CUDA FWHT covers block widths 64 to 512. It runs one row per warp and keeps N/32
values per lane, so wider blocks need more registers per lane than that layout allows.

fwht_cuda_block runs one row per thread block with 256 threads, so each thread keeps
N/256 values instead. Stages below the warp width still shuffle, those up to the block
width go through shared memory, and the rest stay in registers. The butterfly and sign
convention match the existing kernel.

Widths 64 to 512 keep the warp kernel and are untouched. 1024 through 8192 use the new
one, for both F32 and F16 sources.

This is the CUDA counterpart to #29095 (Metal). Stacked on #29096; the last commit is the
one to review here.

Additional information

Tested on an H100 PCIe (compute capability 9.0, CUDA 12.8, -DCMAKE_CUDA_ARCHITECTURES=90):

  • test-backend-ops -o MUL_MAT_HADAMARD 25/25, including the new 1024/2048/4096/8192 cases
    in F32 and F16, and the 16384 case that falls back
  • test-backend-ops -o MUL_MAT 1297/1297, no regressions

A passing suite does not by itself show the new kernel ran, since a Hadamard-hinted matmul that
falls back to the generic path computes the same values. To confirm the wide kernel is actually
selected I temporarily made fwht_cuda_block scale its output by 2 and re-ran, which gave
15/25 with exactly ten failures:

  • the ten wide cases, 1024/2048/4096/8192 and 1024 with seven rows, in both F32 and F16, all failed
  • all fourteen cases at 512 and below passed, which is the warp kernel that change does not touch
  • 16384 also passed, confirming widths outside the supported list still take the generic path

The change was reverted and the binary rebuilt before the numbers above were taken.

8192 allocates 32 KB of shared memory per block, within the 48 KB default limit. If that is too
aggressive for older parts I am happy to cap the width list lower.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. Claude Code was used to write and test this change, following the
    structure of the existing warp kernel and the Metal version, and to format this description to
    the PR template. I reviewed every line and take full responsibility for the changes.

@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Sep 18, 2026
@bri-prism
bri-prism force-pushed the up/fwht-cuda-wide branch 2 times, most recently from 249bd43 to 18187db Compare September 21, 2026 19:16
The CUDA FWHT covers widths 64 to 512. It runs one row per warp and keeps N/32
values per lane, so wider blocks need more registers per lane than that layout
allows.

fwht_cuda_block runs one row per thread block with 256 threads, so each thread
keeps N/256 values. Stages below the warp width still shuffle, those up to the
block width go through shared memory, and the rest stay in registers. Same
butterfly and sign convention as the warp kernel.

Widths 64 to 512 keep the warp kernel. 1024 through 8192 use the new one, for
both F32 and F16 sources. ggml_cuda_op_mul_mat_use_fwht (the shared
supports_op/dispatch predicate added in ggml-org#29096) does not check width, so it
needed no change here: any width it admits that ggml_cuda_op_fwht can't serve
already falls through correctly to the cuBLAS path.

Rebased onto current master with ggml-org#29096's F16 commit underneath it, since this
depends on the same F16 template infrastructure; that commit applied cleanly,
the only conflict was in test-backend-ops.cpp where an unrelated intervening
commit's own test additions landed near this block.

test-backend-ops on an A10 (lambdalabs): MUL_MAT 1297/1297, including all
FWHT/Hadamard cases (18 existing, 4 new F32 wide, 4 new F16 wide, 2 new
many-rows, 1 too-big boundary moved to 16384).
@bri-prism
bri-prism marked this pull request as ready for review September 28, 2026 07:08
@bri-prism
bri-prism requested a review from a team as a code owner September 28, 2026 07:08
bri-prism added a commit to PrismML-Eng/llama.cpp that referenced this pull request Oct 5, 2026
The Vulkan FWHT covers block widths 64 to 512; wider blocks fall back to the
dense matmul. Extend the width list to 8192 and pick the shader per width: the
subgroup shader keeps n/subgroup_size values per invocation, so widths above
512 (or a subgroup narrow enough to exceed 64 values per invocation) use the
shared-memory shader with one row per workgroup and a 256-wide block. The rows
per workgroup become a specialization constant so the dispatch matches.

Pipelines for a width are only created when the device's shared-memory and
workgroup-invocation limits allow it; otherwise that width keeps the fallback.

This is the Vulkan counterpart to ggml-org#29100 (CUDA). Stacked on ggml-org#29101; the last
commit is the one to review.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant