Conversation
bri-prism
force-pushed
the
up/fwht-cuda-wide
branch
2 times, most recently
from
September 21, 2026 19:16
249bd43 to
18187db
Compare
bri-prism
force-pushed
the
up/fwht-cuda-wide
branch
from
September 26, 2026 16:10
18187db to
843fb6e
Compare
The CUDA FWHT covers widths 64 to 512. It runs one row per warp and keeps N/32 values per lane, so wider blocks need more registers per lane than that layout allows. fwht_cuda_block runs one row per thread block with 256 threads, so each thread keeps N/256 values. Stages below the warp width still shuffle, those up to the block width go through shared memory, and the rest stay in registers. Same butterfly and sign convention as the warp kernel. Widths 64 to 512 keep the warp kernel. 1024 through 8192 use the new one, for both F32 and F16 sources. ggml_cuda_op_mul_mat_use_fwht (the shared supports_op/dispatch predicate added in ggml-org#29096) does not check width, so it needed no change here: any width it admits that ggml_cuda_op_fwht can't serve already falls through correctly to the cuBLAS path. Rebased onto current master with ggml-org#29096's F16 commit underneath it, since this depends on the same F16 template infrastructure; that commit applied cleanly, the only conflict was in test-backend-ops.cpp where an unrelated intervening commit's own test additions landed near this block. test-backend-ops on an A10 (lambdalabs): MUL_MAT 1297/1297, including all FWHT/Hadamard cases (18 existing, 4 new F32 wide, 4 new F16 wide, 2 new many-rows, 1 too-big boundary moved to 16384).
bri-prism
force-pushed
the
up/fwht-cuda-wide
branch
from
September 26, 2026 22:38
843fb6e to
7d9305f
Compare
bri-prism
marked this pull request as ready for review
September 28, 2026 07:08
This was referenced Sep 30, 2026
bri-prism
added a commit
to PrismML-Eng/llama.cpp
that referenced
this pull request
Oct 5, 2026
The Vulkan FWHT covers block widths 64 to 512; wider blocks fall back to the dense matmul. Extend the width list to 8192 and pick the shader per width: the subgroup shader keeps n/subgroup_size values per invocation, so widths above 512 (or a subgroup narrow enough to exceed 64 values per invocation) use the shared-memory shader with one row per workgroup and a 256-wide block. The rows per workgroup become a specialization constant so the dispatch matches. Pipelines for a width are only created when the device's shared-memory and workgroup-invocation limits allow it; otherwise that width keeps the fallback. This is the Vulkan counterpart to ggml-org#29100 (CUDA). Stacked on ggml-org#29101; the last commit is the one to review.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
The CUDA FWHT covers block widths 64 to 512. It runs one row per warp and keeps
N/32values per lane, so wider blocks need more registers per lane than that layout allows.
fwht_cuda_blockruns one row per thread block with 256 threads, so each thread keepsN/256values instead. Stages below the warp width still shuffle, those up to the blockwidth go through shared memory, and the rest stay in registers. The butterfly and sign
convention match the existing kernel.
Widths 64 to 512 keep the warp kernel and are untouched. 1024 through 8192 use the new
one, for both F32 and F16 sources.
This is the CUDA counterpart to #29095 (Metal). Stacked on #29096; the last commit is the
one to review here.
Additional information
Tested on an H100 PCIe (compute capability 9.0, CUDA 12.8,
-DCMAKE_CUDA_ARCHITECTURES=90):test-backend-ops -o MUL_MAT_HADAMARD25/25, including the new 1024/2048/4096/8192 casesin F32 and F16, and the 16384 case that falls back
test-backend-ops -o MUL_MAT1297/1297, no regressionsA passing suite does not by itself show the new kernel ran, since a Hadamard-hinted matmul that
falls back to the generic path computes the same values. To confirm the wide kernel is actually
selected I temporarily made
fwht_cuda_blockscale its output by 2 and re-ran, which gave15/25 with exactly ten failures:
The change was reverted and the binary rebuilt before the numbers above were taken.
8192 allocates 32 KB of shared memory per block, within the 48 KB default limit. If that is too
aggressive for older parts I am happy to cap the width list lower.
Requirements
structure of the existing warp kernel and the Metal version, and to format this description to
the PR template. I reviewed every line and take full responsibility for the changes.