Skip to content

cuda: use the vector lightning indexer kernel on MUSA - #29990

Merged
ServeurpersoCom merged 2 commits into
ggml-org:masterfrom
ServeurpersoCom:cuda-lightning-indexer-musa-follow-up
Oct 5, 2026
Merged

ServeurpersoCom merged 2 commits into
ggml-org:masterfrom
ServeurpersoCom:cuda-lightning-indexer-musa-follow-up

Conversation

@ServeurpersoCom

@ServeurpersoCom ServeurpersoCom commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Follow up to #29901, which broke the MUSA Docker build: MUSA archs 21 and 22 allow only 28 KB of static shared memory, and the new tile kernel needs 33 KB.

As suggested by @am17an, MUSA now skips the tile kernel and keeps the vector kernel it used before #29901, so the tile kernel is never compiled for MUSA. CUDA and ROCm run the merged kernel unchanged.

Tested on an R9700 with the MUSA path simulated: the 4 head cases fall back to the vector kernel and test-backend-ops passes. I don't have MUSA hardware, so @yeahdongcn can test the head pass version later, which fits the tile under 28 KB: 8b84aec. The regular MUSA CI only builds arch 31, so only a Docker run will show the fix.

Additional information

cc @CISC
Follow-up #29901

Requirements

MUSA archs 21 and 22 cap static shared memory at 28 KB, and the tile
kernel staged the queries of all four heads next to the key tile for
33 KB. The queries are now staged in passes of
LIGHTNING_INDEXER_TILE_HEADS_PER_PASS heads: two on MUSA for 25 KB,
four elsewhere where the single pass folds to the previous kernel.
@ServeurpersoCom
ServeurpersoCom requested a review from CISC October 5, 2026 10:56
@ServeurpersoCom
ServeurpersoCom requested a review from a team as a code owner October 5, 2026 10:56
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Oct 5, 2026

@am17an am17an left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can just use the vec kernel for this arch

Address review from am17an: the tile kernel stays off MUSA, whose archs
21 and 22 cap static shared memory at 28 KB, below the 33 KB the tile
needs, so MUSA keeps the vector kernel it ran before. This replaces the
head passes, CUDA and ROCm run the merged kernel unchanged.
@ServeurpersoCom

Copy link
Copy Markdown
Contributor Author

we can just use the vec kernel for this arch

Done

@ServeurpersoCom ServeurpersoCom changed the title cuda: stage the lightning indexer queries in head passes for MUSA cuda: use the vector lightning indexer kernel on MUSA Oct 5, 2026
@ServeurpersoCom
ServeurpersoCom requested a review from am17an October 5, 2026 11:24
@ServeurpersoCom
ServeurpersoCom merged commit b809b88 into ggml-org:master Oct 5, 2026
13 of 17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants