Repository navigation
CUDA : looped PAD kernel for more than 65535 rows or slices - #30147
Merged
Merged
Conversation
JohannesGaessler
approved these changes
Oct 8, 2026
gaugarg-nv
approved these changes
Oct 8, 2026
feal87
added a commit
to feal87/myllama.cpp
that referenced
this pull request
Oct 8, 2026
Upstream brings CUDA top-k/argsort/mmq fixes (PR ggml-org#28713, ggml-org#29953, ggml-org#30147, ggml-org#29453), a DFlash output-head fix (ggml-org#30111) and a UI fix (ggml-org#29668). Conflicts and resolution: - top-k.cu / argsort.cu: upstream reworked the top-k selection and the same radix path the fork added. Take upstream's shape-based dispatch and int64_t fixes, keep the fork workspace cap in the shared ggml_cuda_chunk_nrows (GGML_CUDA_ARGSORT_CHUNK_MB) and the fork benchmark-only GGML_CUDA_TOPK_IMPL selector used by bench-topk.py. - mmq.cu: upstream PR ggml-org#29953 fixes the same mul_mat_id padding over-read the fork patched in 1386ec2, but sizes the padding from the chosen J tile in both branches. Take upstream's version. - test-backend-ops.cpp: keep the fork GGML_TOPK_BENCH shapes and add upstream's chunk-spanning cases. Assisted-by: pi (deepseek-flash)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
CUDA
GGML_OP_PADkernel puts the output's row count ingridDim.yand its slice count (ne2*ne3) ingridDim.z. CUDA limits both to 65535. When a padded tensor exceeds either limit, the launch fails withinvalid configuration argumentand ggml aborts. The CPU backend handles the same shapes correctly.ggml/src/ggml-cuda/pad.cu,pad_f32_cuda():The kernel reads
i1 = blockIdx.yandi2/i3fromblockIdx.z, so the grid has to cover the tensor exactly. Anyne1 > 65535orne2*ne3 > 65535(counted on the padded output) cannot be launched.The fix is to do a looped version of the kernel. The grid is capped at 65535 in y and z and blocks loop over rows or slices beyond that. As @JohannesGaessler suggested in last PR this is preferable over the two kernel variant.
Additional information
How to reproduce
New
test-backend-opseval cases:On unmodified
master,build/bin/test-backend-ops -o PAD -b CUDA0passes the existing cases, then aborts on the first new one:How we hit this bug
The cache-aware streaming Nemotron ASR encoder's causal conv subsampling pads a [T, F, 256, B] tensor (256 conv channels, batch B). Batch 255 is the largest that launches on master, batch 256 gives
ne2 * ne3 = 65536and aborts.Perf Impact
Environment: RTX A5000 (sm_86), driver 580.173, CUDA 12.8, at upstream
03aa006.Note: there is 3-20% perf hit with this change
Correctness: all 37 PAD cases pass on CUDA0 (compared against CPU), including the 3 new ones.
Requirements