Skip to content

CUDA: fix NORM/RMS_NORM/L2_NORM for more than 65535 channels or samples - #28175

Open
edenfunf wants to merge 1 commit into
ggml-org:masterfrom
edenfunf:fix-norm-griddim
Open

edenfunf wants to merge 1 commit into
ggml-org:masterfrom
edenfunf:fix-norm-griddim

Conversation

@edenfunf

@edenfunf edenfunf commented Sep 1, 2026 •

Copy link
Copy Markdown

Overview

Fixes #27901, refs #27911.

The CUDA norm kernels use (nrows, nchannels, nsamples) as the launch grid, but gridDim.y/z is limited to 65535. When ne[2] or ne[3] exceeds that, the launch fails with invalid argument.

Affected kernels:

  • norm_f32
  • rms_norm_f32
  • l2_norm_f32
  • rms_norm_mul_rope_f32

The fix follows #25103 and #22944: clamp grid.y/z to 65535 and loop over the remaining channels/samples in the kernel.

Since gridDim.y can no longer be used as the real channel count, nchannels and nsamples are now passed explicitly.

Multi-warp variants also need a trailing __syncthreads() because block_reduce reuses the same shared buffer across loop iterations.

Since the previous version

  • Rebased onto master with CUDA: fuse RMS_NORM + SCALE into one kernel #29393 and integrated the new do_scale path.
  • Removed the relocated ggml_cuda_pdl_lc() in l2_norm_f32 based on review feedback.
  • Merged the rms_norm_mul_rope NMSE override with the existing one on master.
  • Added [[maybe_unused]] for mulc / addc in non-fused instantiations.

Testing

Tested on RTX 5070, CUDA 13.3.

Added 7 cases with ne[2] or ne[3] = 65536. They fail on master and pass against the CPU backend with this patch.

Full test-backend-ops:

  • 17689 / 17689 passed

Also checked additional boundary cases for 1024-thread paths, nrows > 1, sample-axis overflow, 65535 boundary, multiple loop iterations, non-contiguous views, and fused variants.

Performance

Shape master patched
RMS_NORM 4096x512 14.8 us 14.3 us
RMS_NORM 8192x1 2.5 us 2.0 us
NORM 4096x512 17.1 us 17.0 us
RMS_NORM [4,1,65535,1] 136 us 159 us
NORM [4,1,65535,1] 52 us 57 us

Normal shapes are flat or slightly faster.

The slowdown is only on the degenerate ~65k-block cases with very few active threads per block. I tried a few loop variants and none improved it, so this looks like loop overhead rather than a specific implementation detail.

A split fast/slow path could avoid that regression, but would add another kernel instantiation. For now this keeps a single loop-based path, consistent with the existing getrows.cu approach.

Requirements

  • I have read and agree with the [contributing guidelines](...)
  • AI usage disclosure: YES - AI was used for issue investigation and part of the implementation. I reviewed and verified the changes.

@edenfunf
edenfunf requested review from a team and ggerganov as code owners September 1, 2026 14:27
@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Sep 1, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 1, 2026

Copy link
Copy Markdown

Hi @edenfunf, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@ggml-gh-bot ggml-gh-bot Bot added the draft PR will be changed to draft by github-actions bot label Sep 1, 2026
@github-actions
github-actions Bot marked this pull request as draft September 1, 2026 14:32
@github-actions github-actions Bot removed the draft PR will be changed to draft by github-actions bot label Sep 1, 2026
@edenfunf
edenfunf marked this pull request as ready for review September 1, 2026 14:36
@edenfunf

Copy link
Copy Markdown
Author

@ggerganov If you have a chance, could you take a look at this PR? Thanks!

@JohannesGaessler JohannesGaessler left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry for the long radio silence. I took a break from llama.cpp and am currently working through my backlog.

Comment thread ggml/src/ggml-cuda/norm.cu Outdated

x += sample*stride_sample + channel*stride_channel + row*stride_row;
dst += ((sample*nchannels + channel)*nrows + row)*ncols;
ggml_cuda_pdl_lc();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
ggml_cuda_pdl_lc();

The PDL launch completion does not affect correctness has to be placed essentially experimentally. If should not just be moved around like this. So I think the safest bet is to just remove it. cc @aendk

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed ggml_cuda_pdl_lc() from l2_norm_f32 as suggested, so it now matches master.

Also rebased onto master with #29393 and integrated the do_scale path. Full test-backend-ops passes.

@JohannesGaessler JohannesGaessler added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Oct 7, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eval bug: rms_norm_f32 exceeds the CUDA gridDim.y limit (65535) at n_ctx 262144

2 participants