Skip to content

cuda : support arbitrary striding for unary ops on f16, f32, and bf16 - #29781

Merged
JohannesGaessler merged 2 commits into
ggml-org:masterfrom
saady789:cuda-strided-unary
Oct 8, 2026
Merged

JohannesGaessler merged 2 commits into
ggml-org:masterfrom
saady789:cuda-strided-unary

Conversation

@saady789

@saady789 saady789 commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Overview

Adds support for arbitrary 4D strided and non-contiguous tensors to CUDA unary ops across F16, F32, and BF16.
Relates to #14909 and supersedes #28821, #26504, and abandoned #14639.

Unlike prior attempts which only relaxed the check to ggml_is_contiguous_rows or replaced the flat 1D kernel (causing throughput regressions for standard contiguous tensors), this PR:

  1. Implements true arbitrary 4D strided indexing (unary_op_kernel_strided) handling arbitrary byte-offsets without assuming row contiguity.
  2. Preserves the zero-overhead flat 1D contiguous kernel (unary_cuda) when ggml_is_contiguous(src0) is true.
  3. Supports F16, F32, and BF16.

Additional information

Implementation

  • ggml-cuda.cu: Removed the contiguity requirement for unary ops across F16, F32, and BF16 while retaining the BF16 XIELU guard.
  • unary.cu: Added unary_op_kernel_strided to map indices to 4D coordinates and index memory via source byte strides (nb00, nb01, nb02, nb03). Contiguous inputs remain on the fast 1D path.

Testing

  • Validated with test-backend-ops -b CUDA -o UNARY testing both contiguous (v=0) and non-contiguous/view (v=1) modes.

  • I have read and agree with the contributing guidelines

  • AI usage disclosure: YES - Used AI to understand the codebase and file structure.

@saady789
saady789 requested a review from a team as a code owner October 1, 2026 03:47
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Oct 1, 2026
@JohannesGaessler

Copy link
Copy Markdown
Contributor

Is there a concrete use case?

@saady789

saady789 commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Hey @JohannesGaessler thanks for reaching out.

My understanding is that Georgi added PR #19375 to optimize the qwen3next graph by reducing redundant inner tensors, and he also added PR #19511 where he added non-contiguous unary op support for Metal and CPU. While reading through #19375, I observed that we currently don't have non-contiguous unary op support in CUDA. Because of this, the graph we create for qwen3next is creating ggml_cont to do unary ops even for Metal and CPU, even though we already have non-contiguous support for them. If we can implement non-contiguous unary op support for CUDA, we can get rid of ggml_cont on 3 occasions in qwen3next.cpp, while keeping the contiguous path fast.

@saady789

saady789 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Also another observation in #19597 we are also doing ggml_cont and I see there is a comment to remove this once we implement the unary kernel for non contiguous tensors and hence this can be another opportunity to improve

@am17an

am17an commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Need to fix hip tests

@saady789
saady789 force-pushed the cuda-strided-unary branch from 99256fe to 87a66c8 Compare October 5, 2026 04:16
@saady789

saady789 commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

looks like we have to run the CI again

@saady789
saady789 force-pushed the cuda-strided-unary branch from 87a66c8 to 938a993 Compare October 6, 2026 02:51
@saady789

saady789 commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Hey @am17an @JohannesGaessler, that HIP failure was an unrelated 1-VGPR spill in fattn-mma-f16.cuh. Just rebased onto latest master can we run CI again sorry abt that

@am17an

am17an commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

@saady789 sure, also - could create a follow-up PR for removing the conts from the models you mentioned? Should result in a speedup. Thanks!

@saady789

saady789 commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

@saady789 sure, also - could create a follow-up PR for removing the conts from the models you mentioned? Should result in a speedup. Thanks!

thanks I can do that in a follow up pull req. It looks like HIP and webgpu are still failing

Comment thread ggml/src/ggml-cuda/ggml-cuda.cu Outdated
Comment thread ggml/src/ggml-cuda/unary.cu Outdated

@JohannesGaessler JohannesGaessler left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I removed 2 newlines that were added by this PR. But anyways, the "HIP quality check" is mostly used to detect register spills which is only going to happen in matrix multiplications and FlashAttention. A kernel like this is basically impossible to cause it so something else must have broken it (it was already broken for other recent PRs). Similarly for the WebGPU failure: an internal change to the CUDA backend cannot cause a failure in the WebGPU backend so it can be ignored.

@saady789

saady789 commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

thanks @JohannesGaessler we will wait for approval from @am17an since I think his approval is stale after that we can merge

@JohannesGaessler

Copy link
Copy Markdown
Contributor

Please rebase, then I'll merge.

@JohannesGaessler

Copy link
Copy Markdown
Contributor

Sorry, wrong button.

@saady789
saady789 force-pushed the cuda-strided-unary branch from 4cfb33d to 9158693 Compare October 8, 2026 03:15
@saady789

saady789 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Please rebase, then I'll merge.

thanks just rebased and resolved merge conflict

@JohannesGaessler
JohannesGaessler merged commit 08246a2 into ggml-org:master Oct 8, 2026
15 of 17 checks passed
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
…ggml-org#29781)

* cuda : support arbitrary striding for unary ops on f16, f32, and bf16

* Remove added newline

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
(cherry picked from commit 08246a2)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants