Repository navigation
cuda : support arbitrary striding for unary ops on f16, f32, and bf16 - #29781
Conversation
|
Is there a concrete use case? |
|
Hey @JohannesGaessler thanks for reaching out. My understanding is that Georgi added PR #19375 to optimize the qwen3next graph by reducing redundant inner tensors, and he also added PR #19511 where he added non-contiguous unary op support for Metal and CPU. While reading through #19375, I observed that we currently don't have non-contiguous unary op support in CUDA. Because of this, the graph we create for qwen3next is creating |
|
Also another observation in #19597 we are also doing |
|
Need to fix hip tests |
99256fe to
87a66c8
Compare
|
looks like we have to run the CI again |
87a66c8 to
938a993
Compare
|
Hey @am17an @JohannesGaessler, that HIP failure was an unrelated 1-VGPR spill in |
|
@saady789 sure, also - could create a follow-up PR for removing the conts from the models you mentioned? Should result in a speedup. Thanks! |
thanks I can do that in a follow up pull req. It looks like HIP and webgpu are still failing |
JohannesGaessler
left a comment
There was a problem hiding this comment.
Sorry, I removed 2 newlines that were added by this PR. But anyways, the "HIP quality check" is mostly used to detect register spills which is only going to happen in matrix multiplications and FlashAttention. A kernel like this is basically impossible to cause it so something else must have broken it (it was already broken for other recent PRs). Similarly for the WebGPU failure: an internal change to the CUDA backend cannot cause a failure in the WebGPU backend so it can be ignored.
|
thanks @JohannesGaessler we will wait for approval from @am17an since I think his approval is stale after that we can merge |
|
Please rebase, then I'll merge. |
|
Sorry, wrong button. |
4cfb33d to
9158693
Compare
thanks just rebased and resolved merge conflict |
…ggml-org#29781) * cuda : support arbitrary striding for unary ops on f16, f32, and bf16 * Remove added newline --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de> (cherry picked from commit 08246a2)
Overview
Adds support for arbitrary 4D strided and non-contiguous tensors to CUDA unary ops across
F16,F32, andBF16.Relates to #14909 and supersedes #28821, #26504, and abandoned #14639.
Unlike prior attempts which only relaxed the check to
ggml_is_contiguous_rowsor replaced the flat 1D kernel (causing throughput regressions for standard contiguous tensors), this PR:unary_op_kernel_strided) handling arbitrary byte-offsets without assuming row contiguity.unary_cuda) whenggml_is_contiguous(src0)is true.F16,F32, andBF16.Additional information
Implementation
ggml-cuda.cu: Removed the contiguity requirement for unary ops acrossF16,F32, andBF16while retaining theBF16XIELUguard.unary.cu: Addedunary_op_kernel_stridedto map indices to 4D coordinates and index memory via source byte strides (nb00, nb01, nb02, nb03). Contiguous inputs remain on the fast 1D path.Testing
Validated with
test-backend-ops -b CUDA -o UNARYtesting both contiguous (v=0) and non-contiguous/view (v=1) modes.I have read and agree with the contributing guidelines
AI usage disclosure: YES - Used AI to understand the codebase and file structure.