Repository navigation
Added SILU_BACK operation support for Metal gpu. - #25982
Merged
Merged
Conversation
Blackcyan30
marked this pull request as draft
July 22, 2026 02:54
Blackcyan30
marked this pull request as ready for review
July 22, 2026 02:54
4 tasks done
Contributor
Author
|
Hi maintainers of @/ggml-org/ggml-metal , a gentle ping to follow up on my pr, it has been open for a little over a week. It adds f32 contiguous tensor SILU_BACK support for Metal, the focused correctness test and full Metal backend op suite pass on my Apple M1 macbook. I would appreciate a review when time is good for you. |
forforever73
approved these changes
Jul 30, 2026
forforever73
left a comment
Contributor
There was a problem hiding this comment.
Thanks for the contribution. Since this is a training-related op, I don't have any concerns from my side.
ggerganov
approved these changes
Jul 31, 2026
…ion ggml_metal_op_silu_back.
TheTom
pushed a commit
to TheTom/llama-cpp-turboquant
that referenced
this pull request
Aug 3, 2026
* feat(silu_back): implemented silu_back op for f32 * fix(silu_back): removed redundant asserts in ggml-metal-ops.cpp function ggml_metal_op_silu_back.
satindergrewal
pushed a commit
to satindergrewal/llama.cpp
that referenced
this pull request
Aug 12, 2026
* feat(silu_back): implemented silu_back op for f32 * fix(silu_back): removed redundant asserts in ggml-metal-ops.cpp function ggml_metal_op_silu_back.
thecodacus
pushed a commit
to thecodacus/llama.cpp
that referenced
this pull request
Sep 7, 2026
* feat(silu_back): implemented silu_back op for f32 * fix(silu_back): removed redundant asserts in ggml-metal-ops.cpp function ggml_metal_op_silu_back.
zbrad
pushed a commit
to zbrad/llama.cpp
that referenced
this pull request
Sep 10, 2026
* feat(silu_back): implemented silu_back op for f32 * fix(silu_back): removed redundant asserts in ggml-metal-ops.cpp function ggml_metal_op_silu_back.
pl752
pushed a commit
to pl752/llama.cpp
that referenced
this pull request
Sep 15, 2026
* feat(silu_back): implemented silu_back op for f32 * fix(silu_back): removed redundant asserts in ggml-metal-ops.cpp function ggml_metal_op_silu_back.
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
* feat(silu_back): implemented silu_back op for f32 * fix(silu_back): removed redundant asserts in ggml-metal-ops.cpp function ggml_metal_op_silu_back.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Adds the Metal backend for SILU_BACK (CPU/CUDA/Vulkan already exist)
The kernel reads one element from each of the inputs (dy[gid], x[gid]) for a tensor position given by gid, computes the sigmoid and then the derivative and writes one output dx[gid], mirroring the CPU/CUDA math implementations. There is one thread per output element. Supports f32, fully contiguous tensors.
Additional information
test-backend-ops -o SILU_BACK passes on Apple M1 (SILU_BACK test 1/1, full metal suite 12845/12845 pass, checked against the CPU for reference).
#14909 issue reference
Requirements
I have read and agree with the contributing guidelines
AI usage disclosure: Yes
I used AI to help me trace the cpu and the cuda implementations as well as the GPU dispatch architecture and early little code suggestion, so I can understand the architecture being followed as well as the math, it also reviewed the code I wrote for correctness.