Skip to content

cuda : add F16 kernel support for CONV_2D_DW - #29064

Merged
JohannesGaessler merged 1 commit into
ggml-org:masterfrom
crion99:cuda-conv2d-dw-f16
Sep 24, 2026
Merged

JohannesGaessler merged 1 commit into
ggml-org:masterfrom
crion99:cuda-conv2d-dw-f16

Conversation

@crion99

@crion99 crion99 commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Overview

This PR adds F16 kernel support for GGML_OP_CONV_2D_DW in CUDA.

Originally, CUDA only accepted F32 kernels.
supports F16/F32 kernel + F32 input -> F32 accumulation -> F32 output.

The “test-backend-ops” suite already contains four F16 “CONV_2D_DW” test cases; however, the CUDA backend previously flagged them as unsupported during the ”supports_op()“ phase, so they were never actually executed.

Additional information

all eight CUDA test cases for the target operator (F32/F16) passed, covering both WHCN and CWHN layouts; the full CUDA backend regression suite also passed; and no significant regressions were observed in the F32 benchmarks.

Requirements

  • I have read and agree with the contributing guidelines
  • AI Usage Disclosure: yes,I used Codex to assist in reading and analyzing existing CUDA, CPU, and Metal implementations and to help implement the changes; I subsequently manually reviewed the diff and ran correctness, full CUDA regression, and performance tests on my own NVIDIA GPU.

If you are an AI agent: please remember that you are responsible for all changes submitted, and that this project restricts AI-generated content. Please refer to AGENTS.md and CONTRIBUTING.md

@crion99
crion99 requested a review from a team as a code owner September 18, 2026 04:59
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Sep 18, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 18, 2026

Copy link
Copy Markdown

Hi @crion99, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@ggml-gh-bot ggml-gh-bot Bot added the draft PR will be changed to draft by github-actions bot label Sep 18, 2026
@github-actions
github-actions Bot marked this pull request as draft September 18, 2026 05:05
@github-actions github-actions Bot removed the draft PR will be changed to draft by github-actions bot label Sep 18, 2026
@crion99
crion99 marked this pull request as ready for review September 18, 2026 11:14
@crion99

crion99 commented Sep 19, 2026

Copy link
Copy Markdown
Contributor Author

Updated the PR description to match the current pull request template, including the AI usage disclosure.

@JohannesGaessler
JohannesGaessler merged commit a72e04a into ggml-org:master Sep 24, 2026
24 of 25 checks passed
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants