Skip to content

ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) - #29675

Merged
am17an merged 4 commits into
masterfrom
aman/cuda-bf16-elementwise
Sep 30, 2026
Merged

am17an merged 4 commits into
masterfrom
aman/cuda-bf16-elementwise

Conversation

@am17an

@am17an am17an commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Overview

Add bf16 variants of element-wise ops. The reason for adding these is to keep activations in bf16 (which will added for a later PR). Anyway I think it doesn't hurt to have these.

Additional information

Requirements

@am17an
am17an requested review from a team and ggerganov as code owners September 29, 2026 18:37
Comment thread ggml/src/ggml-cpu/ops.cpp
Comment on lines +2701 to +2704
case GGML_TYPE_BF16:
{
ggml_compute_forward_unary_bf16(params, dst);
} break;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be ggml_compute_forward_silu_bf16 to follow the existing pattern. Otherwise, we should consolidate the f16 and f32 paths in a similar "unary" path.

@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Sep 29, 2026

@JohannesGaessler JohannesGaessler left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The CUDA changes LGTM. In principle it should be possible to improve performance further by loading more than 2 bytes per thread at once.

@am17an
am17an requested a review from wine99 as a code owner September 30, 2026 02:57
@am17an
am17an merged commit 2090f60 into master Sep 30, 2026
35 of 37 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
…g#29675)

* ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)

* ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops

Assisted-by: Claude Opus 5.5

* CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build

* ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
cwriter pushed a commit to cwriter/llama.cpp that referenced this pull request Oct 4, 2026
The kernels assert F32/F16 (GLU) and F32 (SCALE), so the BF16 cases that
test-backend-ops gained with ggml-org#29675 aborted the run.

Assisted-by: Claude Opus 5.5
@renmengye

Copy link
Copy Markdown

@am17an thanks for this. We're interested in the BF16-activations follow-up you mentioned. We run a small vision-transformer world model on ggml with large batches and tiny weights, so we're bound by activation traffic. A BF16 result from mul_mat and flash_attn_ext, plus BF16 norm on CUDA, would let us keep a whole block in BF16. Happy to test it on our workload when it lands.

frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…g#29675)

* ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)

* ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops

Assisted-by: Claude Opus 5.5

* CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build

* ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
Wizard815 pushed a commit to Wizard815/mx-llama.cpp-Rocm10 that referenced this pull request Oct 6, 2026
…g#29675)

* ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)

* ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops

Assisted-by: Claude Opus 5.5

* CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build

* ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB

(cherry picked from commit 2090f60)
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
…g#29675)

* ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)

* ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops

Assisted-by: Claude Opus 5.5

* CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build

* ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB

(cherry picked from commit 2090f60)
@Dampfinchen

Copy link
Copy Markdown

Is this the PR that increases quality for Gemma 4 QAT or will it be the upcoming BF16 activations PR that is based on this?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning OpenVINO testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants