Repository navigation
ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) - #29675
Conversation
| case GGML_TYPE_BF16: | ||
| { | ||
| ggml_compute_forward_unary_bf16(params, dst); | ||
| } break; |
There was a problem hiding this comment.
This should be ggml_compute_forward_silu_bf16 to follow the existing pattern. Otherwise, we should consolidate the f16 and f32 paths in a similar "unary" path.
JohannesGaessler
left a comment
There was a problem hiding this comment.
The CUDA changes LGTM. In principle it should be possible to improve performance further by loading more than 2 bytes per thread at once.
…g#29675) * ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) * ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops Assisted-by: Claude Opus 5.5 * CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build * ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
The kernels assert F32/F16 (GLU) and F32 (SCALE), so the BF16 cases that test-backend-ops gained with ggml-org#29675 aborted the run. Assisted-by: Claude Opus 5.5
|
@am17an thanks for this. We're interested in the BF16-activations follow-up you mentioned. We run a small vision-transformer world model on ggml with large batches and tiny weights, so we're bound by activation traffic. A BF16 result from |
…g#29675) * ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) * ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops Assisted-by: Claude Opus 5.5 * CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build * ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
…g#29675) * ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) * ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops Assisted-by: Claude Opus 5.5 * CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build * ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB (cherry picked from commit 2090f60)
…g#29675) * ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) * ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops Assisted-by: Claude Opus 5.5 * CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build * ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB (cherry picked from commit 2090f60)
|
Is this the PR that increases quality for Gemma 4 QAT or will it be the upcoming BF16 activations PR that is based on this? |
Overview
Add bf16 variants of element-wise ops. The reason for adding these is to keep activations in bf16 (which will added for a later PR). Anyway I think it doesn't hurt to have these.
Additional information
Requirements