Skip to content

CUDA: add new op template rms_norm_f32_vec4 for float4 vectorized load/store - #20520

Open
tehsiuhuang wants to merge 2 commits into
ggml-org:masterfrom
tehsiuhuang:float4_rms_norm_2
Open

tehsiuhuang wants to merge 2 commits into
ggml-org:masterfrom
tehsiuhuang:float4_rms_norm_2

Conversation

@tehsiuhuang

@tehsiuhuang tehsiuhuang commented Mar 13, 2026 •

Copy link
Copy Markdown
Contributor
  • Add a separate rms_norm_f32_vec4 kernel using float4 (128-bit) vectorized memory loads/stores
  • Pretty similar to the rms_norm_f32 version but implemented into a separated template to prevent the runtime branch impact and also minimize the regression chance.
  • Host-side dispatch routes to the vec4 kernel when ncols is divisible by 4 and strides are aligned; otherwise falls back to the original rms_norm_f32 kernel.

AI Usage:

  • Help find out all the possible tests and run them to collect test results.
  • Review my implementation and provided suggestion to make the logic/code clean.
  • Wordsmith the PR description :) .

Performance Results

Performance (A100, nrows=512, test-backend-ops perf, 5-run avg):
[512,512]: 427 -> 624 GB/s (+46%)
[768,512]: 626 -> 850 GB/s (+36%)
[1024,512]: 495 -> 645 GB/s (+30%)
[2048,512]: 911 -> 1171 GB/s (+28%)
[3072,512]: 1220 -> 1490 GB/s (+22%)
[5120,512]: 1668 -> 1815 GB/s (+9%)
Scalar fallback (4097,512): 1476 -> 1471 GB/s (no regression)

RTX 4070 Laptop GPU (5-run avg, test-backend-ops perf -o RMS_NORM -b CUDA0):
[512,512]: 474 -> 608 GB/s (+28%)
[768,512]: 580 -> 743 GB/s (+28%)
[1024,512]: 254 -> 313 GB/s (+23%)
[2048,512]: 436 -> 551 GB/s (+26%)
[3072,512]: 555 -> 672 GB/s (+21%)
[5120,512]: 696 -> 811 GB/s (+17%)
[8192,512]: 741 -> 828 GB/s (+12%)
Scalar fallback (4097,512): 678 -> 669 GB/s (no meaningful regression)

Why a separate kernel instead of a runtime branch

A runtime if (use_vec4) inside the existing rms_norm_f32 kernel was tested and
found to cause significant performance regression, even on the scalar fallback path.

Three versions were benchmarked (A100, nrows=512, test-backend-ops perf, 5-run avg):

Dimensions Path Baseline Runtime Branch Separate Kernel
[512,512] vec4 427 GB/s 586 GB/s 627 GB/s
[768,512] vec4 626 GB/s 826 GB/s 839 GB/s
[1024,512] vec4 495 GB/s 429 GB/s 642 GB/s
[2048,512] vec4 911 GB/s 806 GB/s 1174 GB/s
[3072,512] vec4 1220 GB/s 1086 GB/s 1489 GB/s
[4096,512] vec4 1739 GB/s 1292 GB/s 1687 GB/s
[5120,512] vec4 1668 GB/s 1499 GB/s 1825 GB/s
[4097,512] scalar 1476 GB/s 1162 GB/s 1479 GB/s

The runtime branch version regresses at large dimensions -- [4096,512] drops to
0.74x of baseline on the vec4 path, and the scalar fallback [4097,512] drops to
0.79x even though it runs identical code to the baseline. This is because the
CUDA compiler must generate code for both paths within the same function, leading
to suboptimal register allocation and/or instruction cache utilization regardless
of which path is actually taken at runtime.

The separate kernel approach eliminates this entirely -- the compiler optimizes
each kernel independently, resulting in no regression on the scalar path (1.00x)
and consistent speedups on the vec4 path (+22% to +47% at medium dimensions).

Correctness:

RMS_NORM 17/17,
RMS_NORM_MUL_ADD 30/30,
ADD_RMS_NORM 25/25,
RMS_NORM_MUL_ROPE 72/72 passed.

ctest

cd build && ctest --output-on-failure
Those failures are NOT related to the op perf changes
image

End-to-end model benchmark (optional, requires GGUF model) (A100, RTX 4070)

(Looks like RMSNorm is pretty minior in those real model :) )
./bin/llama-bench -m /models/qwen2.5-1.5b-instruct-q4_k_m.gguf -t 1 -ngl 999 -r 5
./bin/llama-bench -m /models/Meta-Llama-3-8B-Instruct-Q4_K_M.gguf -t 1 -ngl 999 -r 5

Command

Build Commands
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --target test-backend-ops -j$(nproc)
cmake --build build --target llama-bench -j$(nproc)
cd build

Correctness tests (A100, RTX 4070)
./bin/test-backend-ops test -o RMS_NORM -b CUDA0
./bin/test-backend-ops test -o RMS_NORM_MUL_ADD -b CUDA0
./bin/test-backend-ops test -o ADD_RMS_NORM -b CUDA0
./bin/test-backend-ops test -o RMS_NORM_MUL_ROPE -b CUDA0

Performance benchmark (isolated RMS_NORM op) (A100, RTX 4070)
./bin/test-backend-ops perf -o RMS_NORM -b CUDA0

@github-actions github-actions Bot added testing Everything test related Nvidia GPU Issues specific to Nvidia GPUs ggml changes relating to the ggml tensor library for machine learning labels Mar 13, 2026
Comment thread ggml/src/ggml-cuda/norm.cu
Add a separate rms_norm_f32_vec4 kernel using float4 (128-bit) vectorized
memory loads/stores. Host-side dispatch routes to the vec4 kernel when
ncols is divisible by 4 and strides are aligned; otherwise falls back to
the original rms_norm_f32 kernel which is completely untouched.

A separate kernel is used instead of a runtime branch inside the existing
kernel to avoid register pressure and instruction cache pollution that
would degrade the scalar path (~22% measured regression with runtime if).

Performance (A100, nrows=512, test-backend-ops perf, 5-run avg):
  [512,512]:  427 -> 624 GB/s (+46%)
  [768,512]:  626 -> 850 GB/s (+36%)
  [1024,512]: 495 -> 645 GB/s (+30%)
  [2048,512]: 911 -> 1171 GB/s (+28%)
  [3072,512]: 1220 -> 1490 GB/s (+22%)
  [5120,512]: 1668 -> 1815 GB/s (+9%)
  Scalar fallback (4097,512): 1476 -> 1471 GB/s (no regression)

Correctness: RMS_NORM 17/17, RMS_NORM_MUL_ADD 30/30,
ADD_RMS_NORM 25/25, RMS_NORM_MUL_ROPE 72/72 passed.
@tehsiuhuang tehsiuhuang changed the title CUDA: add float4 vectorized load/store for rms_norm_f32 CUDA: add new op template rms_norm_f32_vec4 for float4 vectorized load/store Mar 14, 2026
@tehsiuhuang
tehsiuhuang marked this pull request as ready for review March 14, 2026 01:49
@tehsiuhuang
tehsiuhuang requested a review from ggerganov as a code owner March 14, 2026 01:49
@am17an

am17an commented Mar 14, 2026 •

Copy link
Copy Markdown
Contributor

I see no change in performance probably because the alignment of 16 is not checked properly (so this path is never taken)

Model Test t/s master t/s float4_rms_norm_2 Speedup
llama 8B Q4_K_M pp512 5144.39 5128.36 1.00
llama 8B Q4_K_M pp2048 5158.44 5141.04 1.00
llama 8B Q4_K_M tg128 139.57 139.84 1.00

@tehsiuhuang

Copy link
Copy Markdown
Contributor Author

I see no change in performance probably because the alignment of 16 is not checked properly (so this path is never taken)

Model Test t/s master t/s float4_rms_norm_2 Speedup
llama 8B Q4_K_M pp512 5144.39 5128.36 1.00
llama 8B Q4_K_M pp2048 5158.44 5141.04 1.00
llama 8B Q4_K_M tg128 139.57 139.84 1.00

I observed the same results in my end-to-end tests (as detailed in my End-to-end model benchmark section).

You're absolutely right—since RMSNorm accounts for a very small fraction of the total inference time compared to GEMM, any local optimization is easily masked by system noise and the dominant overhead of matrix multiplications.

This is exactly why I shifted to profiling the kernel directly.

@am17an

am17an commented Mar 14, 2026

Copy link
Copy Markdown
Contributor

Are you sure the path is being exercised? What I meant is the the real model test might not have the alignment you need

@tehsiuhuang

tehsiuhuang commented Mar 14, 2026 •

Copy link
Copy Markdown
Contributor Author

Are you sure the path is being exercised? What I meant is the the real model test might not have the alignment you need

thanks for spending time on the code comments. The float4 path is exercised in the real model run via rms_norm_mul_f32_cuda (fused RMS_NORM+MUL). Here are the data and logs.

Command

GGML_CUDA_DEBUG_RMS_NORM=1 ./build/bin/llama-bench -m ~/llama.cpp/models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf -ngl 99 -p 512,2048,0 -n 0,0,128 -dev CUDA0

(Same test setup as your pp512 / pp2048 / tg128; the env var only enables the debug log to stderr.)


Benchmark results (t/s)

ggml_cuda_init: found 1 CUDA devices (Total VRAM: 7805 MiB):
  Device 0: NVIDIA GeForce RTX 4070 Laptop GPU, compute capability 8.9, VMM: yes, VRAM: 7805 MiB (7245 MiB free)
| model                          |       size |     params | backend    | ngl | dev          |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------------ | --------------: | -------------------: |
| llama 8B Q4_K - Medium         |   4.58 GiB |     8.03 B | CUDA       |  99 | CUDA0        |           pp512 |      2047.93 ± 49.19 |
| llama 8B Q4_K - Medium         |   4.58 GiB |     8.03 B | CUDA       |  99 | CUDA0        |          pp2048 |      1711.49 ± 10.41 |
| llama 8B Q4_K - Medium         |   4.58 GiB |     8.03 B | CUDA       |  99 | CUDA0        |           tg128 |         45.07 ± 0.04 |

Path and alignment (first 20 RMS_NORM calls)

From my local debugging messages, all calls go through rms_mul (i.e. rms_norm_mul_f32_cuda). Alignment is satisfied: Every call has x%16=0, dst%16=0, and all conditions ncols%4=1, stride_ok=1, align_x=1, align_dst=1, so the float4 path is taken (vec4=21, scalar=0).

[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #0 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #1 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #2 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #3 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #4 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #5 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #6 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #7 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #8 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #9 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #10 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #11 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #12 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #13 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #14 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #15 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #16 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #17 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #18 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] rms_mul #19 ncols=4096 nrows=512 s_row=4096 s_chan=2097152 s_samp=2097152 x%16=0 dst%16=0 → use_vec4=1 (ncols%4=1 stride_ok=1 align_x=1 align_dst=1)
[GGML_CUDA_DEBUG_RMS_NORM] (further calls omitted; vec4=21 scalar=0 so far)

So in this real model test (Llama 8B Q4_K_M, pp512/pp2048/tg128), the float4 path is taken and alignment is satisfied.

@am17an

am17an commented Mar 14, 2026

Copy link
Copy Markdown
Contributor

Ok, so this PR does not have an effect on end to end model performance. I think merging this would just add to the maintenance burden, we can keep the PR open in case there is some use-case for it later down the line

@tehsiuhuang

tehsiuhuang commented Mar 14, 2026 •

Copy link
Copy Markdown
Contributor Author

Ok, so this PR does not have an effect on end to end model performance. I think merging this would just add to the maintenance burden, we can keep the PR open in case there is some use-case for it later down the line

Sounds like a good plan. Thanks for spending time with me!

@tehsiuhuang

Copy link
Copy Markdown
Contributor Author

Add some data points just in case we want to know the kernel call counts in the future:

nsys profile -o report --stats=true ./bin/llama-bench -m ~/llama.cpp/models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf -ngl 99 -p 512 -n 32 -dev CUDA0

image

@tehsiuhuang
tehsiuhuang requested a review from a team as a code owner April 13, 2026 01:07
@gaugarg-nv

gaugarg-nv commented Apr 13, 2026 •

Copy link
Copy Markdown
Contributor

@tehsiuhuang I think you should consider benchmarking this PR for Gemma-4 models specifically - Gemma-4-E2B, Gemma-4-E4B and Gemma-4-26B-A4B.
Gemma-4 has pre/post attention and FFN norms and QK Norms. For small Gemma-4 models, RMSNorm optimizations can show improvement in e2e speed-up.

@tehsiuhuang

Copy link
Copy Markdown
Contributor Author

@tehsiuhuang I think you you consider benchmarking this PR for Gemma-4 models specifically - Gemma-4-E2B, Gemma-4-E4B and Gemma-4-26B-A4B. Gemma-4 has pre/post attention and FFN norms and QK Norms. For small Gemma-4 models, RMSNorm optimizations can show improvement in e2e speed-up.

Good point... never running on a even smaller model. Let me collect those data cc @am17an

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Nvidia GPU Issues specific to Nvidia GPUs testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants