Repository navigation
CUDA: add new op template rms_norm_f32_vec4 for float4 vectorized load/store - #20520
tehsiuhuang wants to merge 2 commits into
Conversation
540b479 to
3c4a25d
Compare
Add a separate rms_norm_f32_vec4 kernel using float4 (128-bit) vectorized memory loads/stores. Host-side dispatch routes to the vec4 kernel when ncols is divisible by 4 and strides are aligned; otherwise falls back to the original rms_norm_f32 kernel which is completely untouched. A separate kernel is used instead of a runtime branch inside the existing kernel to avoid register pressure and instruction cache pollution that would degrade the scalar path (~22% measured regression with runtime if). Performance (A100, nrows=512, test-backend-ops perf, 5-run avg): [512,512]: 427 -> 624 GB/s (+46%) [768,512]: 626 -> 850 GB/s (+36%) [1024,512]: 495 -> 645 GB/s (+30%) [2048,512]: 911 -> 1171 GB/s (+28%) [3072,512]: 1220 -> 1490 GB/s (+22%) [5120,512]: 1668 -> 1815 GB/s (+9%) Scalar fallback (4097,512): 1476 -> 1471 GB/s (no regression) Correctness: RMS_NORM 17/17, RMS_NORM_MUL_ADD 30/30, ADD_RMS_NORM 25/25, RMS_NORM_MUL_ROPE 72/72 passed.
3c4a25d to
0586379
Compare
|
I see no change in performance probably because the alignment of 16 is not checked properly (so this path is never taken)
|
I observed the same results in my end-to-end tests (as detailed in my End-to-end model benchmark section). You're absolutely right—since RMSNorm accounts for a very small fraction of the total inference time compared to GEMM, any local optimization is easily masked by system noise and the dominant overhead of matrix multiplications. This is exactly why I shifted to profiling the kernel directly. |
|
Are you sure the path is being exercised? What I meant is the the real model test might not have the alignment you need |
thanks for spending time on the code comments. The float4 path is exercised in the real model run via rms_norm_mul_f32_cuda (fused RMS_NORM+MUL). Here are the data and logs. CommandGGML_CUDA_DEBUG_RMS_NORM=1 ./build/bin/llama-bench -m ~/llama.cpp/models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf -ngl 99 -p 512,2048,0 -n 0,0,128 -dev CUDA0(Same test setup as your pp512 / pp2048 / tg128; the env var only enables the debug log to stderr.) Benchmark results (t/s)Path and alignment (first 20 RMS_NORM calls)From my local debugging messages, all calls go through rms_mul (i.e. rms_norm_mul_f32_cuda). Alignment is satisfied: Every call has So in this real model test (Llama 8B Q4_K_M, pp512/pp2048/tg128), the float4 path is taken and alignment is satisfied. |
|
Ok, so this PR does not have an effect on end to end model performance. I think merging this would just add to the maintenance burden, we can keep the PR open in case there is some use-case for it later down the line |
Sounds like a good plan. Thanks for spending time with me! |
|
@tehsiuhuang I think you should consider benchmarking this PR for Gemma-4 models specifically - Gemma-4-E2B, Gemma-4-E4B and Gemma-4-26B-A4B. |
Good point... never running on a even smaller model. Let me collect those data cc @am17an |

AI Usage:
Performance Results
Performance (A100, nrows=512, test-backend-ops perf, 5-run avg):
[512,512]: 427 -> 624 GB/s (+46%)
[768,512]: 626 -> 850 GB/s (+36%)
[1024,512]: 495 -> 645 GB/s (+30%)
[2048,512]: 911 -> 1171 GB/s (+28%)
[3072,512]: 1220 -> 1490 GB/s (+22%)
[5120,512]: 1668 -> 1815 GB/s (+9%)
Scalar fallback (4097,512): 1476 -> 1471 GB/s (no regression)
RTX 4070 Laptop GPU (5-run avg, test-backend-ops perf -o RMS_NORM -b CUDA0):
[512,512]: 474 -> 608 GB/s (+28%)
[768,512]: 580 -> 743 GB/s (+28%)
[1024,512]: 254 -> 313 GB/s (+23%)
[2048,512]: 436 -> 551 GB/s (+26%)
[3072,512]: 555 -> 672 GB/s (+21%)
[5120,512]: 696 -> 811 GB/s (+17%)
[8192,512]: 741 -> 828 GB/s (+12%)
Scalar fallback (4097,512): 678 -> 669 GB/s (no meaningful regression)
Why a separate kernel instead of a runtime branch
A runtime
if (use_vec4)inside the existingrms_norm_f32kernel was tested andfound to cause significant performance regression, even on the scalar fallback path.
Three versions were benchmarked (A100, nrows=512, test-backend-ops perf, 5-run avg):
The runtime branch version regresses at large dimensions -- [4096,512] drops to
0.74x of baseline on the vec4 path, and the scalar fallback [4097,512] drops to
0.79x even though it runs identical code to the baseline. This is because the
CUDA compiler must generate code for both paths within the same function, leading
to suboptimal register allocation and/or instruction cache utilization regardless
of which path is actually taken at runtime.
The separate kernel approach eliminates this entirely -- the compiler optimizes
each kernel independently, resulting in no regression on the scalar path (1.00x)
and consistent speedups on the vec4 path (+22% to +47% at medium dimensions).
Correctness:
RMS_NORM 17/17,
RMS_NORM_MUL_ADD 30/30,
ADD_RMS_NORM 25/25,
RMS_NORM_MUL_ROPE 72/72 passed.
ctest
cd build && ctest --output-on-failure

Those failures are NOT related to the op perf changes
End-to-end model benchmark (optional, requires GGUF model) (A100, RTX 4070)
(Looks like RMSNorm is pretty minior in those real model :) )
./bin/llama-bench -m /models/qwen2.5-1.5b-instruct-q4_k_m.gguf -t 1 -ngl 999 -r 5
./bin/llama-bench -m /models/Meta-Llama-3-8B-Instruct-Q4_K_M.gguf -t 1 -ngl 999 -r 5
Command
Build Commands
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --target test-backend-ops -j$(nproc)
cmake --build build --target llama-bench -j$(nproc)
cd build
Correctness tests (A100, RTX 4070)
./bin/test-backend-ops test -o RMS_NORM -b CUDA0
./bin/test-backend-ops test -o RMS_NORM_MUL_ADD -b CUDA0
./bin/test-backend-ops test -o ADD_RMS_NORM -b CUDA0
./bin/test-backend-ops test -o RMS_NORM_MUL_ROPE -b CUDA0
Performance benchmark (isolated RMS_NORM op) (A100, RTX 4070)
./bin/test-backend-ops perf -o RMS_NORM -b CUDA0