Skip to content

CUDA: Fix data-races when reusing SMEM in block_reduce - #26385

Merged
ORippler merged 7 commits into
ggml-org:masterfrom
ORippler:osimons/fix_race_block_reduce
Aug 3, 2026
Merged

ORippler merged 7 commits into
ggml-org:masterfrom
ORippler:osimons/fix_race_block_reduce

Conversation

@ORippler

@ORippler ORippler commented Jul 31, 2026 •

Copy link
Copy Markdown
Collaborator

Overview

block_reduce introduced in #18785 introduces data-races when SMEM is reused across block_reduce calls. This PR fixes it by double-buffering/adding __syncthreads() as needed. We may look to alternatively always synchronize to make block_reduce inherently safe/self-contained, but that is not needed when not reusing SMEM.

Additional information

Failing softmax tests when using `compute-sanitizer --tool=racecheck` before this PR
(base) osimons@ub-osimons:~/llama.cpp$ compute-sanitizer --tool=racecheck ./build-x64-linux-gcc-reldbg/bin/test-backend-ops -o SOFT_MAX
========= COMPUTE-SANITIZER
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97250 MiB):
  Device 0: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97250 MiB
Testing 2 devices

Backend 1/2: CUDA0
  Device description: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
  Device memory: 97250 MB (96593 MB free)

  SOFT_MAX(type=f32,ne=[16,16,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
  SOFT_MAX(type=f32,ne=[15,15,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
  SOFT_MAX(type=f32,ne=[16,1024,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
  SOFT_MAX(type=f32,ne=[15,1023,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
========= Error: Race reported between Read access at T3 block_reduce<(block_reduce_method)0, (unsigned int)1024, float>(T3, T3 *)+0x940 in common.cuh:643
=========     and Write access at T3 block_reduce<(block_reduce_method)1, (unsigned int)1024, float>(T3, T3 *)+0xaf0 in common.cuh:638 [3516 hazards]
=========
[SOFT_MAX] ERR = 0.850562643 > 0.000001000   SOFT_MAX(type=f32,ne=[1024,16,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): FAIL
========= Error: Race reported between Read access at T3 block_reduce<(block_reduce_method)0, (unsigned int)0, float>(T3, T3 *)+0xd00 in common.cuh:643
=========     and Write access at T3 block_reduce<(block_reduce_method)1, (unsigned int)0, float>(T3, T3 *)+0xff0 in common.cuh:638 [3248 hazards]
=========
[SOFT_MAX] ERR = 0.861961235 > 0.000001000   SOFT_MAX(type=f32,ne=[1023,15,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): FAIL
Passing softmax tests when using `compute-sanitizer --tool=racecheck` after this PR

(base) osimons@ub-osimons:~/llama.cpp$ compute-sanitizer --tool=racecheck ./build-x64-linux-gcc-reldbg/bin/test-backend-ops -o SOFT_MAX
========= COMPUTE-SANITIZER
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97250 MiB):
Device 0: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97250 MiB
Testing 2 devices

Backend 1/2: CUDA0
Device description: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Device memory: 97250 MB (96593 MB free)

SOFT_MAX(type=f32,ne=[16,16,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[15,15,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[16,1024,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[15,1023,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[1024,16,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[1023,15,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[1024,1024,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[1023,1023,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=1.000000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[16,16,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=0.100000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[15,15,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=0.100000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[16,1024,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=0.100000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[15,1023,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=0.100000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[1024,16,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=0.100000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[1023,15,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=0.100000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[1024,1024,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=0.100000,max_bias=0.000000,inplace=0): OK
SOFT_MAX(type=f32,ne=[1023,1023,1,1],mask=0,sinks=0,m_prec=f32,nr23=[1,1],scale=0.100000,max_bias=0.000000,inplace=0): OK

Requirements

ORippler added 3 commits July 31, 2026 22:06
block_reduce currently doesn't resync after reading from SMEM, causing
potential data-races when reusing SMEM for multiple reductions.

One may consider simply always adding this in block_reduce, but this
comes at a potential perf cost
@ORippler
ORippler requested a review from a team as a code owner July 31, 2026 20:25
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Jul 31, 2026
@Kononnable

Copy link
Copy Markdown
Contributor

@ORippler I can confirm that with this changes alone I get deterministic results from llama-completion runs.
It seems that the full stream sync on virtual device I proposed in #26344 just lowered the chances for negative effects of race condition fixed by this PR.

@ORippler ORippler added merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. and removed merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. labels Aug 3, 2026
@ORippler ORippler added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Aug 3, 2026
Comment thread ggml/src/ggml-cuda/softmax.cu Outdated
@ORippler
ORippler requested a review from gaugarg-nv August 3, 2026 09:45
@ORippler
ORippler merged commit 9bd4c09 into ggml-org:master Aug 3, 2026
19 of 22 checks passed
@ORippler
ORippler deleted the osimons/fix_race_block_reduce branch August 3, 2026 12:22
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
* CUDA: Fix data-races when reusing block_reduce

block_reduce currently doesn't resync after reading from SMEM, causing
potential data-races when reusing SMEM for multiple reductions.

One may consider simply always adding this in block_reduce, but this
comes at a potential perf cost

* double-buffering for single-row softmax

* double-buffering for norm as well

* Add comment

* Add explanatory comment to block_reduce

* Specify need for + do memory barrier only in multi-warp scenario

* Implement review-suggestion from @gaugarg-nv
brittlewis12 pushed a commit to brittlewis12/llama.cpp that referenced this pull request Aug 17, 2026
* CUDA: Fix data-races when reusing block_reduce

block_reduce currently doesn't resync after reading from SMEM, causing
potential data-races when reusing SMEM for multiple reductions.

One may consider simply always adding this in block_reduce, but this
comes at a potential perf cost

* double-buffering for single-row softmax

* double-buffering for norm as well

* Add comment

* Add explanatory comment to block_reduce

* Specify need for + do memory barrier only in multi-warp scenario

* Implement review-suggestion from @gaugarg-nv
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
* CUDA: Fix data-races when reusing block_reduce

block_reduce currently doesn't resync after reading from SMEM, causing
potential data-races when reusing SMEM for multiple reductions.

One may consider simply always adding this in block_reduce, but this
comes at a potential perf cost

* double-buffering for single-row softmax

* double-buffering for norm as well

* Add comment

* Add explanatory comment to block_reduce

* Specify need for + do memory barrier only in multi-warp scenario

* Implement review-suggestion from @gaugarg-nv
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* CUDA: Fix data-races when reusing block_reduce

block_reduce currently doesn't resync after reading from SMEM, causing
potential data-races when reusing SMEM for multiple reductions.

One may consider simply always adding this in block_reduce, but this
comes at a potential perf cost

* double-buffering for single-row softmax

* double-buffering for norm as well

* Add comment

* Add explanatory comment to block_reduce

* Specify need for + do memory barrier only in multi-warp scenario

* Implement review-suggestion from @gaugarg-nv
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* CUDA: Fix data-races when reusing block_reduce

block_reduce currently doesn't resync after reading from SMEM, causing
potential data-races when reusing SMEM for multiple reductions.

One may consider simply always adding this in block_reduce, but this
comes at a potential perf cost

* double-buffering for single-row softmax

* double-buffering for norm as well

* Add comment

* Add explanatory comment to block_reduce

* Specify need for + do memory barrier only in multi-warp scenario

* Implement review-suggestion from @gaugarg-nv
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* CUDA: Fix data-races when reusing block_reduce

block_reduce currently doesn't resync after reading from SMEM, causing
potential data-races when reusing SMEM for multiple reductions.

One may consider simply always adding this in block_reduce, but this
comes at a potential perf cost

* double-buffering for single-row softmax

* double-buffering for norm as well

* Add comment

* Add explanatory comment to block_reduce

* Specify need for + do memory barrier only in multi-warp scenario

* Implement review-suggestion from @gaugarg-nv
0xShug0 pushed a commit to 0xShug0/audio.cpp that referenced this pull request Oct 9, 2026
…cpp #26385) (#837)

block_reduce has each warp write its partial result to shared memory,
syncs once and reads the partials back, with no barrier after the
read. When a kernel reduces twice through one buffer, a warp that is
ahead can write its second partial while another warp still reads the
first. soft_max_f32 reduces the row maximum and then the sum of
exponentials that way, so a slow warp can take a partial sum for the
maximum and its part of the row comes out wrong. group_norm_f32 and
the cooperative single-row softmax also reduce twice through one
buffer.

It showed with LFM2-Audio: two sessions in one process sometimes gave
different bytes for the same request, while each alone was bit-exact.
Their kernels share the GPU, the warps of a block drift apart, and the
encoder's attention softmax runs over rows of more than 32 positions,
which take several warps. compute-sanitizer racecheck reports the WAR
hazards in soft_max_f32 on main with a single session too.

This is upstream's fix, from ggml-org/llama.cpp#26385: soft_max_f32
syncs between its two reductions when the block has more than one
warp, group_norm_f32 and the cooperative softmax give their second
reduction a buffer of its own, and block_reduce states that callers
must not reuse the buffer until every read is done. The upstream diff
applies unchanged; only line offsets and the vendored files' CRLF line
endings differ. No other block_reduce caller in ggml-cuda reduces
twice through one buffer.

Checked on an A10 with CUDA 12.8, against main:
- racecheck on the LFM2-Audio encoder over a 99-position question:
  484704 errors in soft_max_f32 on main, none here. Over a single-turn
  S2S request cut at 60 tokens, every kernel checked: 1303436 errors on
  main, none here
- test-backend-ops SOFT_MAX, GROUP_NORM, NORM, RMS_NORM, L2_NORM,
  SUM_ROWS and MEAN pass on both. Under racecheck's timing main fails
  83 of 212 SOFT_MAX cases, the cooperative ones included, and both
  GROUP_NORM cases; this commit passes them all with no hazard
- one session per process, EN and JP: ASR offline, TTS and S2S offline
  and streamed (36 requests): all 82 output files byte-identical to
  main, in two runs each
- under memcheck's timing, main differs in 9 of 9 encoder repeats
  and in 5 of 5 single-turn S2S requests; this commit in none
- two C API sessions in one process, running multi-turn S2S from a
  branch not yet merged: with this commit 0 of 159 turns differ from
  each session alone; without any fix 1 of 162 did, too rare on an idle
  machine to show the fix by itself
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants