Skip to content

🚨 fix(ggml-cuda): backport the block_reduce shared-memory race fix (llama.cpp #26385) - #837

Merged
0xShug0 merged 1 commit into
0xShug0:mainfrom
Liquid4All:ggml-cuda-reduce-race
Oct 9, 2026
Merged

0xShug0 merged 1 commit into
0xShug0:mainfrom
Liquid4All:ggml-cuda-reduce-race

Conversation

@ykhrustalev

Copy link
Copy Markdown
Contributor

Backport of ggml-org/llama.cpp#26385, which fixes a shared-memory race in ggml-cuda kernels that call block_reduce twice.

Problem

  • block_reduce writes each warp's partial to shared memory, syncs once and reads the partials back, with no barrier after the read. A kernel that reduces twice through one buffer lets a warp that is ahead overwrite partials another warp is still reading
  • soft_max_f32 reduces the row maximum and then the sum that way, so a slow warp can take a partial sum for the maximum. group_norm_f32 and the cooperative single-row softmax do the same
  • It showed with lfm2_audio: two sessions in one process sometimes gave different bytes for the same request, while each alone was bit-exact. The encoder's attention softmax has rows over 32 positions, which take several warps. compute-sanitizer racecheck reports the hazards on main with one session too
  • Not specific to lfm2_audio: on the CUDA backend, and in the HIP build of the same sources, any model with a softmax row over 32 or a group norm over groups of 1024 or more elements is exposed

What this PR changes

  • Upstream's diff, unchanged: soft_max_f32 syncs between its two reductions when the block has more than one warp; group_norm_f32 gives its second reduction its own 32 floats; the cooperative softmax gets separate buffers for its maximum and its sum; block_reduce gains a comment stating the contract
  • Adaptations: none to the code. The common.cuh and norm.cu hunks apply at other line offsets (norm.cu carries audio.cpp's channel RMS norm kernels), and the vendored files keep their CRLF line endings
  • The other block_reduce callers (norm, rms_norm, the channel RMS norm, l2_norm, sum_rows and mean) reduce once per kernel, and convrot-linear's own reduction ends with a barrier, so none needs a change

Testing
A10, CUDA 12.8, main (75d0294) against this branch (the HIP build, which compiles the same sources, was not built or run):

  • racecheck, encoder over a 99-position JP question: main 484704 errors, all in soft_max_f32; here none. A single-turn S2S request cut at 60 tokens, every kernel checked: main 1303436 errors; here none
  • ggml's test-backend-ops, built from the vendored source against the same ggml (the project does not build it): SOFT_MAX, GROUP_NORM, NORM, RMS_NORM, L2_NORM, SUM_ROWS and MEAN pass on both. Under racecheck's timing main fails 83 of 212 SOFT_MAX cases (the three cooperative ones among them) and both GROUP_NORM cases; this branch passes all, with no hazard
  • The four CUDA unit tests, lfm2_audio_cuda_test among them, pass on both
  • One session per process, EN and JP F16: ASR offline, seeded TTS and S2S offline and streamed (36 requests per run): all 82 output files byte-identical to main, in two runs each
  • Under memcheck's timing, main's encoder output differs in 9 of 9 repeats and a single-turn S2S reply cut at 60 tokens in 5 of 5; this branch's in none
  • Two C API sessions in one process, running multi-turn S2S from a branch not yet merged: with this commit 0 of 159 turns differ from the session alone; without any fix 1 of 162 differed, too rare on an idle machine to show the fix by itself. racecheck over a turn with history, limited to the softmax, norm and row-reduction kernels: 466172 errors without any fix, none with this commit

…p #26385)

block_reduce has each warp write its partial result to shared memory,
syncs once and reads the partials back, with no barrier after the
read. When a kernel reduces twice through one buffer, a warp that is
ahead can write its second partial while another warp still reads the
first. soft_max_f32 reduces the row maximum and then the sum of
exponentials that way, so a slow warp can take a partial sum for the
maximum and its part of the row comes out wrong. group_norm_f32 and
the cooperative single-row softmax also reduce twice through one
buffer.

It showed with LFM2-Audio: two sessions in one process sometimes gave
different bytes for the same request, while each alone was bit-exact.
Their kernels share the GPU, the warps of a block drift apart, and the
encoder's attention softmax runs over rows of more than 32 positions,
which take several warps. compute-sanitizer racecheck reports the WAR
hazards in soft_max_f32 on main with a single session too.

This is upstream's fix, from ggml-org/llama.cpp#26385: soft_max_f32
syncs between its two reductions when the block has more than one
warp, group_norm_f32 and the cooperative softmax give their second
reduction a buffer of its own, and block_reduce states that callers
must not reuse the buffer until every read is done. The upstream diff
applies unchanged; only line offsets and the vendored files' CRLF line
endings differ. No other block_reduce caller in ggml-cuda reduces
twice through one buffer.

Checked on an A10 with CUDA 12.8, against main:
- racecheck on the LFM2-Audio encoder over a 99-position question:
  484704 errors in soft_max_f32 on main, none here. Over a single-turn
  S2S request cut at 60 tokens, every kernel checked: 1303436 errors on
  main, none here
- test-backend-ops SOFT_MAX, GROUP_NORM, NORM, RMS_NORM, L2_NORM,
  SUM_ROWS and MEAN pass on both. Under racecheck's timing main fails
  83 of 212 SOFT_MAX cases, the cooperative ones included, and both
  GROUP_NORM cases; this commit passes them all with no hazard
- one session per process, EN and JP: ASR offline, TTS and S2S offline
  and streamed (36 requests): all 82 output files byte-identical to
  main, in two runs each
- under memcheck's timing, main differs in 9 of 9 encoder repeats
  and in 5 of 5 single-turn S2S requests; this commit in none
- two C API sessions in one process, running multi-turn S2S from a
  branch not yet merged: with this commit 0 of 159 turns differ from
  each session alone; without any fix 1 of 162 did, too rare on an idle
  machine to show the fix by itself
@0xShug0

0xShug0 commented Oct 8, 2026

Copy link
Copy Markdown
Owner

@ykhrustalev This may take a little longer. For ggml changes, I need to understand the problem, check the solution, and run regression tests. Sometimes there could be a simpler solution with low regression risk (e.g., #542)

@0xShug0 0xShug0 changed the title fix(ggml-cuda): backport the block_reduce shared-memory race fix (llama.cpp #26385) 🚨 fix(ggml-cuda): backport the block_reduce shared-memory race fix (llama.cpp #26385) Oct 9, 2026
@0xShug0

0xShug0 commented Oct 9, 2026

Copy link
Copy Markdown
Owner

@ykhrustalev Thanks, PR merged! No observable performance impact in my tests. I was initially concerned about flash attention, but it bypasses the modified standalone softmax kernel because softmax is fused into the attention kernel.

@0xShug0
0xShug0 merged commit de13064 into 0xShug0:main Oct 9, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants