Skip to content

meta: clear inactive AllReduce shards with FILL, not SCALE - #29793

Merged
ggerganov merged 1 commit into
masterfrom
fix/meta-allreduce-nan
Oct 1, 2026
Merged

ggerganov merged 1 commit into
masterfrom
fix/meta-allreduce-nan

Conversation

@ggerganov

Copy link
Copy Markdown
Member

Overview

Fixes NaN logits in tensor-parallel inference with more than 2 devices: the meta-backend butterfly AllReduce fallback zeroed the shard of a backend whose slice is empty (GGML_TENSOR_FLAG_COMPUTE cleared, tensor never written) by computing GGML_OP_SCALE with 0.0f, but that uninitialized memory can contain Inf/NaN. Use GGML_OP_FILL instead.

Additional information

Repro: GGML_CUDA_DEVICES=4 ./bin/test-llama-archs -a qwen4exp -v 4 -> nan logit

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. pi:llama.cpp/MiMo-V2.6-Flash-MOPD

@github-actions github-actions Bot added the ggml changes relating to the ggml tensor library for machine learning label Oct 1, 2026
The butterfly allreduce fallback zeroed the data of a backend whose slice
was empty (GGML_TENSOR_FLAG_COMPUTE cleared, so the tensor was never
written) by scaling it with 0.0f. The buffer still held uninitialized
memory, and per IEEE-754 0.0f * inf == nan and 0.0f * nan == nan, so the
garbage became NaNs which the reduction then summed into every backend's
result - NaN logits for the whole graph.

This showed up with >2 devices, e.g. GGML_CUDA_DEVICES=4 on qwen4exp:
4-way tensor split granularity leaves some devices with empty slices, and
both the NCCL and internal CUDA AllReduce paths (used with 2 devices)
zero such shards with an actual memset instead.

GGML_OP_FILL writes 0.0f to every element without reading the old
contents and runs on the backend's own stream, keeping it ordered after
the subgraph compute and before the reduction. ggml_fill currently
supports F32/F16 only.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
@ggerganov
ggerganov force-pushed the fix/meta-allreduce-nan branch from 55e6291 to 97c34e5 Compare October 1, 2026 09:10

@JohannesGaessler JohannesGaessler left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you, this is definitely a way better solution. I think to remember though that I did it like this because in some other part of the code memory was zeroed like this and I had assumed there was (as of yet) no better solution. So it may make sense to get a clanker to check the codebase for more instances of this.

@ggerganov
ggerganov merged commit 5503b04 into master Oct 1, 2026
31 of 32 checks passed
@ggerganov
ggerganov deleted the fix/meta-allreduce-nan branch October 1, 2026 10:42
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants