Skip to content

ggml: add moe_sum operator for efficient MoE expert aggregation - #19362

Closed
blime4 wants to merge 6 commits into
ggml-org:masterfrom
blime4:moe-sum-upstream
Closed

blime4 wants to merge 6 commits into
ggml-org:masterfrom
blime4:moe-sum-upstream

Conversation

@blime4

@blime4 blime4 commented Feb 5, 2026 •

Copy link
Copy Markdown

Add a new GGML_OP_MOE_SUM operator that efficiently aggregates outputs
from multiple experts in Mixture of Experts (MoE) models.

Input: [hidden_dim, n_expert_used, n_tokens]
Output: [hidden_dim, n_tokens]

Performance

Benchmark on Qwen3-30B-A3B (Q4_K_M) with NVIDIA A40 (At the moment I only have an A40 graphics card):

Test ggml_moe_sum ggml_add loop Improvement
pp256 145.42 t/s 143.90 t/s +1.06%
tg512 137.81 t/s 135.85 t/s +1.44%

Implementation

  • CPU: Optimized cache-friendly loop order
  • CUDA: Specialized kernels for different workloads
  • Automatic fallback to ggml_add loop for non-contiguous tensors
  • Runtime disable: LLAMA_DISABLE_MOE_SUM=1

shaobo.xie added 5 commits February 5, 2026 15:34
Add a new operator GGML_OP_MOE_SUM that efficiently aggregates outputs
from multiple experts in MoE models by summing along the expert dimension.

Input format: [hidden_dim, n_expert_used, n_tokens]
Output format: [hidden_dim, n_tokens]

CPU implementation:
- Optimized cache-friendly loop order (expert -> token -> hidden_dim)
- Multi-threaded parallelization across tokens
- Specialized F32 implementation for better performance
- 1.28x faster than naive add_loop approach

CUDA implementation:
- Warp-per-token kernels for large token counts
- Specialized F16 vectorized kernel for large batches
- Small-token kernels for edge cases
- 1.50x faster than naive add_loop approach

Tests:
- 96 test cases covering F32/F16, various expert counts (2,4,8),
  hidden dimensions (64-4096), and token counts (16-256)
- Relaxed error threshold for F16 (1e-6 vs 1e-7 for F32) due to
  limited precision when summing multiple expert outputs
Replace the loop of ggml_add operations with ggml_moe_sum when the experts
tensor is contiguous. This is more efficient, especially for GPU kernels.

- Fast path: Use ggml_moe_sum for contiguous tensors with n_expert_used > 1
- Fallback: Keep the ggml_add loop for non-contiguous tensors or single expert
@blime4
blime4 requested review from CISC and ggerganov as code owners February 5, 2026 13:15
@CISC

CISC commented Feb 5, 2026

Copy link
Copy Markdown
Member

Repeating adds are already fused on most backends, including CUDA:

if (node->op == GGML_OP_ADD) {
int n_fuse = 0;
ggml_op ops[8];
std::fill(ops, ops + 8, GGML_OP_ADD);
for (; n_fuse <= 6; ++n_fuse){
if (!ggml_can_fuse(cgraph, i + n_fuse, ops + n_fuse, 2)) {
break;
}
if (cgraph->nodes[i + n_fuse] != cgraph->nodes[i + n_fuse + 1]->src[0]) {
break;
}
if (!ggml_are_same_layout(cgraph->nodes[i + n_fuse]->src[1], cgraph->nodes[i + n_fuse + 1]->src[1])) {
break;
}
}
n_fuse++;
if (n_fuse > 1) {
for (int j = 0; j < n_fuse - 1; ++j) {
node->src[j + 2] = cgraph->nodes[i + j + 1]->src[1];
}
cgraph->nodes[i + n_fuse - 1]->data = node->data;
ggml_cuda_op_fused_add(*cuda_ctx, node, n_fuse);
i += n_fuse - 1;
continue;
}
}

@am17an

am17an commented Feb 5, 2026

Copy link
Copy Markdown
Contributor

Also this was attempted in before #16857 and it leads to various problems I haven't diagnosed yet. In any case, this should be a fusion optimization not a full fledged ggml operator.

@github-actions github-actions Bot added documentation Improvements or additions to documentation testing Everything test related Nvidia GPU Issues specific to Nvidia GPUs ggml changes relating to the ggml tensor library for machine learning labels Feb 5, 2026
@blime4

blime4 commented Feb 5, 2026

Copy link
Copy Markdown
Author

Thank you both @CISC and @am17an for taking the time to review and for your helpful feedback!

Also, regarding @am17an your comment that "this should be a fusion optimization not a full fledged ggml operator" - I'm still learning the codebase architecture and would really appreciate your guidance on this. What exactly makes a "fusion optimization" different from a "full fledged ggml operator" in the llama.cpp design philosophy? And why would the fusion approach be preferred here? I want to make sure I understand the proper way to contribute to this project.

@ngxson

ngxson commented Feb 5, 2026 •

Copy link
Copy Markdown
Collaborator

IMO calling it "moe_sum" is misleading because it has nothing to do with MoE. Indeed, this is just the equivalent of pytorch sum() with dim=1: [x, y, z] -> [x, 1, z]

Probably it's more useful to have an operator sum_dim, or better, extend the current sum_rows to accept custom dim

This allows disabling the CUDA implementation of ggml_moe_sum to
compare performance with ggml_cuda_op_fused_add.

When GGML_DISABLE_MOE_SUM_CUDA is defined:
- moesum.cu becomes empty (no CUDA kernel)
- ggml_moe_sum falls back to CPU implementation
- Setting LLAMA_DISABLE_MOE_SUM=1 will use ggml_add loop
  which triggers ggml_cuda_op_fused_add

Usage for comparison:
- ggml_moe_sum (CUDA): default (both flags unset)
- ggml_cuda_op_fused_add: -DGGML_DISABLE_MOE_SUM_CUDA=1 -DLLAMA_DISABLE_MOE_SUM=1
@blime4 blime4 closed this Feb 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning Nvidia GPU Issues specific to Nvidia GPUs testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants