Repository navigation
Conversation
Add a new operator GGML_OP_MOE_SUM that efficiently aggregates outputs from multiple experts in MoE models by summing along the expert dimension. Input format: [hidden_dim, n_expert_used, n_tokens] Output format: [hidden_dim, n_tokens] CPU implementation: - Optimized cache-friendly loop order (expert -> token -> hidden_dim) - Multi-threaded parallelization across tokens - Specialized F32 implementation for better performance - 1.28x faster than naive add_loop approach CUDA implementation: - Warp-per-token kernels for large token counts - Specialized F16 vectorized kernel for large batches - Small-token kernels for edge cases - 1.50x faster than naive add_loop approach Tests: - 96 test cases covering F32/F16, various expert counts (2,4,8), hidden dimensions (64-4096), and token counts (16-256) - Relaxed error threshold for F16 (1e-6 vs 1e-7 for F32) due to limited precision when summing multiple expert outputs
Replace the loop of ggml_add operations with ggml_moe_sum when the experts tensor is contiguous. This is more efficient, especially for GPU kernels. - Fast path: Use ggml_moe_sum for contiguous tensors with n_expert_used > 1 - Fallback: Keep the ggml_add loop for non-contiguous tensors or single expert
|
Repeating adds are already fused on most backends, including CUDA: llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu Lines 3624 to 3653 in 3795cc1 |
|
Also this was attempted in before #16857 and it leads to various problems I haven't diagnosed yet. In any case, this should be a fusion optimization not a full fledged ggml operator. |
|
Thank you both @CISC and @am17an for taking the time to review and for your helpful feedback! Also, regarding @am17an your comment that "this should be a fusion optimization not a full fledged ggml operator" - I'm still learning the codebase architecture and would really appreciate your guidance on this. What exactly makes a "fusion optimization" different from a "full fledged ggml operator" in the llama.cpp design philosophy? And why would the fusion approach be preferred here? I want to make sure I understand the proper way to contribute to this project. |
|
IMO calling it "moe_sum" is misleading because it has nothing to do with MoE. Indeed, this is just the equivalent of pytorch sum() with Probably it's more useful to have an operator |
This allows disabling the CUDA implementation of ggml_moe_sum to compare performance with ggml_cuda_op_fused_add. When GGML_DISABLE_MOE_SUM_CUDA is defined: - moesum.cu becomes empty (no CUDA kernel) - ggml_moe_sum falls back to CPU implementation - Setting LLAMA_DISABLE_MOE_SUM=1 will use ggml_add loop which triggers ggml_cuda_op_fused_add Usage for comparison: - ggml_moe_sum (CUDA): default (both flags unset) - ggml_cuda_op_fused_add: -DGGML_DISABLE_MOE_SUM_CUDA=1 -DLLAMA_DISABLE_MOE_SUM=1
68d480a to
d176ae1
Compare
Add a new GGML_OP_MOE_SUM operator that efficiently aggregates outputs
from multiple experts in Mixture of Experts (MoE) models.
Input: [hidden_dim, n_expert_used, n_tokens]
Output: [hidden_dim, n_tokens]
Performance
Benchmark on Qwen3-30B-A3B (Q4_K_M) with NVIDIA A40 (At the moment I only have an A40 graphics card):
Implementation