Repository navigation
CUDA: fix MMQ memory fault if n_expert >> n_ubatch - #29941
Merged
JohannesGaessler merged 1 commit intoOct 4, 2026
Merged
JohannesGaessler merged 1 commit into
JohannesGaessler merged 1 commit into
Conversation
am17an
approved these changes
Oct 4, 2026
Contributor
Author
IMbackK
approved these changes
Oct 4, 2026
pwilkin
approved these changes
Oct 4, 2026
1 task
5 tasks
Wizard815
pushed a commit
to Wizard815/mx-llama.cpp-Rocm10
that referenced
this pull request
Oct 6, 2026
(cherry picked from commit dd26678)
Wizard815
added a commit
to Wizard815/mx-llama.cpp-Rocm10
that referenced
this pull request
Oct 6, 2026
Picking dd26678 brought in upstream's launch block as well as its one-line size fix, but the fork declares ids_src1/ids_dst/expert_bounds further down, so the block referenced them before their declarations: mmq.cu:251: error: use of undeclared identifier 'ids_src1' The fork already has its own launch inside the chunked workspace path, so the inserted copy was both a duplicate and out of order. Remove it. Upstream's actual change (J_max sized by ne12 rather than ne11) stays. Assisted-by: Hermes Agent
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
My attention was drawn to these failing test cases by @am17an :
With the CUDA backend they result in an illegal memory access. The issue seems to be a bug in determining how much padding is needed for the temporary q8_1 buffer in the MMQ kernel. It's supposed to be as many extra columns as is the maximum tile width
J. For dense models this is determined bounded bysrc1->ne[1]. For MoE models however the tensor layout is different and it is instead bounded bysrc1->ne[2]. So the MMQ code is using the wrong value which can result in OOB memory accesses if the number of experts is much larger than the physical batch size. And the fix is simply to use the correct value instead.I decided against adding test cases to
test-backend-ops.cppbecause the tensors needed for a reproduction are relatively large and the conditions for this defect to manifest as a bug were highly specific.Requirements