Repository navigation
Normalize biases before encoding in gather_qmm_rhs - #4056
Conversation
zcbenz
left a comment
There was a problem hiding this comment.
Thanks for the fix, I can verify the bug and the fix, but it is strange that gather_qmm_rhs actually uses ensure_row_contiguous inside it:
mlx/mlx/backend/metal/quantized.cpp
Lines 1612 to 1613 in f599c02
Do you have an idea what might go wrong?
|
Thank you for pointing that out — it led to the actual root cause. The biases Updated the PR: the biases normalization is hoisted next to |
Proposed changes
Fixes #4055.
GatherQMM::eval_gpunormalizesw,scalesandbiaseswithensure_row_contiguous_matrix, which validates only the last two axes. A rowslice of an
[E, 2R, D]expert weight passes that check while keeping theleading stride of the original array, so the sorted path reads every expert
after the first from the wrong offset — silently, with no error.
The gather kernels index the leading axes themselves, so this switches those
three inputs to
ensure_row_contiguous, which already exists in the file andis already used for
indices.xis left as is; the failure I can demonstrateis in the weight path.
Verified on M5 Pro / macOS 26.5.1: the new test fails before the change and
passes after, and
Ran 794 tests ... OK (skipped=46)for the full suite.Checklist
Put an
xin the boxes that apply.pre-commit run --all-filesto format my code / installed pre-commit prior to committing changes