Repository navigation
docs(mmq): the MMQ padding must cover the padded y tile — 903 re-cut; short on CDNA3, gfx1151 and RDNA4, not NVIDIA - #448
Merged
Conversation
…t the fix Upstream merged llama.cpp#29941 (dd266785c) and closed #27044 for it. On master dd266785c as shipped, with only new test cases added, MUL_MAT_ID(q4_K, 256 experts, 10 used, 576 rows, 120 tokens, k=1536) aborts with an illegal memory access in 10 of 10 runs on sm_120: the stock pool, no sanitizer, no debug code. The shape was found with a model of the VMM pool and the stream-k fixup buffer. The fixup buffer is why #445's short-batch shapes never faulted, and why #29847's shape at 100 tokens shows no compute-sanitizer errors upstream. A CPU check over 4.7 million shapes counts every rule's shortfall: #29941 leaves about 304,000 shapes short by up to 63 blocks. #27044's line (compat 903) is short with one expert per token and in 192 gfx1151 shapes, and pads no ids_dst. Padding for the widest tile, as before #24127, covers all of them: mmq-successor-fix.patch, with ids_dst and the NVFP4 scales padded the same way. This is raw material. llama.cpp prohibits AI-written posts, so the upstream PR text is the maintainer's to write. The GPU driver's end-to-end table follows in a later commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The maintainer's decision (2026-10-05): 903 carries the successor fix instead of llama.cpp#27044's get_J_max(ne12*n_expert_used). That line covered the broadcast fault it was written for. But it was short with one expert per token, and on gfx1151 for 2-3 experts at 5-7 tokens, and it padded no ids_dst. Upstream's #29941 (get_J_max(ne12)) is short below 128 tokens, and master dd266785c aborts on a test case. 903 now pads src1_q8_1 in both branches, ids_dst, and NVFP4's src1_scale for the widest tile that has a config, get_J_max(type, fallback, cc, 512): 128 on NVIDIA, as get_mmq_x_max_host() padded before #24127. The patch is cut against b11081, whose ids branch still reads ne11. The full series applies to a clean b11081. The compat README entry, the retirement register row, the README rows, the submission material's outcome, the investigation doc's header and the successor material follow. Production's MoE GGUFs use 6 or 8 experts per token, so neither of #27044's src1 gaps touched them, and the ids_dst read landed in the pool's next allocation. The change matters for one-expert models, for gfx1151's small batches, and wherever a buffer ends at the edge of the pool's mapping, as master's failing test case shows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
End-to-end on a served MoE model, the way the fault was first found. On
qwen3.6:35b-a3b-q4_K_M with one 3072x1728 image (2048-token ubatch),
sm_120, b11081:
raw ne11, VMM pool (shipping), num_ctx 8192 and 33792 -> HTTP 200 (pool masks the read)
raw ne11, exact-size src1 alloc (debug 910), num_ctx 8192 -> illegal memory access at
"decoding image batch 1/2", runner core-dumped
new 903 widest-tile, exact-size src1 alloc -> HTTP 200, decoded cleanly
The broadcast gate/up MUL_MAT_ID has ne11==1, so upstream's get_J_max(ne11)
is 0 and src1 gets no tail padding. The shipping VMM pool maps memory past
the buffer, so the read is latent (no crash) -- which is why it first
looked like an intermittent crash gated on num_ctx>32768. The debug 910
patch (env MMQ_EXACT, not shipped) gives src1 an exact allocation so the
read crosses the boundary deterministically; the raw kernel then crashes
and the widest-tile 903 does not.
compute-sanitizer cannot instrument the runner through ollama serve
(ollama re-execs with a rebuilt env that drops the injection), so the
precise memcheck line is from the test-backend-ops path; on the real model
the exact allocation turns the read into the crash above.
Logs: docs/maxusai/tasks/mmq-successor-results/real-model/.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…+ sub-128 Real-model check requested: does the merged #29941 (ne12) still fail on a large image? No. get_J_max(ne12) with ne12=2048 (the image ubatch) pads the widest tile, so #29941 fully covers the src1 broadcast read that crashes the raw ne11 kernel -- qwen3.6:35b-a3b decodes the image cleanly on #29941+exact. Its residual over-reads are ids_dst (never padded, any batch) and src1 below 128 tokens, both in the test-backend-ops exact matrix; neither hard-crashes the large-image path. The widest-tile 903 closes all of them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
#29953 removes the rounding inconsistency between the launched tile width and the padding, which is the right shape of fix. It is still short in two places, both reproduced on sm_120 without a sanitizer. 1. src1. A tile's y load is unbounded in l and copies the whole padded shared-memory y tile from global memory, GGML_PAD(J*sizeof(block_q8_1_mmq), nthreads*sizeof(int)) bytes from the tile's first column — the expression mmq_get_nbytes_shared() already uses. That is more than J blocks whenever nthreads*4 does not divide J*144: at J=16 it needs 21 blocks, at J=112 it needs 113. #29953 pads J_best blocks. 2. ids_dst, which no rule so far pads at all: a tile loads J entries from the expert's first row with no bound, so the last expert's last tile reads up to J-1 entries past ne12*n_expert_used. This also corrects our own #448: padding by the widest tile covers every NVIDIA shape checked but is short by up to 5 blocks in 25,384 gfx1151 shapes. New here: - tasks/mmq-amend-29953.patch, 8 lines on top of #29953, closing both; tasks/mmq-fix-amended.patch, the equivalent standalone change against master. - tasks/mmq-debug-alloc.cuh: MMQ_DEBUG_ALLOC=guard[:src1|:ids_dst] places a buffer at the top of its own VMM mapping with the next granule unmapped, so an over-read faults with no sanitizer; =exact keeps the earlier cudaMalloc mode. MMQ_DEBUG_PRINT=1 has the build report its own J, nthreads and padding. - tasks/mmq-rules-check.cu: the CPU check, now six rules and the corrected requirement. tasks/build-check.sh builds it with the gencode list that ggml_cuda_highest_compiled_arch() needs — without it the earlier counts modelled a machine that does not exist. - tasks/mmq-tests.patch adds ids16 and ids64; ids16 (q4_0, 256 experts, 16 used, 640x16x2560) is the reproduction: stock it passes under every rule, under guard:src1 it aborts for ne11, ne12 and #29953, under guard:ids_dst it aborts for all four and passes only with #448. - tasks/mmq-rules-gpu.sh and tasks/mmq-rules-table.py: the seven-rule matrix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…native The amendment writes GGML_PAD(J*B, nthreads*4) a second time in mmq.cu, which is the duplication that let the two sizes drift apart in the first place. Naming it once in mmq.cuh and having mmq_get_nbytes_shared() use it removes that. Flagged explicitly as not what the measurements were taken with. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A variant present in results.tsv but missing from VARIANT_LABEL was filtered out of every table with no warning, so the two amended rules did not appear in the matrix at all. It now refuses to render instead. Also registers s448p and p29953fix, and gives mmq-variant.py a hint for the missing p29953.patch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The widest-tile rule merged earlier today assumed a tile reads J blocks of src1. It reads GGML_PAD(J*sizeof(block_q8_1_mmq), nthreads*sizeof(int)) bytes: the y load is unbounded in l and copies the whole padded shared-memory y tile, the same size mmq_get_nbytes_shared() gives it. On sm_120 J=16 needs 21 blocks and J=112 needs 113; only J=128 happens to equal J-1, which is why the shapes that first exposed this bug were covered either way. So the plain widest-tile rule was short by up to 5 blocks in 25,384 gfx1151 shapes (q2_K and q3_K, where J=80 with nthreads=256 needs 85 against the widest tile's 80). The ROCm host serves gfx1151, so that is not hypothetical for the fork. 903 now takes the maximum padded tile over every config that exists, which leaves no shape short on any of the five architectures checked. Verified against the pin, b11081: - the CPU check built against b11081 gives counts identical to dd266785c for all six rules, so the config tables did not move between them (tasks/mmq-successor-results/check-rules-b11081.txt); - b11081's y load, ids_dst load and mmq_get_nbytes_shared are character-identical to master's, so the analysis transfers; - the full compat series applies to a clean b11081 worktree with the new 903, and the resulting mmq.cu compiles for sm_120. The check now calls ggml_cuda_mmq_get_config() in its four-argument form, which is b11081's only signature and master's with prec_src1 defaulted, so one source builds against either pin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Across sm_75/86/89/120 and gfx1151 x ten quantization types, the widest-tile and widest-padded-tile rules differ in exactly one combination: gfx1151, non-fallback, q2_K, where J_pad goes 80 -> 86 blocks. Everywhere else both give 128 blocks, so closing the 25,384 short gfx1151 shapes is free except for those 864 bytes per call. tasks/mmq-jpad-cost.cu prints the comparison. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… fault 12 cases x 7 padding rules x 4 modes on sm_120. Two columns come out clean: under a guard page on ids_dst all twelve cases abort for ne11, ne12, #29953 and #27044 and pass for the three rules that pad it, and under exact-size allocations with memcheck all four published rules report errors (43 to 11,939) against 0 for the three amended ones. src1 separates them: #29953 aborts on ids16 2/2 and on j100_b1 1/2, #27044 on one113 2/2, ne11 on all twelve, ne12 on eight. Correction to the earlier claim that j100 could not realise the worst case. What a run reads past the data is ceil(T/B) - r blocks with r the last non-empty expert's row count, so it needs r = 1, and that is set by the mean rows per expert: ids16 is 1.0 (every run), j100 is 2.0 and t100 3.1 (sometimes), ids64 is 25.0 (never, and it passes guard:src1 under every rule but ne11). I had put j100 at 25 along with ids64; it is 1000 rows over 512 experts, not 256. j100_b1's 1 of 2 is recorded as not a result pending repeats with the amended rule as a control in the same harness. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…at 749 rows #29953's shortfall at J=112 is one 144-byte block, which only bites when the routing puts a single row in the last non-empty expert, so one observation was not a result. Ten repeats per cell, with the amendment run in the same harness as the control, aborts out of ten under guard:src1: case mean rows/expert ne12 #29953 #29953+amendment ids16 1.0 10/10 10/10 0/10 j100_b1 2.0 10/10 4/10 0/10 j100_b0 2.0 10/10 1/10 0/10 So the one-block shortfall reproduces at 4 in 10 and 1 in 10, and the amendment is clean 10 of 10 on both, which is what separates the padding from the harness. ids16 stays the case to hand over: 5 blocks short and one row per expert, so it aborts every run. Also adds the combined guard mode (both buffers at once) for five cases: all four published rules abort on all five, all three amended rules pass. The two addenda ran from tasks/mmq-rules-addenda.sh. The earlier chained versions never started - they waited on `pgrep -f "mmq-rules-gpu.sh llama.cpp-e2e"`, which matched the shell whose argv contained the heredoc that created them, so each waiter matched itself. No pgrep waiting now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The counts that justify amending 903 come from nvcc compiling the host-side AMD branch of ggml_cuda_mmq_get_config(), not from a ROCm build or a GPU. Asked of the ROCm host in #449; until it answers, the gfx1151 numbers are a prediction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…s are left hclsys reported both gaps on #29953 from a GB10 the same day, independently and with the same formula for the src1 one, and the author pushed a fix. At head 3070d927f: nthreads_best carried out of the selection loop, src1 padded by GGML_PAD(J_best*sizeof(block_q8_1_mmq), nthreads_best*sizeof(int)), and ids_dst at ne_get_rows + J_best-1. That is our amendment in all but spelling, so there is nothing to send upstream about those two and mmq-amend-29953.patch is now historical. Two routes to the same formula is a cross-check on both. Still open there: src1_scale is untouched, and the stream-k fixup write-back is called with j_max == J, so it reads y_scale_tile[j] for the whole tile from y_scale + col_low + jt*J -- up to J_best-1 floats past the end, the same bound the PR just fixed for ids_dst. Every NVFP4 native-FP4 config uses stream-k on sm_120 and sm_121 (tasks/mmq-nvfp4-streamk-probe.cu), so the path is not exotic, and hclsys's sweep covered q4_0/q4_K/q6_K/q8_0 but not NVFP4. tasks/mmq-amend-29953-yscale.patch is the two-line fix against 3070d927f. It is read, not measured: native FP4 needs a 120a build and nothing here has one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
RETRACTION. The claim that src1_scale is read up to J_best-1 floats past its end does not survive measurement, and is withdrawn. On an sm_120a build -- native FP4 confirmed in play, the build reports prec=q4 and allocates the scale buffer -- a shape chosen to run the stream-k fixup kernel (8 experts, 1 used, 640x9x512: 40 tiles, 21% of the waves on 188 SMs) gives 0 memcheck errors with src1_scale in an exact-size allocation, and passes the guard page 3/3. Shrinking the allocation to one float reports 9 errors immediately, as writes from the quantize kernel, so the buffer is instrumented and live. Why the source's unbounded read does not materialise is not established, so the claim goes rather than becoming a guess. Nothing is outstanding for upstream from this work. Two of my own errors it corrects: - "no native FP4 on this host" was wrong. blackwell_mma_available() only needs highest_compiled_arch >= 1200, and __CUDA_ARCH_LIST__ is 1200 for 120 and 120a alike. The claim should never have been left unmeasured. - "every NVFP4 config uses stream-k so the path is not exotic" was beside the point: stream-k is always selected, but the fixup kernel runs only when the tiles fill under 90% of the waves. tasks/mmq-fixup-shape-search.cu finds shapes that do; the first one tried was at 97% and never ran it. - the report printed prec_src1 as "q8" for anything that was not F32, so it labelled q4 as q8 and the first run looked like the path was not taken at all. sm_75 (2080 Ti), new: dense321 is a plain MUL_MAT with no MUL_MAT_ID, and the published #29953 aborts on it 0/3 while ne12 passes -- master's get_J_max(321) over-pads to 128 blocks where #29953 gives exactly 80 against 85 needed. So the padded-tile defect is not MoE-specific, and in the dense branch the published fixup was a regression against master. Against today's head it is covered. ids16 reproduces identically on sm_75, so the sm_120 results were not architecture-specific. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…449 The ROCm host answered the ask. hipcc's config tables are byte-identical to nvcc's, so the gfx1151 column is no longer modelled, and under a guard page the widest-tile 903 aborts 6/6 on q2_K at J=80 with one row in the last expert while the amended form passes 6/6, every passing run matching the CPU backend. The amendment is confirmed on hardware. Four things I had wrong: 1. The NVIDIA MMVQ limit is not "Turing and newer". Volta, and Ada Lovelace and newer, always take MMVQ for MUL_MAT_ID up to MMVQ_MAX_BATCH_SIZE for every type; only Turing and Ampere use the per-type table. The check counted 496 shapes that MMVQ actually takes on sm_89/sm_120, so every total moves (4,717,824 -> 4,717,328) and #29953's q2_K J=8 worst case holds on sm_75/86 but not sm_89/120. 2. The gfx1151 shortfall is q2_K ONLY, not q2_K and q3_K: q3_K's non-fallback table reaches J=128 there. mmq-jpad-cost.cu had been printing exactly one row all along and I wrote two types anyway. 3. The cost is 720 bytes, not 864. The patch floors nbytes_pad_y/144 and so reserves 85 blocks, not the 86 a ceiling gives; floor(T/B) >= ceil(T/B) - 1 is the requirement, so it is sufficient. The check and the cost tool now score the floor, i.e. the rule the code actually implements. 4. #449 told them the guard allocator would not port to HIP. It does: HIP VMM is supported on that APU with 4 KiB granularity and vendors/hip.h maps the cuMem* calls. Their three edits are now in mmq-debug-alloc.cuh behind GGML_USE_HIP, and the CUDA build still compiles. Also recorded from their reply: counted is not reachable (2,648 of the 25,384 reach MMQ there, 2,616 realisably); test_mul_mat_id cannot produce the gap at all because its routing is uniform, which is a limit of this harness; main's 903 has its own gfx1151 gap at 6-of-128 q4_K and 5 tokens; and #29953's head is clean there too, so once the pin includes it 903's padding is redundant on ROCm. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… too, by 7 blocks Extending the check to sm_70/80/90, CDNA3 and RDNA4 (9,431,816 ids-branch shapes) shows #448 as published is short on three AMD architectures, not just gfx1151: CDNA3 576,000 shapes, gfx1151 25,256, RDNA4 25,256, and zero on all seven NVIDIA ones. nthreads is 512 on CDNA rather than 256, so the padded tile is larger and the worst case grows from 5 blocks to 7 (q4_0, non-fallback, J=64: GGML_PAD(64*144, 2048) = 10,240 needs 71 against the widest tile's 64). The dense branch is short in 1,960 CDNA3 shapes. #29953's worst case grows from 6 blocks to 12 for the same reason. The amended rule is covered on all ten, which it is by construction: it takes the max padded tile over every config that exists for (type, fallback, cc), the launch can only pick one of those, and smpbo only removes candidates. So the sweeps confirm the construction rather than carrying it. The five added architectures are modelled, not measured, and flagged as such. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… VMM fix Answers #449 from the ROCm host (amd-server/rocm, gfx1151, ROCm 7.2.1). fa5ad16 already took the corrections the reply listed. This adds what they rest on. - The tools: the b11081 forms of 903 with the debug hooks (read by SHA), the gfx1151 sweep with the dispatch cut from the source, the production tensors, a hand-routed MUL_MAT_ID, the HIP probes, and the build/run/table harness. - The results, under mmq-successor-results/rocm-gfx1151/: hipcc vs nvcc by sha256, the 12 cases x 4 forms x 3 modes, the hand-routed cases, the production scan, llama.cpp#29953 at 3070d927f, and the host profile. - mmq-debug-alloc.cuh keeps its VMM reservations on HIP. On gfx1151, a range that is freed and reserved again reads and writes the wrong memory (mmq-hip-vmm-reuse.hip). Every published run makes one guarded allocation per buffer and is unaffected. A process with several, such as the per-expert fallback, returned wrong results without the fix. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…/RDNA4 reach Adds what #450's review and the re-tightening question asked for. - mmq-padding-invariant.cu checks every config of all ten tables against four padding rules: 2,597 configs in about 10 ms, without a GPU. It reproduces #448's per-shape verdicts and adds GCN, where the widest-tile rule is short in 28 configs by up to 7 blocks. It calls the table functions directly, so nvcc and hipcc evaluate it alike. - mmq-padding-guard.cu is a compile-time prototype: one static_assert per (table, type), with the requirement written from the kernel's load loop. Under hipcc, re-tightening to J blocks fails the build in all ten tables, and the widest-tile rule fails in cdna, gcn, rdna3_5 and rdna4. The padded-tile rules build. - mmq-gfx1151-reach.cu takes --arch. On CDNA3, all 595,200 counted widest-tile shapes reach MMQ and are realisable; on RDNA4, 23,746 of 25,384. Both are modelled. - multi-call.txt: 175 guarded dense calls in one process pass with and without the VA fix, so the reuse defect needs a free-and-reserve inside one computation. The leaked reservations do not show above the suite's own allocations. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
903's src1 padding rule moves into ggml_cuda_mmq_get_J_pad() in mmq.cuh, which the allocation calls. The y-tile size moves into ggml_cuda_mmq_get_nbytes_y_tile(), which mmq_get_nbytes_shared() now uses too. A compile-time guard at the end of mmq.cu checks the helper against every config of all ten config tables. It makes one static_assert per (table, type) and states the requirement from the kernel's load loop. Measured on gfx1151 (hipcc, ROCm 7.2.1): - It builds. mmq.cu takes 3.6 s, against 2.0 s without the guard. - Re-tightening the rule to the widest tile fails the build in cdna, gcn, rdna3_5 and rdna4. Shrinking the y-tile helper to J blocks does too. - The device reports the same J, nthreads, need and pad as the amended 903 on every hand-routed shape, and the twelve cases pass stock. - The full compat series applies in order to b11081. HIP needs __HIP_DEVICE_COMPILE__ to find the host pass. vendors/hip.h defines __CUDA_ARCH__ in every HIP pass, and the first draft keyed off it: the guard compiled out silently and passed both mutations. Not built with nvcc. The CUDA host needs to check the constexpr lambda and the static_asserts there. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
4 of 5 tasks
docs(mmq): Add gfx1151 measurements for #449 and keep HIP VMM reservations
refactor(compat): Pad through one helper in 903 and guard it at compile time
glennneuber
marked this pull request as ready for review
October 5, 2026 00:20
Author
|
Merged to main as Verified from
Merged over five pending checks, all of them the Go test/race matrix restarted by marking this ready. Nothing was failing, this PR touches no Go code, and Where this leaves the subject:
ai-server/mlx-cuda |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Material for the MMQ
MUL_MAT_IDtail-padding defect, covering llama.cpp#29953 — the fixup to #29941, whose author asked for a reproduction they could not get locally. Reproduced, and it corrected this PR's own earlier claim.Important
Upstream fixed both of these on 2026-10-05, independently.
hclsysreported the same two gaps on #29953 from a GB10 the same day, with the same formula for the src1 one, and the author pushed a fix: at head3070d927fsrc1 is padded byGGML_PAD(J_best*sizeof(block_q8_1_mmq), nthreads_best*sizeof(int))andids_dstbyJ_best-1. So nothing here is an upstream ask, andmmq-amend-29953.patchis historical.A third claim — that the NVFP4 y scales are read out of bounds the same way — is RETRACTED. It came from reading the source and does not survive measurement on an
sm_120abuild with native FP4 confirmed in play: a shape chosen to run the stream-k fixup kernel gives 0 memcheck errors withsrc1_scalein an exact-size allocation. See the retraction section in the material;mmq-amend-29953-yscale.patchis kept as the record of what was tried and is not to be sent anywhere.What this PR is for: the fork's compat 903, which needed the same correction as upstream and is re-cut and verified against the pin; the sm_75 result showing the defect is not MUL_MAT_ID-specific; and the record of how it was found.
The 903 amendment is confirmed on gfx1151 hardware (#449): hipcc's config tables are byte-identical to nvcc's, and under a guard page the widest-tile form aborts 6/6 on q2_K at
J = 80while the amended form passes 6/6. That reply also corrected four things here — the NVIDIA MMVQ limit (496 shapes miscounted), q2_K-only rather than q2_K and q3_K, 720 B rather than 864 B of cost, and that the guard allocator does port to HIP. All folded in.Nothing here is for posting upstream as written. llama.cpp prohibits AI-written posts, so any reply is the maintainer's to write, from
docs/maxusai/upstream-mmq-29953-material.md. The code (the patches, the test cases, the checks, the harness) is AI-generated, which upstream allows with disclosure.The arithmetic everyone got wrong, us included
A tile's src1 load is unbounded in
l:so it copies
GGML_PAD(J*sizeof(block_q8_1_mmq), nthreads*sizeof(int))bytes from the tile's first column — the same expressionmmq_get_nbytes_shared()already uses to size the shared-memory tile it lands in. The requirement isceil(that/144) - 1blocks, notJ - 1. On sm_120 withnthreads = 256:J = 16needs 21,J = 112needs 113, andJ = 128needs 127 — exactlyJ - 1, which is why every large-batch case is covered and a reproduction has to land on a narrower tile.Five rules, 4,717,824 shapes
ids_dstne11(#24127)ne12(#29941, master)ne12*n_expert_used(#27044)#29953 trades breadth for depth — short in more shapes than master but by at most 6 blocks instead of 63, because it fixes the rounding and leaves only the padded-tile term.
ids_dstis padded by no rule but this one: a tile loadsJentries from the expert's first row with no bound.The reproduction
test_mul_mat_id(Q4_0, F32, 256, 16, /*b =*/ false, 640, 16, 2560)—J = 16, src1 needs 21 blocks,ids_dst15 entries. Stock it passes under every rule; the debug allocator exposes it:MMQ_DEBUG_ALLOCne11ne12exact+ memcheckguard:src1guard:ids_dstguardmode places a buffer so its last byte is the last byte of its own VMM mapping with the next granule reserved and never mapped, so an over-read faults on the device — no sanitizer, deterministic, which is what makes this reproducible where memcheck sees nothing. The:buffersuffix attributes the fault.The over-read is a function of the routing as well as the shape: a run reads
ceil(T/144) - rblocks past, whereris the last non-empty expert's row count, so the worst case needsr = 1.ids16realises it (one row per expert); theJ = 112cases haver ≈ 25and do not fault. Third masking layer, after the VMM pool and the stream-k fixup buffer.Changes
docs/maxusai/upstream-mmq-29953-material.md— the facts, with the measurement behind each claim.upstream-mmq-successor-material.mdkeeps the history and now carries a CAUTION banner over the two claims this supersedes.docs/maxusai/tasks/:mmq-amend-29953.patch— 8 lines on top of #29953 (verified against5bd8b0013), closing both gaps. The recommendation.mmq-fix-amended.patch— the equivalent standalone change against master, for compat 903.mmq-debug-alloc.cuh—MMQ_DEBUG_ALLOC=guard[:src1|:ids_dst](guard page) and=exact(owncudaMalloc);MMQ_DEBUG_PRINT=1has the build report its ownJ,nthreadsand padding. Debug only, never shipped.mmq-rules-check.cu+build-check.sh— the CPU check, six rules, no GPU. The gencode list is load-bearing:ggml_cuda_highest_compiled_arch()reads__CUDA_ARCH_LIST__, and without itget_J_max(512)comes back 64 on sm_120 where the device uses 128. This PR's earlier counts were from such a build.mmq-tests.patch— twelvetest_mul_mat_idcases,ids16andids64new.mmq-variant.py,mmq-rules-gpu.sh,mmq-rules-table.py— seven padding rules built from one checkout, every case run four ways, table from the generator (ADR 0012).mmq-successor-results/—check-rules.txt,case-shapes.txt(the device's own per-case report),matrix.md, the real-model logs.For the fork
Compat 903 needs re-cutting. As merged here it pads by the widest tile, which covers every NVIDIA shape checked but is short on three AMD architectures — not just gfx1151:
nthreadsis 512 on CDNA rather than 256, so the padded tile is larger:J = 64needsGGML_PAD(64*144, 2048) = 10,240bytes, 71 blocks, against the widest tile's 64. The dense branch is short in 1,960 CDNA3 shapes too. gfx1151 is measured (#449); CDNA3 and RDNA4 are modelled, and flagged as such.The amended rule is covered on all ten architectures checked, and is sufficient by construction: it takes the max padded tile over every config that exists for
(type, fallback, cc), the launch can only pick one of those, andsmpboonly removes candidates.mmq-fix-amended.patchis the corrected form — the maintainer's call.Nothing production serves is affected today. Its MoE GGUFs (8 of 128 for gemma4:26b-a4b, 8 of 256 for qwen3.6:35b-a3b, 6 of 128 plus a shared one for nemotron3:33b) run
J = 128at the image ubatch sizes in use, where the padded tile equalsJblocks and #29941's padding already covers src1. Theids_dstread happens on every MoE call but lands inexpert_bounds.Test plan
MMQ_DEBUG_PRINT=1(case-shapes.txt), which is how the gencode fault was caught.ids16on sm_120: stock passes under all five rules;guard:src1aborts forne11,ne12and #29953;guard:ids_dstaborts for all four and passes only with the widest-tile rule.mmq-amend-29953.patchapplies to #29953's head and produces the same tree as thep29953fixvariant.qwen3.6:35b-a3b-q4_K_M, 2048-token image ubatch, exact-size src1): rawne11core-dumps at the image decode; #29941 and the widest-tile 903 decode cleanly.J = 128there, so the padded-tile term is zero.mmq-rules-gpu.shend to end: 7 rules × 12 cases × 4 modes on sm_120, CUDA 12.8, plus a combined-guard pass and ten repeats on the marginal cells — 749 rows,matrix.mdin the results directory. Under a guard page onids_dstall twelve cases abort forne11,ne12, #29953 and #27044 and pass for the three rules that pad it; under exact-size allocations with memcheck all four published rules report errors (43 to 11,939) against 0 for the three amended ones.J = 112cells repeated ten times with the amendment as an in-harness control. #29953 aborts 4/10 onj100_b1and 1/10 onj100_b0(the one-block shortfall, which needs the routing to leave one row in the last non-empty expert); the amendment is 0/10 on both.ids16is 10/10 deterministic.mmq.cuandmmq.cuh; no test file. And the suite structurally cannot catch it — the pool masks the read (VMM and legacy alike), andtest_mul_mat_id's routing is uniform so wide tiles never leave ~1 row in the last expert. All 12 cases here pass stock on CUDA and 144/144 on ROCm. A re-tightening would go unnoticed, which is what happened in #29953's dense branch.dd266785c; widest-tile 903 aborts 6/6 and amended passes 6/6 underguard:src1on a hand-routed q2_KJ = 80shape, each passing run matching the CPU backend; 144/144 pass stock; #29953's head clean there too (72/72 + 36/36).check_source_paths.py(no new failures), name scan.ai-server/mlx-cuda🤖 Generated with Claude Code