Repository navigation
docs(upstream): llama.cpp#27044 was reviewed; test the maintainer's variant on sm_120 - #445
Conversation
…t still needs a test On 2026-10-04 the upstream CUDA maintainer said #27044 "looks 90% correct". He opened #29941, which pads the MMQ ids path with ne12 instead of ne12*n_expert_used, and asked whether it works for us. The submission record now holds the facts for the maintainer's reply: - the confirmations, and #29847's test-backend-ops reproducer; - how the two lines differ. From 512 tokens up they give the same padding, which covers every crash reported. Below that, #29941 can pad too little where the launch rounds J up; - what this means for compat 903; - the tests that are left for the CUDA host. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd the tests behind them The CUDA host ran #445's checks against llama.cpp master 05043961 on an RTX PRO 6000 Blackwell (sm_120, CUDA 13.0). It built three variants that differ only in the ids-path padding argument, and ran each case under compute-sanitizer memcheck with the src1 buffer in its own exact-size allocation. - #29941 (ne12) reads past its buffer at 65 and 100 tokens. #27044 (ne12*n_expert_used) is clean in every case. - At 508 and 2040 tokens, #29847's cases and the original fault, both lines are clean. master fails everywhere. - The stock pool hides the over-read in every short-batch case. - master's mm_ids_helper fails to launch on this host, with 1 KB of static shared memory and a dynamic limit raised to the device maximum. A one-line workaround, common to all three builds, gets past it. Adds the GPU-free check (tasks/mmq-ids-padding-test.cu); the llama.cpp patch with the six cases, the debug switch and the workaround (tasks/mmq-ids-padding-gpu.patch); and the driver that builds the variants and runs the matrix (tasks/mmq-ids-padding-gpu.sh). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The CUDA host ran #445's tests on sm_120, at the maintainer's request. #29941 reads past its buffer below 128 tokens; #27044 is clean in every case. The record and the test code are in Setup.
Memcheck errors when the src1 buffer has its own exact-size allocation (✓ test passed, ✗ aborted):
With the stock memory pool, every cell is 0 ✓ except What it shows.
Two local changes, common to all three builds:
Test code, in
The results section is in
|
…pstream's The record said the failure is a separate upstream problem. It is not shown to be. mmid.cu is byte-identical in b11081, b11232 and master 05043961, and production's b11081 build runs this path on the same card. Production's sm_120a PTX for the helper declares no static shared memory, while these builds' native sm_120a code reports 1 KB. The cause more likely lies in these builds' configuration. The padding results do not change, because the workaround is the same in all three builds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
A correction to my comment above: the
The padding results do not change: the workaround was the same in all three builds. The record now says this, in
|
…heck starts at MMVQ's real limit Three corrections found in review: - On NVIDIA, no tile wider than 128 has a config, so both lines pad 128 blocks from 128 tokens up. #29941 pads less only below 128 tokens, not below 512. - At 100 tokens on sm_120 there is no config at J = 104. The launch takes J = 112, so #29941 can fall up to 15 blocks short, not 7. - The CPU check assumed MMVQ takes every MUL_MAT_ID batch up to 8 tokens. On Turing and newer it does not for q2_K (7) or q3_K (5), so 960 shapes were missing. The check now starts each type above its real limit. Corrected counts: 2,439,360 shapes, #29941 short in 268,440, #27044 in none. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review of #445 at Corrected:
Checked:
The #27044 reply stays the maintainer's to write.
|
A hand-off to the CUDA host (
ai-server/mlx-cuda), on the maintainer's word. Our upstream MMQ fix, ggml-org/llama.cpp#27044, has been reviewed. The reply needs tests on sm_120, the hardware of the original crash, undercompute-sanitizer.What happened upstream (2026-10-04)
ne12instead of ourne12*n_expert_used.test-backend-opsreproducer, and suggested adding its two cases to #27044.What this PR records
docs/maxusai/upstream-mmq-submission-material.mdgets a section with the facts:master, not from a measurement:For the CUDA host
compute-sanitizer --tool memcheck, runmaster,master+ 903 andmaster+ #29941 on these cases:test_mul_mat_idcases;CONTRIBUTING.md; see the top of the record). The maintainer writes the reply from the results.Checks
check_source_paths.py --changed-since origin/mainis clean.check_no_names.py, with a local deny-list file: no match.amd-server/rocm-gfx1151🤖 Generated with Claude Code