Repository navigation
speculative: fix failed to decode mtmd chunk with DFlash - #28587
Conversation
When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset. Stop copying them to allow the drafter to continue.
|
Hi @jesdga95, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
/bot review |
Automated code reviewReview of Will slow the review (point 1) The guard keys on the wrong model. (point 2) The skip is broader than the comment claims. (point 3) The new comment contradicts the comment six lines above it, which states that skipping embedding batches "leaves a hole in the draft's cache and the next injection fails to initialize". The new code now deliberately skips them and asserts "the draft can jump over the gap". If the hole is genuinely tolerable for DFlash, please update the older comment to carve out this case (or reference it); as written a reader gets two opposing claims about the same situation and cannot tell which is true. Nits (point 4) Comment precision, (point 5) Follow-up worth checking after this lands: Otherwise the change is minimal, single-purpose, ASCII-clean, and in the right spot; disclosure is filled in. This review was generated automatically by pi coding agent using |
|
This is a wrong fix, vision input works well from my testing. |
|
@ruixiang63 I am able to reproduce this every single time in the checkpoint I posted, if you don't mind, what is your checkpoint? large ubatches might make this is a non issue (try with -ub 64, although I could reproduce with -ub 1024). I'm addressing the comments from the review either way. There is also this second error that's the same root issue when -ub is big enough: |
|
Did you build the current master branch to test? |
limit M-RoPE skip to images only, allow audio to pass through. Clean up comments to align to the updated implementation
|
@ruixiang63 rebuilt on current master (5d806aa), same errors: Small image, fails on the text prefill right after the image: Larger image, fails while injecting the image itself, third chunk: Here's the image if it matters:
Setup: Qwen3.8-27B-UD-Q5_K_S + mmproj-qwen3.8-27b-F16 + Qwen3.8-27B-DFlash2-Q8_0, -ub 512: |
|
I can't reproduce that. cc @ngxson to take a look. |
|
@ruixiang63 I have a theory on why you might not be able to repro it, the z-lab DFlash2 for Qwen 3.8 27B has Tested on a new clean build from master (5d806aa):
So full repro: use a high resolution image and send it as the first message with this dflash file, and |
Thanks! Will give it a try. |
|
I faced the same crash on 3× gfx1030 GPUs with ROCm. This fix prevents the crash for me, but DFlash performance appears to degrade significantly after vision input - TG drops to ~11 t/s, which is about the same as without DFlash enabled. |
|
@nouveau-nvc0 thanks for testing. The slowdown after an image is a different bug which I was also able to reproduce. One-line fix here: jesdga95@0c5dff206 (use |
|
Thanks for identifying this and proposing a fix. Vision inputs are not well supported by current speculative decoding logic, and the lack of vision-specific training can also reduce the acceptance rate for DFlash/DFlash2. That said, I would prefer to avoid a hotfix, since the current solution in dflash is also a workaround. The proper long-term fix appears to be #24669. I suggest waiting for that PR to be merged and then verifying whether it resolves the issue definitively. |
This fix worked. Reasoning performance is back to my normal 20–26 t/s (1 slot) |
|
/bot review |
Automated code reviewThis follow-up review covers the changes since commit Previous review points - all addressed (point 1) The wrong-model gating ( (point 2) Audio and other 1d-media embedding batches (whose positions advance via (point 3) The contradictory comment is resolved - the old text about "leaves a hole ... and the next injection fails to initialize" is replaced with "embeddings are injected too, except the pinned ones skipped below", which now matches the code. Nits (point 4) (point 5) Known follow-up, not for this PR: The change remains minimal, single-purpose, ASCII-clean, and the new comment is concise and accurate (the injected rows really are all written at one draft position: the This review was generated automatically by pi coding agent using |
* speculative: fix failed to decode mtmd chunk with DFlash When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset. Stop copying them to allow the drafter to continue. * address PR feedback limit M-RoPE skip to images only, allow audio to pass through. Clean up comments to align to the updated implementation
* speculative: fix failed to decode mtmd chunk with DFlash When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset. Stop copying them to allow the drafter to continue. * address PR feedback limit M-RoPE skip to images only, allow audio to pass through. Clean up comments to align to the updated implementation
* speculative: fix failed to decode mtmd chunk with DFlash When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset. Stop copying them to allow the drafter to continue. * address PR feedback limit M-RoPE skip to images only, allow audio to pass through. Clean up comments to align to the updated implementation
* speculative: fix failed to decode mtmd chunk with DFlash When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset. Stop copying them to allow the drafter to continue. * address PR feedback limit M-RoPE skip to images only, allow audio to pass through. Clean up comments to align to the updated implementation
* speculative: fix failed to decode mtmd chunk with DFlash When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset. Stop copying them to allow the drafter to continue. * address PR feedback limit M-RoPE skip to images only, allow audio to pass through. Clean up comments to align to the updated implementation
* speculative: fix failed to decode mtmd chunk with DFlash When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset. Stop copying them to allow the drafter to continue. * address PR feedback limit M-RoPE skip to images only, allow audio to pass through. Clean up comments to align to the updated implementation (cherry picked from commit fa67698)
* speculative: fix failed to decode mtmd chunk with DFlash When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset. Stop copying them to allow the drafter to continue. * address PR feedback limit M-RoPE skip to images only, allow audio to pass through. Clean up comments to align to the updated implementation

Overview
When using DFlash w/ vision models, the drafter memory fails to allocate new tokens because images report a fixed offset and errors out (Tested on Qwen3.8 27B Q5_S, RTX 5090 CUDA 13.3):
Stop copying them to allow the drafter to continue.
Additional information
Repro:
llama-serverwith a target +--mmproj+ a DFlash draft, use an image bigger than whatever -ub is set at.Requirements