Skip to content

llama : re-reserve the sched when the nextn extraction flags change - #30020

Merged
pwilkin merged 1 commit into
ggml-org:masterfrom
pwilkin:fix-mtp-nextn-reserve
Oct 6, 2026
Merged

pwilkin merged 1 commit into
ggml-org:masterfrom
pwilkin:fix-mtp-nextn-reserve

Conversation

@pwilkin

@pwilkin pwilkin commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

#29958 fixed reallocation in the non-MTP path, but the MTP path still failed under GGML_SCHED_DEBUG_REALLOC because the initial graph was built before the nextn extraction flags were set.

Requirements

The speculative MTP init enables NextN extraction on the target and draft
contexts after both were created and their schedulers reserved. With
unmasked extraction the trunk graph keeps every token through the last
layer instead of cropping to the output rows, so the first decode
reallocates to that batch's shape and the next, wider batch trips
GGML_SCHED_DEBUG_REALLOC. Invalidate the reserve when the flags change so
the next compute re-reserves with the new graph shape.

Assisted-by: Claude
@pwilkin
pwilkin requested a review from ggerganov as a code owner October 5, 2026 20:13
@ggerganov ggerganov self-assigned this Oct 5, 2026
@ggerganov

Copy link
Copy Markdown
Member

Yes, I think we need that. I am also trying to consolidate the logic across all the graphs to make it easier to follow: #30017

@ggerganov

ggerganov commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Let's do the following order:

cc @am17an @ServeurpersoCom

@pwilkin
pwilkin merged commit 1a3011c into ggml-org:master Oct 6, 2026
12 checks passed
@am17an

am17an commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

wouldn't this cause a big re-reserve during inference? Possibly this will be fixed via #30017

@ggerganov

Copy link
Copy Markdown
Member

wouldn't this cause a big re-reserve during inference? Possibly this will be fixed via #30017

Hm, did I miss some case? The nextn flags are only changed once upon speculative context construction and shouldn't change during inference?

wanghqc added a commit to qualcomm/llama.cpp that referenced this pull request Oct 8, 2026
…change (ggml-org#30020)"

This reverts commit 1a3011c on this branch only.

With it, Qwen3.8-27B with its built-in MTP head (--spec-type draft-mtp) produces a
different greedy output on every request through llama-server on the X2-90, while
upstream master's OpenCL backend and this branch without the change are
deterministic. The re-reserve changes the compute-buffer layout and exposes a
branch-specific read of unwritten memory that is still being located. E4B with
its MTP assistant keeps the same speed and output without the change.
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
…gml-org#30020)

The speculative MTP init enables NextN extraction on the target and draft
contexts after both were created and their schedulers reserved. With
unmasked extraction the trunk graph keeps every token through the last
layer instead of cropping to the output rows, so the first decode
reallocates to that batch's shape and the next, wider batch trips
GGML_SCHED_DEBUG_REALLOC. Invalidate the reserve when the flags change so
the next compute re-reserves with the new graph shape.

Assisted-by: Claude
(cherry picked from commit 1a3011c)
wanghqc added a commit to qualcomm/llama.cpp that referenced this pull request Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants