Skip to content

spec : fix n-gram drafts rejected at temp > 0 after truncation - #29924

Merged
ggerganov merged 1 commit into
ggml-org:masterfrom
praneshgo:pgonegandla/spec-ngram-truncate-fix
Oct 4, 2026
Merged

ggerganov merged 1 commit into
ggml-org:masterfrom
praneshgo:pgonegandla/spec-ngram-truncate-fix

Conversation

@praneshgo

Copy link
Copy Markdown
Contributor

Overview

Fix the regression where, at temperature > 0, n-gram drafts cut short near the end of a reply were wrongly rejected outright, slowing generation down. The regression affects n-gram drafters and is fixed with this PR; draft-simple and draft-mtp are not affected by the regression.
Regression introduced by #27694 commit 1fb7ef3e3.

Root cause

  • The server passes a candidate list (result_q) to the drafters whenever temperature > 0. use_spec_rejection() checks only temp > 0.0f, and --spec-draft-sampling doesn't affect this.

  • draft-simple and draft-mtp either clear the list in greedy mode or fill one candidate list per drafted token in probabilistic mode. N-gram drafters (and EAGLE3/DFlash/DSpark) never touch it.

  • common_speculative_draft() truncates drafts longer than dp.n_max, which is the remaining room before max_tokens or the end of the context. It also resizes result_q to n_max. On an untouched, empty list this creates n_max empty candidate lists.

  • The verifier then sees a non-empty list whose size matches the draft, so it switches to rejection sampling. With q = 0 for every drafted token, it rejects the whole draft.

  • Effect: near the end of a reply, every truncated n-gram draft is fully rejected.

  • Both greedy (default) and probabilistic modes are affected, temperature 0 is not affected, draft-simple and draft-mtp on their own are not affected.

Fix

  • When a draft is truncated, the candidate list is now trimmed only if it isn't empty, i.e. only if the drafter actually produced candidates.
  • An empty list stays empty, so n-gram drafts are verified by greedy sample-and-match, as before #27694.
  • Drafters that do produce candidates (draft-simple and draft-mtp in probabilistic mode) still have their list trimmed together with the draft, as #27694 intended.

Testing

Tested on RTX 5090, following decode perf data is with builds corresponding to pre-PR 134b2bb, ToT 836d571, ToT + fix (this PR). Temperature 0.7 over 8 seeds:

Llama-3.1-8B-Instruct Q4_K_M (dense):

--spec-type build median t/s worst seed t/s mean drafted mean acceptance
ngram-mod pre-PR 509.6 419.1 1048 0.493
ToT 375.1 334.7 1368 0.365
fix 500.7 419.0 1048 0.493

Qwen3.6-35B-A3B Q4_K_M (hybrid, built-in MTP):

--spec-type build median t/s worst seed t/s mean drafted mean acceptance
ngram-mod,draft-mtp (np=1) ToT 301.7 203.1 1306 0.496
fix 333.9 304.9 1168 0.553
ngram-mod ToT 227.8 184.9 849 0.377
fix 292.8 262.0 780 0.514
ngram-mod,draft-mtp (np=2) ToT 187.0 145.9 1254 0.538
fix 329.4 225.4 1041 0.670
ngram-mod,draft-mtp (--spec-draft-sampling probabilistic) ToT 307.6 218.4 1233 0.522
fix 344.1 314.2 1214 0.529

Requirements

@praneshgo

Copy link
Copy Markdown
Contributor Author

@ggerganov, @danbev, @gaugarg-nv, @lbyte
Can you please review this PR? Thank you.

@lbyte

lbyte commented Oct 3, 2026

Copy link
Copy Markdown

Verified, the regression is fixed.
Thanks!

@ggerganov
ggerganov merged commit 8330e96 into ggml-org:master Oct 4, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants