Skip to content

server: correct accepted tokens when need draft token replay - #26320

Merged
ggerganov merged 2 commits into
ggml-org:masterfrom
ruixiang63:spec_metrics_fix
Jul 31, 2026
Merged

ggerganov merged 2 commits into
ggml-org:masterfrom
ruixiang63:spec_metrics_fix

Conversation

@ruixiang63

Copy link
Copy Markdown
Member

Overview

When the target context can't roll back a partially-accepted speculative draft in place — i.e. it falls back to a full checkpoint restore (COMMON_CONTEXT_SEQ_RM_TYPE_FULL, or RS with n_rollback > n_rs_seq) — the server restores the checkpoint and replays the accepted prefix plus the target's own correction token as the next draft.

On that replayed step every token is accepted, so the per-request speculative stats counted the target-generated correction token as an accepted draft token, inflating draft acceptance, per-position acceptance and mean len by 1 for each checkpoint restore.

This marks the slot when a checkpoint replay is queued (is_draft_replay) and excludes that single correction token from the acceptance statistics on the replayed step.

This change only affects reported per-request stats only.

Additional information

Requirements

@ruixiang63
ruixiang63 requested a review from a team as a code owner July 30, 2026 13:42
@ruixiang63

Copy link
Copy Markdown
Member Author

@ggerganov Can you please take a look?

@ruixiang63 ruixiang63 changed the title spec: correct accepted tokens when need draft token replay server: correct accepted tokens when need draft token replay Jul 30, 2026
@ggerganov ggerganov self-assigned this Jul 30, 2026
@ggerganov ggerganov added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Jul 31, 2026
@ggerganov
ggerganov merged commit 0005475 into ggml-org:master Jul 31, 2026
17 of 26 checks passed
@ruixiang63
ruixiang63 deleted the spec_metrics_fix branch July 31, 2026 15:51
huaxel pushed a commit to huaxel/CachyLLama that referenced this pull request Aug 2, 2026
…g#26320)

* spec: correct accepted tokens when need draft token replay

* cont : naming

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
rcmorano pushed a commit to rcmorano/ROCmFPX that referenced this pull request Aug 6, 2026
Port of ggml-org/llama.cpp#26320 (upstream commit 0005475).

When the context does not support partial draft acceptance, the draft is
truncated and the state restored, and the replayed token was then counted as
an accepted draft token. That inflated both n_draft_accepted and
n_accepted_per_pos by one for every replay, so the acceptance rate and mean
acceptance length reported by the server (and used to tune --spec-draft-n-max
and --spec-draft-p-min) read higher than reality.

Adds a spec_is_replay flag on the slot: cleared in reset(), set on the
truncate-and-restore path, and consumed when tallying acceptance.

This only corrects reported speculative metrics. Generated output and decode
throughput are unaffected.

This tree shares no git ancestry with upstream (it begins at a published
source snapshot), so cherry-pick could not apply the commit directly; the
change was applied by hand at the four equivalent sites.

Co-authored-by: Ruixiang Wang <wangruixiang07@outlook.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
…g#26320)

* spec: correct accepted tokens when need draft token replay

* cont : naming

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
…g#26320)

* spec: correct accepted tokens when need draft token replay

* cont : naming

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
…g#26320)

* spec: correct accepted tokens when need draft token replay

* cont : naming

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…g#26320)

* spec: correct accepted tokens when need draft token replay

* cont : naming

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…g#26320)

* spec: correct accepted tokens when need draft token replay

* cont : naming

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. server

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants