Skip to content

speculative-simple : pass ctx_other to the draft context and fail gracefully - #26883

Closed
bri-prism wants to merge 1 commit into
ggml-org:masterfrom
PrismML-Eng:fix-spec-simple-dflash-ctx-other
Closed

bri-prism wants to merge 1 commit into
ggml-org:masterfrom
PrismML-Eng:fix-spec-simple-dflash-ctx-other

Conversation

@bri-prism

@bri-prism bri-prism commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Overview

llama-speculative-simple creates the draft context without ctx_other and never null-checks the result. For drafters that borrow tensors from the target model (DFlash models without their own token embeddings or output head, such as the deepseek-ai dspark checkpoints), llama_init_from_model returns null, and the unchecked ctx_dft reaches llama_decode, which crashes.

This PR makes the following changes:

  • Passes the target context as ctx_other when creating the draft context.
  • Exits with a clear error message if the draft context still cannot be created.
  • Warns, for DFlash drafters, that this example does not stage target-model features, so draft acceptance will be much lower than through llama-server, which uses the full framework path.

Additional information

The crash can be reproduced with:

llama-speculative-simple -m Qwen3-8B-bf16.gguf -md dspark_qwen3_8b_block7.gguf \
  --spec-type draft-dspark -ngl 99 -ngld 99 -fa on --temp 0 -n 128 -p "..."

The failure occurs at prompt evaluation. The "dflash requires ctx_other" message emitted during context creation is easy to miss, because the same message also appears benignly during memory fitting.

The feature-staging gap itself is described in the companion issue #26884. I am happy to work on full DFlash support in this example if there is interest.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. Claude Code was used to help develop and test the fix and to format this description to the PR template. I reviewed every line and take full responsibility for the changes.

…cefully

Drafters that borrow tensors from the target model (e.g. DFlash models
without their own token embeddings / output head) require ctx_other at
context creation. The example creates the draft context without it, so
llama_init_from_model returns null, and the unchecked ctx_dft is later
passed to llama_decode, which crashes.

Pass the target context as ctx_other, error out cleanly if the draft
context still cannot be created, and warn that this example does not
stage target-model features for dflash drafters (acceptance will be much
lower than with llama-server, which uses the full framework path).
@ggml-gh-bot

ggml-gh-bot Bot commented Aug 11, 2026

Copy link
Copy Markdown

Hi @bri-prism, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@ggml-gh-bot ggml-gh-bot Bot added the draft PR will be changed to draft by github-actions bot label Aug 11, 2026
@github-actions
github-actions Bot marked this pull request as draft August 11, 2026 04:16
@github-actions github-actions Bot removed the draft PR will be changed to draft by github-actions bot label Aug 11, 2026
@bri-prism
bri-prism marked this pull request as ready for review August 11, 2026 05:36
@bri-prism bri-prism closed this Aug 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant