Summary
We maintain a Rust llama.cpp host (in-process via the llama-cpp-rs bindings — no llama-server subprocess). We cannot use any of llama.cpp's speculative-decoding machinery (including the new DeepSeek V4 DSpark path from #25784) because it lives entirely in the common/ layer, which is tied to llama-server's main loop.
The gap
As of 596a579 (and latest master), the core llama.h API exposes no speculative params:
llama_model_params has no draft_model / spec fields.
llama_context_params has no spec fields.
- There is no
llama_speculative_* init function in the public API.
DSpark (and MTP) support is implemented in common/speculative.cpp / the server layer only. A library consumer that links llama.cpp and drives llama_decode directly cannot:
- load a draft model alongside the target, or
- enable
--spec-type draft-dspark style speculation, or
- benefit from DSpark's Markov/confidence-head draft at all.
Ask
A minimal core-API speculative surface for in-process hosts, e.g.:
llama_model_params.draft_model (+ spec type / draft-n-max) re-added, OR
- a
llama_speculative_* API that manages draft+target contexts and exposes a llama_decode_speculative-style call.
We (and likely other language bindings that bypass llama-server) would be happy to test, provide repro cases, or contribute to a design/implementation. We already drive the bundled llama.cpp in production for DeepSeek V4 Flash 0731 (284B/13B MoE, deepseek4 arch) and are currently forced to use our own draft→verify loop, which cannot use DSpark's confidence head.
Context
Thanks for considering — happy to help move it forward.
Summary
We maintain a Rust llama.cpp host (in-process via the llama-cpp-rs bindings — no llama-server subprocess). We cannot use any of llama.cpp's speculative-decoding machinery (including the new DeepSeek V4 DSpark path from #25784) because it lives entirely in the
common/layer, which is tied to llama-server's main loop.The gap
As of
596a579(and latest master), the corellama.hAPI exposes no speculative params:llama_model_paramshas nodraft_model/ spec fields.llama_context_paramshas no spec fields.llama_speculative_*init function in the public API.DSpark (and MTP) support is implemented in
common/speculative.cpp/ the server layer only. A library consumer that links llama.cpp and drivesllama_decodedirectly cannot:--spec-type draft-dsparkstyle speculation, orAsk
A minimal core-API speculative surface for in-process hosts, e.g.:
llama_model_params.draft_model(+ spec type / draft-n-max) re-added, ORllama_speculative_*API that manages draft+target contexts and exposes allama_decode_speculative-style call.We (and likely other language bindings that bypass llama-server) would be happy to test, provide repro cases, or contribute to a design/implementation. We already drive the bundled llama.cpp in production for DeepSeek V4 Flash 0731 (284B/13B MoE,
deepseek4arch) and are currently forced to use our own draft→verify loop, which cannot use DSpark's confidence head.Context
Thanks for considering — happy to help move it forward.