Repository navigation
chat: add new template for DeepSeek V4 Flash 0731 - #26398
Conversation
Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change. - Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present. - Add structured output response-format instructions to the V4 templates and pass the schema into template rendering. - Add a separate Flash 0731 template for the updated high and max reasoning effort mapping. - Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests. Official references: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py Assisted-by: Codex
| }); | ||
| }, | ||
| [&](context & ctx) { | ||
| ctx.set_val("enable_thinking", mk_val<value_bool>(true)); |
There was a problem hiding this comment.
this makes sense to me, but I can revert if not desired
Try my quants: https://huggingface.co/tarruda/DeepSeek-V4-Flash-0731-GGUF And compare with the official deepseek API |
|
@ionizing-plasma TBH this thinking style doesn't seem to be 100% consistent, a few times I got standard short reasoning even though I've set the |
|
@ionizing-plasma it seems to be possible to steer the model into that way of thinking by customizing the template. I've created a version that adds Clearly that is not part of official encoding so won't be adding to this PR, but I saw great performance out of Nex N2 when using that thinking style, so I will be playing with it locally. |
There was a problem hiding this comment.
the conversion script automatically pick up the original template, not the new 0731, would you mind having a look to see the conversion need to be updated too? (i.e. switch the file based on input model name)
There was a problem hiding this comment.
Since the model architecture is the same, should I just use "0731" being present in the model name?
The OpenAI "reasoning_effort" body field was parsed but only "none" was ever
acted on; every other level was dropped with the comment "model-specific and
not yet handled". Templates that DO implement effort levels therefore never
saw the value, and the knob was inert over the OAI API -- the only way to
reach it was the non-standard chat_template_kwargs escape hatch.
Forward the remaining levels verbatim into chat_template_kwargs and let the
template decide what they mean. Templates that ignore the variable are
unaffected, "none" keeps its existing enable_thinking=false meaning, and an
explicit chat_template_kwargs entry still wins.
Measured on DeepSeek-V4-Flash-0731 (UD-IQ3_XXS), rendered prompt length:
before low/high/max/unset all 50 chars -- completely inert
after unset/low 50 (template acts on high|max only)
high 526 (+ "Reasoning Effort: Absolute maximum ..." block)
max 576 (+ "Reasoning Effort: Beyond maximum ..." block)
none 51 (thinking disabled, unchanged behaviour)
and generated reasoning grows 217 -> 318 chars from default to max.
Note this is NOT the fix from upstream ggml-org#26398, which was the starting point:
that PR repairs llama.cpp's own bundled DSv4 jinja templates. The template we
serve is the one embedded in the GGUF, which already implements the effort
levels correctly -- so the bundled-template changes are redundant here and the
real defect was one layer down, in the server's body parsing. Verified against
the live server that the other two defects ggml-org#26398 lists (structured output,
drop_thinking default) do not reproduce on this template.
Test asserts on the plumbing rather than any model's prompt wording: a jinja
template that echoes reasoning_effort, covering the OAI field reaching the
template, absent staying absent, "none" not being treated as a level, and an
explicit kwarg overriding the field.
|
Has anyone else noticed that DSv4 Flash has to reprocess its output for every subsequent agentic turn? If it thinks for 5000 tokens then calls a tool, the next turn will also have 5000 additional tokens it needs to process on prefill. I found a way to fix this by modifying the chat template. This diff is against the chat template that Unsloth includes with their GGUFs: The issue was an extra I believe the template in this PR will suffer the same reprocessing issue unless the fix is applied. (Someone could also fix the parser instead of the template, perhaps? I'm not sure what the best approach is here.) |
That does not happen with this template |
|
Ok... so now I've spent 30 minutes to reproduce it with this PR and show that it does happen with this PR and this exact template. repro.sh (I had to rename it to a .txt extension to attach it to this comment) Running against a $ export SERVER_BIN=./llama-server
$ export MODEL=./deepseek-v4-flash-0731-ud-q2_k_xl.gguf
$ export TEMPLATE=./deepseek-ai-DeepSeek-V4-Flash-0731.jinja
$ ./repro.sh
Starting llama-server...
Turn 1: generating a tool call...
{
"finish_reason": "tool_calls",
"assistant_content": "\n\n",
"prompt_tokens": 343,
"completion_tokens": 241
}
Turn 2: returning the tool result...
{
"reproduced": true,
"first_completion_tokens": 241,
"second_prompt_tokens": 606,
"expected_cached_history_tokens": 583,
"actual_cached_tokens": 339,
"avoidably_reprocessed_history_tokens": 244,
"genuinely_new_turn_tokens": 23,
"total_reprocessed_tokens": 267,
"prompt_ms": 972.062,
"response": "OK"
}So, it wasted the processing time to reprocess 244 tokens on the second turn. Running with the template I provided above: $ export TEMPLATE=./deepseek-v4-flash-0731-chat-template.jinja
$ ./repro.sh
Starting llama-server...
Turn 1: generating a tool call...
{
"finish_reason": "tool_calls",
"assistant_content": "\n\n",
"prompt_tokens": 343,
"completion_tokens": 241
}
Turn 2: returning the tool result...
{
"reproduced": false,
"first_completion_tokens": 241,
"second_prompt_tokens": 606,
"expected_cached_history_tokens": 583,
"actual_cached_tokens": 583,
"avoidably_reprocessed_history_tokens": 0,
"genuinely_new_turn_tokens": 23,
"total_reprocessed_tokens": 23,
"prompt_ms": 288.51,
"response": "OK"
}No wasted tokens. There is an issue with the template in this PR. |
|
Fundamentally, the issue is that the model seems to be outputting two newlines which the parser isn't catching, which is what we see in |
|
@coder543 dsv4 is using a manual parser, I'm gonna try to fix that. Thanks for catching. |
aldehir
left a comment
There was a problem hiding this comment.
Parser edits are fine, but it seems this could also be handled with | trim, which is common across templates.
|
@coder543 can you verify the parser updates address the cache issue. |
|
@aldehir yes, the round trip rendering error is fixed now! |
|
@coder543 thanks for validating it! |
|
When we merge this, how is the new chat template going to be propagated to the GGUFs? |
|
@ggerganov it will have to be re-embedded into existing GGUFs (gguf-new-metadata). New GGUFs should pick it up automatically. Note that this is not a critical template fix, main things it enables is supporting reasoning levels and adding the prompt DSv4 was trained to emit structured output. But the later bugfix caught by @coder543 will actually benefit existing GGUFs. |
I'm not sure how that works. I thought when converting the GGUFs with |
Deepseek v4 doesn't have a jinja template. They use a specialized parsing/encoding python module: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py. Note very familiar with convert_hf_to_gguf.py, but I imagine it picks up a template from llama.cpp repo if the official repo doesn't have a template. |
|
I already tested this by converting safetensors with this branch and confirmed the new template is picked up. |
| model_id_hint = self.remote_hf_model_id or self.dir_model.name | ||
| is_0731 = "0731" in model_id_hint | ||
| template_name = "deepseek-ai-DeepSeek-V4-Flash-0731.jinja" if is_0731 else "deepseek-ai-DeepSeek-V4.jinja" | ||
| template_path = Path(__file__).parent.parent / "models" / "templates" / template_name |
|
@coder543 it seems the prompt reprocessing issue is still there, I saw this running locally:
That is: the model generated close to 60k tokens of thinking, then it did a tool call. When processing the tool call, it had to reprocess the entire prompt. Going to see if I can reproduce consistently and come up with a fix. |
|
@tarruda Which client do you used? Btw, you should set these envs for some extra logs: export LLAMA_TRACE=1
export LLAMA_SERVER_SLOTS_DEBUG=1 |
|
I'm actually using my own client, so it could be a bug on it too: https://github.com/tarruda/neoagent I will turn those llama traces on, but will also enable request logging on my client so I can compare individual requests to see if it changes something in the payload prefix. |
|
Interesting. I had run my repro.sh script on the latest version of this branch, and it reported no issue anymore. On my DGX Spark, I get better performance using a fork of |
|
The issue was pretty consistent across all clients I tested (mine, codex and pi) and it always happened in streamed thinking blocks. I believe the main reason it was not noticed before is that most times it doesn't think long enough before a tool call for the extra time to be noticeable. This time it reasoned for 60k tokens, and given the slow promt processing on my mac, it was easy to spot. Apparently it was a tokenization bug. I've pushed a fix to my branch: 7018277 Still not sure because I used deepseek v4 flash itself to assist in debugging this and it provided the fix plus this summary: The patch looks like it fixed for me. Will use for a week to see if there are any side effects and can create a PR if that looks right for @ggerganov. |
17f390a でパーサ検証用に追加した Unsloth 版 DeepSeek-V4-Flash テンプ レートのスナップショットだが、llama-server/llama-cli のコードからも テストからも一切参照されない孤立ファイルだった (--chat-template-file での明示指定なし、既定探索の対象でもない)。 本家に models/templates/deepseek-ai-DeepSeek-V4-Flash-0731.jinja が 存在し (PR ggml-org#26398)、パーサ検証が必要な場合はこちらで代替できるため fork 独自のスナップショットを削除する。
* common/chat: update DeepSeek V4 templates Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change. - Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present. - Add structured output response-format instructions to the V4 templates and pass the schema into template rendering. - Add a separate Flash 0731 template for the updated high and max reasoning effort mapping. - Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests. Official references: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py Assisted-by: Codex * Fix deepseek v4 0731 template selection * remove unneeded lower normalization * Fix DSML parser to consume the tool call separator * address aldehir requests * address aldehir comment
* common/chat: update DeepSeek V4 templates Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change. - Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present. - Add structured output response-format instructions to the V4 templates and pass the schema into template rendering. - Add a separate Flash 0731 template for the updated high and max reasoning effort mapping. - Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests. Official references: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py Assisted-by: Codex * Fix deepseek v4 0731 template selection * remove unneeded lower normalization * Fix DSML parser to consume the tool call separator * address aldehir requests * address aldehir comment
* common/chat: update DeepSeek V4 templates Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change. - Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present. - Add structured output response-format instructions to the V4 templates and pass the schema into template rendering. - Add a separate Flash 0731 template for the updated high and max reasoning effort mapping. - Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests. Official references: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py Assisted-by: Codex * Fix deepseek v4 0731 template selection * remove unneeded lower normalization * Fix DSML parser to consume the tool call separator * address aldehir requests * address aldehir comment
* common/chat: update DeepSeek V4 templates Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change. - Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present. - Add structured output response-format instructions to the V4 templates and pass the schema into template rendering. - Add a separate Flash 0731 template for the updated high and max reasoning effort mapping. - Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests. Official references: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py Assisted-by: Codex * Fix deepseek v4 0731 template selection * remove unneeded lower normalization * Fix DSML parser to consume the tool call separator * address aldehir requests * address aldehir comment
* common/chat: update DeepSeek V4 templates Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change. - Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present. - Add structured output response-format instructions to the V4 templates and pass the schema into template rendering. - Add a separate Flash 0731 template for the updated high and max reasoning effort mapping. - Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests. Official references: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py Assisted-by: Codex * Fix deepseek v4 0731 template selection * remove unneeded lower normalization * Fix DSML parser to consume the tool call separator * address aldehir requests * address aldehir comment
* common/chat: update DeepSeek V4 templates Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change. - Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present. - Add structured output response-format instructions to the V4 templates and pass the schema into template rendering. - Add a separate Flash 0731 template for the updated high and max reasoning effort mapping. - Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests. Official references: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py Assisted-by: Codex * Fix deepseek v4 0731 template selection * remove unneeded lower normalization * Fix DSML parser to consume the tool call separator * address aldehir requests * address aldehir comment



Overview
Fix the existing dsv4 preview template to match the reference encoding and add new template for the 0731 which uses different prompts for reasoning levels.
Additional information
References for these changes:
Requirements
cc @pwilkin @aldehir @am17an