Skip to content

chat: add new template for DeepSeek V4 Flash 0731 - #26398

Merged
aldehir merged 6 commits into
ggml-org:masterfrom
tarruda:dsv4-template-fixes
Aug 3, 2026
Merged

aldehir merged 6 commits into
ggml-org:masterfrom
tarruda:dsv4-template-fixes

Conversation

@tarruda

@tarruda tarruda commented Aug 1, 2026

Copy link
Copy Markdown
Contributor

Overview

Fix the existing dsv4 preview template to match the reference encoding and add new template for the 0731 which uses different prompts for reasoning levels.

Additional information

  • Fix the existing dsv4 preview template to include handling for max effort reasoning, structured output and the default value of drop_thinking to True (which is the default behavior in deepseek encoding).
  • Add new template for 0731, which only changes the prompts used for triggering max effort reasoning and now has an explicit "high" setting.
  • Also added tests

References for these changes:

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes I used codex to assist, but I reviewed the changes and the official encoding files. Also have been testing these changes for weeks in the preview dsv4 and started testing the new 0731.

cc @pwilkin @aldehir @am17an

Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change.

- Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present.
- Add structured output response-format instructions to the V4 templates and pass the schema into template rendering.
- Add a separate Flash 0731 template for the updated high and max reasoning effort mapping.
- Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests.

Official references:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py

Assisted-by: Codex
@tarruda
tarruda requested review from a team, CISC and pwilkin as code owners August 1, 2026 10:55
@github-actions github-actions Bot added testing Everything test related jinja parser Issues related to the jinja parser labels Aug 1, 2026
Comment thread common/jinja/caps.cpp
});
},
[&](context & ctx) {
ctx.set_val("enable_thinking", mk_val<value_bool>(true));

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this makes sense to me, but I can revert if not desired

@tarruda

tarruda commented Aug 1, 2026

Copy link
Copy Markdown
Contributor Author

It seems like the "max effort" reasoning in DSv4 0731 is much more verbose than in the preview version. I also noticed an interesting pattern: DSv4 now has "caveman reasoning" similar to Nex N2:

image

@aa956 aa956 mentioned this pull request Aug 1, 2026
11 of 13 tasks
@ionizing-plasma

Copy link
Copy Markdown

It seems like the "max effort" reasoning in DSv4 0731 is much more verbose than in the preview version. I also noticed an interesting pattern: DSv4 now has "caveman reasoning" similar to Nex N2:

I thought maybe it was how you worded the input and tested the same prompt, and did not get back caveman reasoning traces (this is with the unsloth IQ4_XS):

ds4_thinking_tetris

@tarruda

tarruda commented Aug 1, 2026

Copy link
Copy Markdown
Contributor Author

I thought maybe it was how you worded the input and tested the same prompt, and did not get back caveman reasoning traces (this is with the unsloth IQ4_XS):

Try my quants: https://huggingface.co/tarruda/DeepSeek-V4-Flash-0731-GGUF

And compare with the official deepseek API

@tarruda

tarruda commented Aug 1, 2026 •

Copy link
Copy Markdown
Contributor Author

@ionizing-plasma TBH this thinking style doesn't seem to be 100% consistent, a few times I got standard short reasoning even though I've set the reasoning_effort = max (but this also happens in the official API)

@tarruda

tarruda commented Aug 1, 2026

Copy link
Copy Markdown
Contributor Author

@ionizing-plasma it seems to be possible to steer the model into that way of thinking by customizing the template. I've created a version that adds We need answer right after <think> in high/max reasoning: https://gist.github.com/tarruda/d60797a81277d781f5e70209e007cc34

Clearly that is not part of official encoding so won't be adding to this PR, but I saw great performance out of Nex N2 when using that thinking style, so I will be playing with it locally.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the conversion script automatically pick up the original template, not the new 0731, would you mind having a look to see the conversion need to be updated too? (i.e. switch the file based on input model name)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the model architecture is the same, should I just use "0731" being present in the model name?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

TrevorS added a commit to TrevorS/llama.cpp that referenced this pull request Aug 2, 2026
The OpenAI "reasoning_effort" body field was parsed but only "none" was ever
acted on; every other level was dropped with the comment "model-specific and
not yet handled". Templates that DO implement effort levels therefore never
saw the value, and the knob was inert over the OAI API -- the only way to
reach it was the non-standard chat_template_kwargs escape hatch.

Forward the remaining levels verbatim into chat_template_kwargs and let the
template decide what they mean. Templates that ignore the variable are
unaffected, "none" keeps its existing enable_thinking=false meaning, and an
explicit chat_template_kwargs entry still wins.

Measured on DeepSeek-V4-Flash-0731 (UD-IQ3_XXS), rendered prompt length:

  before   low/high/max/unset all 50 chars -- completely inert
  after    unset/low 50 (template acts on high|max only)
           high 526  (+ "Reasoning Effort: Absolute maximum ..." block)
           max  576  (+ "Reasoning Effort: Beyond maximum ..."  block)
           none 51   (thinking disabled, unchanged behaviour)

and generated reasoning grows 217 -> 318 chars from default to max.

Note this is NOT the fix from upstream ggml-org#26398, which was the starting point:
that PR repairs llama.cpp's own bundled DSv4 jinja templates. The template we
serve is the one embedded in the GGUF, which already implements the effort
levels correctly -- so the bundled-template changes are redundant here and the
real defect was one layer down, in the server's body parsing. Verified against
the live server that the other two defects ggml-org#26398 lists (structured output,
drop_thinking default) do not reproduce on this template.

Test asserts on the plumbing rather than any model's prompt wording: a jinja
template that echoes reasoning_effort, covering the OAI field reaching the
template, absent staying absent, "none" not being treated as a level, and an
explicit kwarg overriding the field.
@coder543

coder543 commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

Has anyone else noticed that DSv4 Flash has to reprocess its output for every subsequent agentic turn? If it thinks for 5000 tokens then calls a tool, the next turn will also have 5000 additional tokens it needs to process on prefill.

I found a way to fix this by modifying the chat template. This diff is against the chat template that Unsloth includes with their GGUFs:

215c215,220
<       {{- '\n\n<' + dsml_token + 'tool_calls>\n' -}}
---
>       {#- llama.cpp preserves the DSML separator in parsed content; do not add it twice. -#}
>       {%- set content_has_separator = message['content'] is defined and message['content'] is string and message['content'].endswith('\n\n') -%}
>       {%- if not content_has_separator -%}
>         {{- '\n\n' -}}
>       {%- endif -%}
>       {{- '<' + dsml_token + 'tool_calls>\n' -}}

The issue was an extra \n\n that would get rendered during agentic turns, causing large amounts of wasted additional prefill in my testing. This modification to the template avoids duplicating the \n\n separator.

I believe the template in this PR will suffer the same reprocessing issue unless the fix is applied. (Someone could also fix the parser instead of the template, perhaps? I'm not sure what the best approach is here.)

@tarruda

tarruda commented Aug 2, 2026

Copy link
Copy Markdown
Contributor Author

Has anyone else noticed that DSv4 Flash has to reprocess its output for every subsequent agentic turn?

That does not happen with this template

@coder543

coder543 commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor

Ok... so now I've spent 30 minutes to reproduce it with this PR and show that it does happen with this PR and this exact template.

repro.sh
deepseek-v4-flash-0731-chat-template.txt

(I had to rename it to a .txt extension to attach it to this comment)

Running against a llama-server binary that is built from this PR using this PR's template:

$ export SERVER_BIN=./llama-server
$ export MODEL=./deepseek-v4-flash-0731-ud-q2_k_xl.gguf
$ export TEMPLATE=./deepseek-ai-DeepSeek-V4-Flash-0731.jinja
$ ./repro.sh
Starting llama-server...
Turn 1: generating a tool call...
{
  "finish_reason": "tool_calls",
  "assistant_content": "\n\n",
  "prompt_tokens": 343,
  "completion_tokens": 241
}
Turn 2: returning the tool result...
{
  "reproduced": true,
  "first_completion_tokens": 241,
  "second_prompt_tokens": 606,
  "expected_cached_history_tokens": 583,
  "actual_cached_tokens": 339,
  "avoidably_reprocessed_history_tokens": 244,
  "genuinely_new_turn_tokens": 23,
  "total_reprocessed_tokens": 267,
  "prompt_ms": 972.062,
  "response": "OK"
}

So, it wasted the processing time to reprocess 244 tokens on the second turn.

Running with the template I provided above:

$ export TEMPLATE=./deepseek-v4-flash-0731-chat-template.jinja
$ ./repro.sh
Starting llama-server...
Turn 1: generating a tool call...
{
  "finish_reason": "tool_calls",
  "assistant_content": "\n\n",
  "prompt_tokens": 343,
  "completion_tokens": 241
}
Turn 2: returning the tool result...
{
  "reproduced": false,
  "first_completion_tokens": 241,
  "second_prompt_tokens": 606,
  "expected_cached_history_tokens": 583,
  "actual_cached_tokens": 583,
  "avoidably_reprocessed_history_tokens": 0,
  "genuinely_new_turn_tokens": 23,
  "total_reprocessed_tokens": 23,
  "prompt_ms": 288.51,
  "response": "OK"
}

No wasted tokens.

There is an issue with the template in this PR.

@coder543

coder543 commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

Fundamentally, the issue is that the model seems to be outputting two newlines which the parser isn't catching, which is what we see in assistant_content in the repro, then the template reinserts those when the turn is rendered the next time around. This causes four newlines, and invalidates the cache from that point on. Someone could fix the parser, or someone make the template more robust. I don't know if DSv4 is using the autoparser or a custom parser. I fixed the template on my machine to avoid the issue, but I'm not sure what the preferred fix is here.

@tarruda

tarruda commented Aug 3, 2026

Copy link
Copy Markdown
Contributor Author

@coder543 dsv4 is using a manual parser, I'm gonna try to fix that. Thanks for catching.

@tarruda

tarruda commented Aug 3, 2026

Copy link
Copy Markdown
Contributor Author

@coder543 a8aa9d4

@aldehir aldehir left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Parser edits are fine, but it seems this could also be handled with | trim, which is common across templates.

Comment thread common/chat.cpp Outdated
Comment thread common/chat.cpp Outdated
Comment thread tests/test-chat.cpp Outdated
Comment thread tests/test-chat.cpp Outdated
Comment thread common/chat.cpp Outdated
@aldehir

aldehir commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

@coder543 can you verify the parser updates address the cache issue.

@coder543

coder543 commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

@aldehir yes, the round trip rendering error is fixed now!

@tarruda

tarruda commented Aug 3, 2026

Copy link
Copy Markdown
Contributor Author

@coder543 thanks for validating it!

@ggerganov

Copy link
Copy Markdown
Member

When we merge this, how is the new chat template going to be propagated to the GGUFs?

@tarruda

tarruda commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor Author

@ggerganov it will have to be re-embedded into existing GGUFs (gguf-new-metadata). New GGUFs should pick it up automatically.

Note that this is not a critical template fix, main things it enables is supporting reasoning levels and adding the prompt DSv4 was trained to emit structured output.

But the later bugfix caught by @coder543 will actually benefit existing GGUFs.

@ggerganov

Copy link
Copy Markdown
Member

New GGUFs should pick it up automatically.

I'm not sure how that works. I thought when converting the GGUFs with convert_hf_to_gguf.py it picks up the template from the source transformers repo?

@tarruda

tarruda commented Aug 3, 2026

Copy link
Copy Markdown
Contributor Author

it picks up the template from the source transformers repo?

Deepseek v4 doesn't have a jinja template. They use a specialized parsing/encoding python module: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py.

Note very familiar with convert_hf_to_gguf.py, but I imagine it picks up a template from llama.cpp repo if the official repo doesn't have a template.

@tarruda

tarruda commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor Author

I already tested this by converting safetensors with this branch and confirmed the new template is picked up.

Comment thread conversion/deepseek.py
Comment on lines +502 to +505
model_id_hint = self.remote_hf_model_id or self.dir_model.name
is_0731 = "0731" in model_id_hint
template_name = "deepseek-ai-DeepSeek-V4-Flash-0731.jinja" if is_0731 else "deepseek-ai-DeepSeek-V4.jinja"
template_path = Path(__file__).parent.parent / "models" / "templates" / template_name

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, I see it now. Thanks

@aldehir
aldehir merged commit 0ef6e55 into ggml-org:master Aug 3, 2026
24 of 28 checks passed
@tarruda
tarruda deleted the dsv4-template-fixes branch August 3, 2026 23:01
@tarruda

tarruda commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

@coder543 it seems the prompt reprocessing issue is still there, I saw this running locally:

image

That is: the model generated close to 60k tokens of thinking, then it did a tool call. When processing the tool call, it had to reprocess the entire prompt.

Going to see if I can reproduce consistently and come up with a fix.

@ggerganov

Copy link
Copy Markdown
Member

@tarruda Which client do you used?

Btw, you should set these envs for some extra logs:

export LLAMA_TRACE=1
export LLAMA_SERVER_SLOTS_DEBUG=1

@tarruda

tarruda commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

I'm actually using my own client, so it could be a bug on it too: https://github.com/tarruda/neoagent

I will turn those llama traces on, but will also enable request logging on my client so I can compare individual requests to see if it changes something in the payload prefix.

@coder543

coder543 commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Interesting. I had run my repro.sh script on the latest version of this branch, and it reported no issue anymore. On my DGX Spark, I get better performance using a fork of ds4, so that’s what I’ve been playing with the past few days, so I haven’t tested in practice whether the issue is still recurring on llama-server.

@tarruda

tarruda commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

The issue was pretty consistent across all clients I tested (mine, codex and pi) and it always happened in streamed thinking blocks. I believe the main reason it was not noticed before is that most times it doesn't think long enough before a tool call for the extra time to be noticeable. This time it reasoned for 60k tokens, and given the slow promt processing on my mac, it was easy to spot.

Apparently it was a tokenization bug. I've pushed a fix to my branch: 7018277

Still not sure because I used deepseek v4 flash itself to assist in debugging this and it provided the fix plus this summary:

## Summary

With the DeepSeek V4 Flash 0731 template, a coding harness can observe the
following behavior:

1. The model generates a large `<think>` block and then a tool call.
2. The harness executes the tool and sends the result back.
3. llama-server reprocesses the thinking block on the second turn instead of
   reusing the KV cache prefix.

The root cause is a tokenization mismatch, not a template or parser bug:

- During turn 1, the model samples token `6883` (`Ġ~`, the BPE encoding of
  ` ~`) inside the thinking block, and that token ID is appended to the slot
  context.
- On turn 2, llama.cpp re-tokenizes the same text and produces `223` (`Ġ`) +
  `96` (`~`) instead of `6883`.
- KV cache reuse compares token IDs, so the common prefix stops at the first
  divergent token and everything after it (the rest of the thinking block) is
  reprocessed.

The patch looks like it fixed for me. Will use for a week to see if there are any side effects and can create a PR if that looks right for @ggerganov.

ishikawa added a commit to ishikawa/llama.cpp that referenced this pull request Aug 9, 2026
17f390a でパーサ検証用に追加した Unsloth 版 DeepSeek-V4-Flash テンプ
レートのスナップショットだが、llama-server/llama-cli のコードからも
テストからも一切参照されない孤立ファイルだった (--chat-template-file
での明示指定なし、既定探索の対象でもない)。

本家に models/templates/deepseek-ai-DeepSeek-V4-Flash-0731.jinja が
存在し (PR ggml-org#26398)、パーサ検証が必要な場合はこちらで代替できるため
fork 独自のスナップショットを削除する。
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
* common/chat: update DeepSeek V4 templates

Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change.

- Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present.
- Add structured output response-format instructions to the V4 templates and pass the schema into template rendering.
- Add a separate Flash 0731 template for the updated high and max reasoning effort mapping.
- Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests.

Official references:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py

Assisted-by: Codex

* Fix deepseek v4 0731 template selection

* remove unneeded lower normalization

* Fix DSML parser to consume the tool call separator

* address aldehir requests

* address aldehir comment
brittlewis12 pushed a commit to brittlewis12/llama.cpp that referenced this pull request Aug 17, 2026
* common/chat: update DeepSeek V4 templates

Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change.

- Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present.
- Add structured output response-format instructions to the V4 templates and pass the schema into template rendering.
- Add a separate Flash 0731 template for the updated high and max reasoning effort mapping.
- Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests.

Official references:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py

Assisted-by: Codex

* Fix deepseek v4 0731 template selection

* remove unneeded lower normalization

* Fix DSML parser to consume the tool call separator

* address aldehir requests

* address aldehir comment
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
* common/chat: update DeepSeek V4 templates

Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change.

- Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present.
- Add structured output response-format instructions to the V4 templates and pass the schema into template rendering.
- Add a separate Flash 0731 template for the updated high and max reasoning effort mapping.
- Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests.

Official references:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py

Assisted-by: Codex

* Fix deepseek v4 0731 template selection

* remove unneeded lower normalization

* Fix DSML parser to consume the tool call separator

* address aldehir requests

* address aldehir comment
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* common/chat: update DeepSeek V4 templates

Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change.

- Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present.
- Add structured output response-format instructions to the V4 templates and pass the schema into template rendering.
- Add a separate Flash 0731 template for the updated high and max reasoning effort mapping.
- Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests.

Official references:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py

Assisted-by: Codex

* Fix deepseek v4 0731 template selection

* remove unneeded lower normalization

* Fix DSML parser to consume the tool call separator

* address aldehir requests

* address aldehir comment
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* common/chat: update DeepSeek V4 templates

Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change.

- Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present.
- Add structured output response-format instructions to the V4 templates and pass the schema into template rendering.
- Add a separate Flash 0731 template for the updated high and max reasoning effort mapping.
- Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests.

Official references:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py

Assisted-by: Codex

* Fix deepseek v4 0731 template selection

* remove unneeded lower normalization

* Fix DSML parser to consume the tool call separator

* address aldehir requests

* address aldehir comment
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* common/chat: update DeepSeek V4 templates

Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change.

- Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present.
- Add structured output response-format instructions to the V4 templates and pass the schema into template rendering.
- Add a separate Flash 0731 template for the updated high and max reasoning effort mapping.
- Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests.

Official references:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.py
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py

Assisted-by: Codex

* Fix deepseek v4 0731 template selection

* remove unneeded lower normalization

* Fix DSML parser to consume the tool call separator

* address aldehir requests

* address aldehir comment
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion jinja parser Issues related to the jinja parser testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants