Skip to content

chat : fix reasoning leak with force-opened bare <think> templates - #24674

Merged
pwilkin merged 2 commits into
ggml-org:masterfrom
newjordan:chat-fix-forced-open-reasoning-leak
Jul 13, 2026
Merged

pwilkin merged 2 commits into
ggml-org:masterfrom
newjordan:chat-fix-forced-open-reasoning-leak

Conversation

@newjordan

@newjordan newjordan commented Jun 16, 2026 •

Copy link
Copy Markdown
Contributor

Fixes a reasoning leak in the PEG chat auto-parser for templates that force-open thinking by prefilling a bare <think> (e.g. Nex-N2-mini). The reasoning start tag inferred from prior turns is <think>\n, which the prefix split fails to find in the bare-<think> generation prompt, so the tag is swallowed and the reasoning trace lands in content instead of reasoning_content.

Trims the start tag used for the prefix split so the bare <think> is matched (per @aldehir's suggestion).

@newjordan
newjordan requested review from a team and pwilkin as code owners June 16, 2026 01:29
@github-actions github-actions Bot added the testing Everything test related label Jun 16, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Jun 16, 2026

Copy link
Copy Markdown

Hi @newjordan, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

  • AI-generated content: This project does not accept PRs, descriptions or commit messages that are fully or predominantly AI-generated. If you have used AI to assist you in writing code, please make sure to disclose that explicitly.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@newjordan

Copy link
Copy Markdown
Contributor Author

Hey there, this is a bug fix on soemthing we found from testing that is upstream of our model that can affect models with the think vs think>n - there is no lazy or gracefu ldefualt there and some models break on loading for users on the Nex2 - this was a deliberate PR, I use AI, but this decision to send this PR was hopefully a courtesy to LLama and I sincrely do not wish to produce noise or bad work.

@tarruda

tarruda commented Jun 20, 2026

Copy link
Copy Markdown
Contributor

I confirm that this fixes Nex N2 for me.

There's a workaround in the chat template, but changing the template can affect the final performance.

Quoting one of N2 devs (https://huggingface.co/nex-agi/Nex-N2-mini/discussions/4#6a2b6c0e143fc2731e78c958):

After investigation, the root cause turned out to be in llama.cpp's reasoning parser, not the template.

Adding \n after does work around it, but the model was trained strictly on the current template, so deviating from it at inference time may hurt output quality. We'd rather keep the template as-is."

So this fix appears to be necessary for N2 to work properly in llama.cpp

@aldehir

aldehir commented Jun 20, 2026

Copy link
Copy Markdown
Contributor

This should be sufficient for this issue:

diff --git a/common/chat-auto-parser-generator.cpp b/common/chat-auto-parser-generator.cpp
index 37ca55c..53bc85c 100644
--- a/common/chat-auto-parser-generator.cpp
+++ b/common/chat-auto-parser-generator.cpp
@@ -147,7 +147,8 @@ common_peg_arena autoparser::build_parser(const generation_params & inputs, cons
         } else {
             parser = content.build_parser(ctx);
         }
-        return pure_content ? p.prefix(generation_prompt, reasoning.start) + parser : p.prefix(generation_prompt, reasoning.start) << parser;
+        const std::string reasoning_start = trim_whitespace(reasoning.start);
+        return pure_content ? p.prefix(generation_prompt, reasoning_start) + parser : p.prefix(generation_prompt, reasoning_start) << parser;
     });
 }

@tarruda

tarruda commented Jun 20, 2026

Copy link
Copy Markdown
Contributor

@newjordan can you push the simplified version suggested by @aldehir ?

@nenkoru

nenkoru commented Jun 20, 2026

Copy link
Copy Markdown

Applied the patch proposed by @aldehir. Worked fine with the Nex-N2-Mini model.

Reasoning is adaptive for this model. So I was able to force it to produce thinking by a logic task[1] in open-webui. No custom templates needed. Also works as expected in OpenCode.

[1]

Let (S = {1, 2, 3, 4, 5, 6, 7, 8, 9, 10}). We want to find the number of subsets of (S) that do not contain any two consecutive integers. Find the total number of such subsets (including the empty set).

@newjordan

Copy link
Copy Markdown
Contributor Author

Hey guys,

Yes I will PR now. However, I did some work and ran into this on exact implementation -

regresses an existing case: test-chat aborts on
▎ NVIDIA-Nemotron-Nano-v2.jinja with a pure tool call (…), leaking a bare into content (expected empty)

Sorry for being obtuse, I try and do a lot of testing before sending anything over and I keep getting an error that if I submit just the short fix, it breaks a nemo config upstream. I will submit just the PR as requested, with this as a notation in case it is an issue. I have the fix to this also prepared.

@newjordan
newjordan force-pushed the chat-fix-forced-open-reasoning-leak branch from c68c44b to 69dc863 Compare June 21, 2026 02:23
The reasoning start tag inferred from prior turns can carry trailing
whitespace (e.g. <think>\n) while a force-open template prefills a bare
<think>. Trim the tag used for the prefix split so the bare prefill is
matched instead of being swallowed into content.
@newjordan
newjordan force-pushed the chat-fix-forced-open-reasoning-leak branch from 69dc863 to 461630c Compare June 21, 2026 02:32
@aldehir

aldehir commented Jun 21, 2026

Copy link
Copy Markdown
Contributor

Sorry for being obtuse, I try and do a lot of testing before sending anything over and I keep getting an error that if I submit just the short fix, it breaks a nemo config upstream. I will submit just the PR as requested, with this as a notation in case it is an issue. I have the fix to this also prepared.

I'll look into it, but it seems to have surfaced an existing bug with Nemotron Nano v2.

@aldehir aldehir left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should be good to go.

cc @pwilkin for review

@tarruda

tarruda commented Jul 12, 2026

Copy link
Copy Markdown
Contributor

While this is not merged, I've uploaded a chat template workaround that "fools" the autoparser into using the correct delimiter: https://huggingface.co/tarruda/Nex-N2-Pro-GGUF/blob/main/chat_template.jinja#L102-L107

@pwilkin pwilkin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All right, since we already do << for the prefix, this won't change semantics, so it should be fine.

@pwilkin
pwilkin merged commit 91c631b into ggml-org:master Jul 13, 2026
25 checks passed
RehanQasim-dev pushed a commit to aifoundry-org/llama.cpp that referenced this pull request Jul 23, 2026
…gml-org#24674)

* chat : fix reasoning leak with force-opened bare <think> templates

The reasoning start tag inferred from prior turns can carry trailing
whitespace (e.g. <think>\n) while a force-open template prefills a bare
<think>. Trim the tag used for the prefix split so the bare prefill is
matched instead of being swallowed into content.

* chat : fix Nemotron Nano v2 regression

---------

Co-authored-by: Alde Rojas <hello@alde.dev>
RehanQasim-dev pushed a commit to aifoundry-org/llama.cpp that referenced this pull request Jul 23, 2026
…gml-org#24674)

* chat : fix reasoning leak with force-opened bare <think> templates

The reasoning start tag inferred from prior turns can carry trailing
whitespace (e.g. <think>\n) while a force-open template prefills a bare
<think>. Trim the tag used for the prefix split so the bare prefill is
matched instead of being swallowed into content.

* chat : fix Nemotron Nano v2 regression

---------

Co-authored-by: Alde Rojas <hello@alde.dev>
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
…gml-org#24674)

* chat : fix reasoning leak with force-opened bare <think> templates

The reasoning start tag inferred from prior turns can carry trailing
whitespace (e.g. <think>\n) while a force-open template prefills a bare
<think>. Trim the tag used for the prefix split so the bare prefill is
matched instead of being swallowed into content.

* chat : fix Nemotron Nano v2 regression

---------

Co-authored-by: Alde Rojas <hello@alde.dev>
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
…gml-org#24674)

* chat : fix reasoning leak with force-opened bare <think> templates

The reasoning start tag inferred from prior turns can carry trailing
whitespace (e.g. <think>\n) while a force-open template prefills a bare
<think>. Trim the tag used for the prefix split so the bare prefill is
matched instead of being swallowed into content.

* chat : fix Nemotron Nano v2 regression

---------

Co-authored-by: Alde Rojas <hello@alde.dev>
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…gml-org#24674)

* chat : fix reasoning leak with force-opened bare <think> templates

The reasoning start tag inferred from prior turns can carry trailing
whitespace (e.g. <think>\n) while a force-open template prefills a bare
<think>. Trim the tag used for the prefix split so the bare prefill is
matched instead of being swallowed into content.

* chat : fix Nemotron Nano v2 regression

---------

Co-authored-by: Alde Rojas <hello@alde.dev>
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…gml-org#24674)

* chat : fix reasoning leak with force-opened bare <think> templates

The reasoning start tag inferred from prior turns can carry trailing
whitespace (e.g. <think>\n) while a force-open template prefills a bare
<think>. Trim the tag used for the prefix split so the bare prefill is
matched instead of being swallowed into content.

* chat : fix Nemotron Nano v2 regression

---------

Co-authored-by: Alde Rojas <hello@alde.dev>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants