Skip to content

NPU route: reasoning_content, as llama-server returns it - #73

Merged
bong-water-water-bong merged 2 commits into
mainfrom
npu-reasoning
Sep 25, 2026
Merged

bong-water-water-bong merged 2 commits into
mainfrom
npu-reasoning

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Checklist item 6: the NPU route returned a thinking model's reasoning inside content. Lemonade's clients expect reasoning_content, as llama-server returns it.

Change: app/think_split.h splits the output:

  • it handles tags split across tokens;
  • it handles a think block the chat template opened (Qwen3.6);
  • a literal <think> after the answer stays in the content;
  • whitespace around the block is dropped.

It works for streamed and non-streamed responses. "reasoning_format": "none" keeps the raw text, as in llama-server.

Checks:

  • tests/think_split_test.cpp: 6 cases, run in CI.
  • Qwen3.6-35B-A3B on the NPU (strixhalo): content Paris with reasoning_content of 471 characters, the same streamed and non-streamed, across 3 non-streamed repeats. reasoning_format: none returns the raw thinking in content.

🤖 Generated with Claude Code

…eturns it

app/think_split.h splits a thinking model's output (tags split across tokens, a think block
the template opened, a literal <think> after the answer stays content); "reasoning_format":
"none" keeps the raw text. tests/think_split_test.cpp runs in CI. Qwen3.6-35B-A3B on the NPU:
content "Paris", reasoning_content 471 chars, streamed and not (3/3 repeats).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 21d37fa

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit 21d37fa)

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 PR contains tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Potential Buffer Overflow

The ThinkSplit::feed function may potentially lead to a buffer overflow or undefined behavior if the tag string is empty or if buf.size() is less than tag.size(). This could happen if the tag variable is constructed incorrectly or if the buf is manipulated in an unexpected way. The code should validate that tag.size() is not zero and that buf.size() is greater than or equal to tag.size() before performing string operations.

    const std::string tag = in_think ? "</think>" : "<think>";
    const size_t p = buf.find(tag);
    const bool opens = p != std::string::npos && !in_think && !content_started &&
                       buf.find_first_not_of(" \t\r\n") >= p;  // only whitespace before it
    if (p != std::string::npos && (in_think || opens)) {
        out(buf.substr(0, p), emit);
        buf.erase(0, p + tag.size());
        in_think = !in_think;
        continue;
    }
    // hold back the longest suffix that could start the tag
    size_t keep = 0;
    for (size_t k = std::min(tag.size() - 1, buf.size()); k > 0; --k)
        if (buf.compare(buf.size() - k, k, tag, 0, k) == 0) { keep = k; break; }
    out(buf.substr(0, buf.size() - keep), emit);
    buf.erase(0, buf.size() - keep);
    return;
}
Logic Error in Tag Detection

The logic for detecting whether a tag opens a thinking block (opens) may not correctly handle all edge cases. Specifically, the condition buf.find_first_not_of(" \t\r\n") >= p might not correctly identify when the tag is preceded only by whitespace, especially if the buffer contains only whitespace or if the tag is at the beginning of the buffer. This could lead to incorrect splitting of content.

const bool opens = p != std::string::npos && !in_think && !content_started &&
                   buf.find_first_not_of(" \t\r\n") >= p;  // only whitespace before it
if (p != std::string::npos && (in_think || opens)) {

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit 21d37fa

@bong-water-water-bong
bong-water-water-bong merged commit a47e3e1 into main Sep 25, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the npu-reasoning branch September 25, 2026 11:14
bong-water-water-bong pushed a commit that referenced this pull request Oct 2, 2026
…512 tokens, deterministic flash attention

llama.cpp fork since cebcd70 (balanced mode, figures from each PR):
- #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX;
  #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000,
  tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029.
- #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156.
- #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959).
- #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files.
- #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits
  (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B.

Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Oct 2, 2026
…tic FA, attention sinks) (#295)

* Pin llama.cpp 2bd7f58: gpt-oss on HRX, TQ1_0/TQ2_0, MLA prompts past 512 tokens, deterministic flash attention

llama.cpp fork since cebcd70 (balanced mode, figures from each PR):
- #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX;
  #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000,
  tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029.
- #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156.
- #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959).
- #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files.
- #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits
  (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B.

Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp 4485916: attention sinks on HRX (gpt-oss)

#68 runs gpt-oss's sink logits on HRX as an exact post-correction of the flash-attention output.
Docs: docs/hrx.md "Our patches". Registry pin updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp f5b7f4a: PrismML tile bytes (PQ2_0 small-model decode)

#74: the low-token SwiGLU read every PQ2_0 / PTQ1_0 row from row 0; Ternary-Bonsai-1.7B PQ2_0 now
matches the CPU. Found by the release format matrix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant