Repository navigation
NPU route: reasoning_content, as llama-server returns it - #73
Conversation
…eturns it app/think_split.h splits a thinking model's output (tags split across tokens, a think block the template opened, a literal <think> after the answer stays content); "reasoning_format": "none" keeps the raw text. tests/think_split_test.cpp runs in CI. Qwen3.6-35B-A3B on the NPU: content "Paris", reasoning_content 471 chars, streamed and not (3/3 repeats). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍(Review updated until commit 21d37fa)Here are some key observations to aid the review process:
|
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Persistent review updated to latest commit 21d37fa |
…512 tokens, deterministic flash attention llama.cpp fork since cebcd70 (balanced mode, figures from each PR): - #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX; #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000, tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029. - #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156. - #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959). - #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files. - #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B. Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tic FA, attention sinks) (#295) * Pin llama.cpp 2bd7f58: gpt-oss on HRX, TQ1_0/TQ2_0, MLA prompts past 512 tokens, deterministic flash attention llama.cpp fork since cebcd70 (balanced mode, figures from each PR): - #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX; #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000, tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029. - #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156. - #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959). - #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files. - #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B. Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin llama.cpp 4485916: attention sinks on HRX (gpt-oss) #68 runs gpt-oss's sink logits on HRX as an exact post-correction of the flash-attention output. Docs: docs/hrx.md "Our patches". Registry pin updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin llama.cpp f5b7f4a: PrismML tile bytes (PQ2_0 small-model decode) #74: the low-token SwiGLU read every PQ2_0 / PTQ1_0 row from row 0; Ternary-Bonsai-1.7B PQ2_0 now matches the CPU. Found by the release format matrix. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Checklist item 6: the NPU route returned a thinking model's reasoning inside
content. Lemonade's clients expectreasoning_content, as llama-server returns it.Change:
app/think_split.hsplits the output:<think>after the answer stays in the content;It works for streamed and non-streamed responses.
"reasoning_format": "none"keeps the raw text, as in llama-server.Checks:
tests/think_split_test.cpp: 6 cases, run in CI.Pariswithreasoning_contentof 471 characters, the same streamed and non-streamed, across 3 non-streamed repeats.reasoning_format: nonereturns the raw thinking incontent.🤖 Generated with Claude Code