…tic FA, attention sinks) (#295)
* Pin llama.cpp 2bd7f58: gpt-oss on HRX, TQ1_0/TQ2_0, MLA prompts past 512 tokens, deterministic flash attention
llama.cpp fork since cebcd70 (balanced mode, figures from each PR):
- #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX;
#73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000,
tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029.
- #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156.
- #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959).
- #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files.
- #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits
(Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B.
Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Pin llama.cpp 4485916: attention sinks on HRX (gpt-oss)
#68 runs gpt-oss's sink logits on HRX as an exact post-correction of the flash-attention output.
Docs: docs/hrx.md "Our patches". Registry pin updated.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Pin llama.cpp f5b7f4a: PrismML tile bytes (PQ2_0 small-model decode)
#74: the low-token SwiGLU read every PQ2_0 / PTQ1_0 row from row 0; Ternary-Bonsai-1.7B PQ2_0 now
matches the CPU. Found by the release format matrix.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The NPU route runs FastFlowLM's Q4NX models, and Lemonade's
onebitrecipe now downloadsFastFlowLM/*-NPU2(#72, fork #7). FastFlowLM's README asks projects to acknowledge it withPowered by [FastFlowLM](https://github.com/ROCm/FastFlowLM). This adds:🤖 Generated with Claude Code