Skip to content

Pin llama.cpp f5b7f4a: release pin (gpt-oss, ternary, MLA, deterministic FA, attention sinks) - #295

Merged
bong-water-water-bong merged 4 commits into
mainfrom
pin-llama-f5b7f4a
Oct 2, 2026
Merged

bong-water-water-bong merged 4 commits into
mainfrom
pin-llama-f5b7f4a

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Release pin for Sunday 2026-10-04: llama.cpp fork 2bd7f58 -> f5b7f4a. Three commits, one per step:

Each commit updates docs/hrx.md "Our patches" and the registry pin.

Release format matrix through 1bit serve --device hrx (balanced mode). All 18 rows load with no CPU fallback, and repeated greedy answers are identical. Highlights (decode tok/s):

format model tok/s
Q4_K_M Qwen3-0.6B 317
Q1_0 Bonsai-1.7B 194
Q1_0 Bonsai-8B 56
TQ1_0 Ternary-Bonsai-1.7B 114
TQ2_0 Ternary-Bonsai-1.7B 159
PQ2_0 Ternary-Bonsai-1.7B 206 (CPU-identical answer, at f5b7f4a)
PQ2_0 Ternary-Bonsai-2-27B 19.2
PTQ1_0 Ternary-Bonsai-2-27B 16.0
UD-Q4_K_XL Qwen3.8-27B 11.9
MLA Q4_K_M GLM-4.7-Flash, 761-token prompt 21
MXFP4 gpt-oss-20b 39 (at f5b7f4a)
  • gpt-oss-20b vs CPU: perplexity 463.0 vs 459.0; mean KLD 0.027 (max 1.99); same top token 87.4%. Our claim is "runs on HRX, KLD 0.027 vs CPU", not bit-close.
  • Too few bits for a 0.6B: Qwen3-0.6B IQ1_S / IQ2_XS / Q2_K give the same poor text on CPU. They load and run, and the quality is the model's.

🤖 Generated with Claude Code

bong-water-water-bong and others added 3 commits October 2, 2026 15:51
…512 tokens, deterministic flash attention

llama.cpp fork since cebcd70 (balanced mode, figures from each PR):
- #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX;
  #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000,
  tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029.
- #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156.
- #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959).
- #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files.
- #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits
  (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B.

Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#68 runs gpt-oss's sink logits on HRX as an exact post-correction of the flash-attention output.
Docs: docs/hrx.md "Our patches". Registry pin updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#74: the low-token SwiGLU read every PQ2_0 / PTQ1_0 row from row 0; Ternary-Bonsai-1.7B PQ2_0 now
matches the CPU. Found by the release format matrix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit d91683b

@bong-water-water-bong
bong-water-water-bong merged commit 0baf286 into main Oct 2, 2026
10 of 11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin-llama-f5b7f4a branch October 2, 2026 18:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant