Skip to content

Pin llama.cpp d60cc4f: PrismML ternary types native on HRX, graph program cache cap - #280

Merged
bong-water-water-bong merged 1 commit into
mainfrom
pin-llama-d60cc4f
Oct 2, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
pin-llama-d60cc4f

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

llama.cpp fork since bd5b297:

  • NPU lane: fresh hw_context per generate(), fixes the idle timeout #62: PrismML's PQ2_0 / PTQ1_0 as native ggml types (142 / 143) with HRX decode on the K-quant
    kernels; Q1_0 decode moves onto them. Ternary-Bonsai-2-27B (balanced mode, llama-bench -fa 1):
    PTQ1_0 5.53 GiB 14.3 tok/s, PQ2_0 6.70 GiB 15.8 tok/s, vs 14.13 GiB / 15.3 for the Q4_0 copy;
    KLD vs CPU 0.000129 for all three. CPU decode bit-identical to PrismML's build on every tensor of
    both 27B files. test-backend-ops -b HRX0 1020/1020.
  • pr-agent: a bot comment no longer cancels a review #63: GGML_HRX_GRAPH_PROGRAM_CACHE (default 64) bounds HRX server memory: ZAYA1-8B over 40
    varying-length requests peaks at 8.7 GiB instead of growing past 24.6 GiB; answers identical.

serve: a file in PrismML's ternary types goes to HRX with --device auto, rotated (prism.hadamard)
or not (Ternary-Bonsai-1.7B); another --device is refused with the converter's name. Smoke test on
strixhalo: 1bit serve -m Ternary-Bonsai-2-27B-PTQ1_0.gguf routed to HRX and answered "Paris" 3/3.
tests/prism_route.sh covers the new routes; ctest 19/19 (build without HRX).

Docs: docs/hrx.md Ternary Bonsai section (native types, table), the cache cap section, "Our patches".
Registry regenerated (no mapping changes).

🤖 Generated with Claude Code

…gram cache cap; serve runs PQ2_0/PTQ1_0 files on HRX

llama.cpp fork since bd5b297:
- #62: PrismML's PQ2_0 / PTQ1_0 as native ggml types (142 / 143) with HRX decode on the K-quant
  kernels; Q1_0 decode moves onto them. Ternary-Bonsai-2-27B (balanced mode, llama-bench -fa 1):
  PTQ1_0 5.53 GiB 14.3 tok/s, PQ2_0 6.70 GiB 15.8 tok/s, vs 14.13 GiB / 15.3 for the Q4_0 copy;
  KLD vs CPU 0.000129 for all three. CPU decode bit-identical to PrismML's build on every tensor of
  both 27B files. test-backend-ops -b HRX0 1020/1020.
- #63: GGML_HRX_GRAPH_PROGRAM_CACHE (default 64) bounds HRX server memory: ZAYA1-8B over 40
  varying-length requests peaks at 8.7 GiB instead of growing past 24.6 GiB; answers identical.

serve: a file in PrismML's ternary types goes to HRX with --device auto, rotated (prism.hadamard)
or not (Ternary-Bonsai-1.7B); another --device is refused with the converter's name. Smoke test on
strixhalo: `1bit serve -m Ternary-Bonsai-2-27B-PTQ1_0.gguf` routed to HRX and answered "Paris" 3/3.
tests/prism_route.sh covers the new routes; ctest 19/19 (build without HRX).

Docs: docs/hrx.md Ternary Bonsai section (native types, table), the cache cap section, "Our patches".
Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Oct 2, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 624cb7b

@bong-water-water-bong
bong-water-water-bong merged commit 6f324e4 into main Oct 2, 2026
10 of 11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin-llama-d60cc4f branch October 2, 2026 08:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant