Skip to content

Pin llama.cpp cebcd70: the Hadamard rotation on HRX (Bonsai pp512 13.5 -> 90.9), MXFP4 weights - #282

Merged
bong-water-water-bong merged 1 commit into
mainfrom
pin-llama-cebcd70
Oct 2, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
pin-llama-cebcd70

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

llama.cpp fork since d60cc4f:

  • site: Docs7 analytics #64: a kernel for the MUL_MAT ggml hints as GGML_HINT_SRC0_IS_HADAMARD. Bonsai's 1024-point rotation
    of the 17408-wide FFN input ran on the CPU in every layer at prompt sizes. Ternary-Bonsai-2-27B,
    balanced mode: pp512 13.5 -> 90.9 tok/s, tg128 15.8 -> 19.0. test-hrx-hadamard: six shapes vs CPU
    and an exact product, x4, bitwise-identical repeats; 8 sanitizer-clean check.cases.
  • ComfyUI.cpp inside the engine: 1bit comfy #65: MXFP4 in the shared dequantizer, the E8M0 scale built exactly as GGML_E8M0_TO_FP32_HALF.
    test-hrx-mxfp4 bit-exact vs ggml on HRX0 (subnormal scales e = 0/1 flush to zero);
    test-backend-ops -b HRX0 1051/1051. MUL_MAT_ID admits only multiples of 256 (its kernels declare
    it): gpt-oss-20b's 2880-wide experts fall back to the CPU instead of failing the JIT.

Docs: docs/hrx.md Ternary Bonsai section and "Our patches". Registry regenerated (no mapping changes).

🤖 Generated with Claude Code

…5 -> 90.9), MXFP4 weights

llama.cpp fork since d60cc4f:
- #64: a kernel for the MUL_MAT ggml hints as GGML_HINT_SRC0_IS_HADAMARD. Bonsai's 1024-point rotation
  of the 17408-wide FFN input ran on the CPU in every layer at prompt sizes. Ternary-Bonsai-2-27B,
  balanced mode: pp512 13.5 -> 90.9 tok/s, tg128 15.8 -> 19.0. test-hrx-hadamard: six shapes vs CPU
  and an exact product, x4, bitwise-identical repeats; 8 sanitizer-clean check.cases.
- #65: MXFP4 in the shared dequantizer, the E8M0 scale built exactly as GGML_E8M0_TO_FP32_HALF.
  test-hrx-mxfp4 bit-exact vs ggml on HRX0 (subnormal scales e = 0/1 flush to zero);
  test-backend-ops -b HRX0 1051/1051. MUL_MAT_ID admits only multiples of 256 (its kernels declare
  it): gpt-oss-20b's 2880-wide experts fall back to the CPU instead of failing the JIT.

Docs: docs/hrx.md Ternary Bonsai section and "Our patches". Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Oct 2, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit c98a448

@bong-water-water-bong
bong-water-water-bong merged commit 4b5d828 into main Oct 2, 2026
10 of 11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin-llama-cebcd70 branch October 2, 2026 11:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant