Skip to content

Pin llama.cpp dd74f6b: IQ1_S/IQ1_M on HRX0 (#53), packed ternary decode for Bonsai (#54) - #268

Merged
bong-water-water-bong merged 2 commits into
mainfrom
hrx-iq1-ternary
Oct 1, 2026
Merged

bong-water-water-bong merged 2 commits into
mainfrom
hrx-iq1-ternary

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves third_party/llama.cpp from cde002d to dd74f6b, which adds two fork PRs:

  • llama.cpp #53: IQ1_S and IQ1_M weights run on HRX0 (shared dequantizer plus K-quant decode kernels) instead of falling back to the CPU. Qwen3-4B, balanced power mode, KLD equal to the CPU:
    • IQ1_S: pp512 70.8 → 81.7 tok/s, tg128 14.1 → 18.6.
    • IQ1_M: pp512 70.8 → 80.4, tg128 14.1 → 18.1.
    • An interleaved A/B on 27B UD-Q4_K_XL (with and without --mtp) and ZAYA1-8B showed no change for the other formats.
  • llama.cpp #54: packed ternary decode, opt-in with GGML_HRX_TERNARY_Q4_0.
    • The K-quant decode kernels read exact-ternary Q4_0 weights from a 2-bit copy made at load.
    • The load checks every block and refuses a model that isn't ternary.

What changes in the engine:

  • 1bit serve sets GGML_HRX_TERNARY_Q4_0=1 on HRX for files stamped onebit.ternary_q4_0 (written by tools/ternary_to_q4_0.py), unless the user has set it. It also prints the memory cost. Prompt batches keep the Q4_0 copy, so the packed copy is extra GPU memory.
  • docs/hrx.md: the IQ1 numbers, and the packed ternary path with its cost.
  • registry/architectures.json regenerated for the pin. tools/registry_build.py --check-pins passes, and tools/check_pins.py origin/main reports the pin as ahead.

Bonsai-2-27B (exact Q4_0), HRX0, balanced power mode:

Q4_0 kernels packed ternary
tg128 9.26 tok/s 15.35 tok/s
pp512 13.1 tok/s 13.1 tok/s
GPU peak over idle 17.3 GiB 21.3 GiB (+4.0)
KLD vs CPU logits 0.0019 0.00013

serve.cpp passes a syntax-only compile here; the full build runs in CI.

🤖 Generated with Claude Code

…de for Bonsai (#54)

- third_party/llama.cpp: cde002d -> dd74f6b, adding llama.cpp #53 and #54.
  - #53: IQ1_S and IQ1_M weights run on HRX0 (shared dequantizer, K-quant decode) instead of the CPU.
  - #54: exact-ternary Q4_0 weights decode from a 2-bit copy made at load, opt-in with
    GGML_HRX_TERNARY_Q4_0.
- 1bit serve sets GGML_HRX_TERNARY_Q4_0=1 for files stamped onebit.ternary_q4_0
  (tools/ternary_to_q4_0.py), unless the user set it. It prints that the packed copy costs about 30% of
  the file in extra GPU memory.
- docs/hrx.md: the IQ1 numbers, and the packed ternary path with its memory cost (+4 GiB for
  Bonsai-2-27B) and decode speed (9.3 -> 15.4 tok/s).
- registry/architectures.json regenerated for the pin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 52fc5ec

@bong-water-water-bong bong-water-water-bong left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review (PR-Agent duty). Looks good.

  • Pin: dd74f6b is the tip of 1bit/hrx-vulkan-patched (checked with ls-remote) and covers #53 and #54. The registry diff is the pin line only.
  • serve: sets GGML_HRX_TERNARY_Q4_0=1 only when the device is hrx and the file carries the onebit.ternary_q4_0 = 128 stamp written by tools/ternary_to_q4_0.py, and a value the user set wins. The notice's "=0 turns it off" is true: ternary_q4_0_enabled() treats an empty value or a leading '0' as off.
  • Memory: stating the +4 GiB (about 30% of the file) in the startup notice and the docs is enough for now. A hard pre-load check can wait until serve has a general GGUF size walk; no need to add one just for this.

Merging when checks pass.

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit 52fc5ec)

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

53 - Partially compliant

Compliant requirements:

  • Discord invite link updated to permanent invite
  • Community link added to README

Non-compliant requirements:

  • None

Requires further human verification:

  • None

54 - Partially compliant

Compliant requirements:

  • Canonical URLs implemented
  • Link previews implemented
  • Structured data (JSON-LD) implemented
  • robots.txt and sitemap.xml implemented
  • IndexNow support implemented

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Incorrect ternary Q4_0 detection

The ternary_q4_0 function in app/serve.cpp checks for a GGUF metadata key onebit.ternary_q4_0 with a value of 128. However, the PR description states that the tools/ternary_to_q4_0.py stamps files with onebit.ternary_q4_0 set to 128, and the documentation in docs/hrx.md also mentions this. If the value is indeed 128, then the check is correct, but if it's a different value, this could lead to incorrect behavior. The function should be verified to ensure it correctly identifies ternary Q4_0 models.

bool ternary_q4_0(const std::string& model) {
    return model.size() > 5 && model.compare(model.size() - 5, 5, ".gguf") == 0 &&
           gguf_int(model, "onebit.ternary_q4_0") == 128;
}
Memory cost estimation

The comment in app/serve.cpp states that the packed copy is about 30% of the file size in extra GPU memory. This is a specific claim that should be verified with actual measurements on the hardware. If the actual memory usage differs significantly from this estimate, it could lead to incorrect resource allocation or performance issues.

// The Q4_0 copy stays resident for prompt batches, so the packed copy is extra GPU memory: about 30% of
// the file (+4.0 GiB for Bonsai-2-27B, measured; only the decode kernels' projections are packed).
if (device == "hrx" && ternary_q4_0(o.model) && !std::getenv("GGML_HRX_TERNARY_Q4_0")) {
    env.push_back("GGML_HRX_TERNARY_Q4_0=1");
    std::fprintf(stderr, "1bit serve: ternary Q4_0 file: HRX0 decodes from a 2-bit copy of its weights, "
                         "about 30%% of the file size in extra GPU memory (GGML_HRX_TERNARY_Q4_0=0 turns it off)\n");
}

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

Persistent review updated to latest commit 52fc5ec

@bong-water-water-bong
bong-water-water-bong merged commit 5c81776 into main Oct 1, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the hrx-iq1-ternary branch October 1, 2026 18:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant