Repository navigation
Pin llama.cpp dd74f6b: IQ1_S/IQ1_M on HRX0 (#53), packed ternary decode for Bonsai (#54) - #268
Conversation
…de for Bonsai (#54) - third_party/llama.cpp: cde002d -> dd74f6b, adding llama.cpp #53 and #54. - #53: IQ1_S and IQ1_M weights run on HRX0 (shared dequantizer, K-quant decode) instead of the CPU. - #54: exact-ternary Q4_0 weights decode from a 2-bit copy made at load, opt-in with GGML_HRX_TERNARY_Q4_0. - 1bit serve sets GGML_HRX_TERNARY_Q4_0=1 for files stamped onebit.ternary_q4_0 (tools/ternary_to_q4_0.py), unless the user set it. It prints that the packed copy costs about 30% of the file in extra GPU memory. - docs/hrx.md: the IQ1 numbers, and the packed ternary path with its memory cost (+4 GiB for Bonsai-2-27B) and decode speed (9.3 -> 15.4 tok/s). - registry/architectures.json regenerated for the pin. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
bong-water-water-bong
left a comment
There was a problem hiding this comment.
Review (PR-Agent duty). Looks good.
- Pin: dd74f6b is the tip of
1bit/hrx-vulkan-patched(checked with ls-remote) and covers #53 and #54. The registry diff is the pin line only. - serve: sets
GGML_HRX_TERNARY_Q4_0=1only when the device ishrxand the file carries theonebit.ternary_q4_0 = 128stamp written bytools/ternary_to_q4_0.py, and a value the user set wins. The notice's "=0 turns it off" is true:ternary_q4_0_enabled()treats an empty value or a leading '0' as off. - Memory: stating the +4 GiB (about 30% of the file) in the startup notice and the docs is enough for now. A hard pre-load check can wait until serve has a general GGUF size walk; no need to add one just for this.
Merging when checks pass.
PR Reviewer Guide 🔍(Review updated until commit 52fc5ec)Here are some key observations to aid the review process:
|
|
Persistent review updated to latest commit 52fc5ec |
Moves
third_party/llama.cppfromcde002dtodd74f6b, which adds two fork PRs:--mtp) and ZAYA1-8B showed no change for the other formats.GGML_HRX_TERNARY_Q4_0.What changes in the engine:
1bit servesetsGGML_HRX_TERNARY_Q4_0=1on HRX for files stampedonebit.ternary_q4_0(written bytools/ternary_to_q4_0.py), unless the user has set it. It also prints the memory cost. Prompt batches keep the Q4_0 copy, so the packed copy is extra GPU memory.docs/hrx.md: the IQ1 numbers, and the packed ternary path with its cost.registry/architectures.jsonregenerated for the pin.tools/registry_build.py --check-pinspasses, andtools/check_pins.py origin/mainreports the pin as ahead.Bonsai-2-27B (exact Q4_0), HRX0, balanced power mode:
serve.cpppasses a syntax-only compile here; the full build runs in CI.🤖 Generated with Claude Code