Repository navigation
Site SEO: canonical URLs, preview card, JSON-LD, robots.txt, sitemap.xml, IndexNow - #54
Merged
Merged
Conversation
…emap.xml, IndexNow Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
bong-water-water-bong
added a commit
that referenced
this pull request
Oct 1, 2026
…de for Bonsai (#54) (#268) - third_party/llama.cpp: cde002d -> dd74f6b, adding llama.cpp #53 and #54. - #53: IQ1_S and IQ1_M weights run on HRX0 (shared dequantizer, K-quant decode) instead of the CPU. - #54: exact-ternary Q4_0 weights decode from a 2-bit copy made at load, opt-in with GGML_HRX_TERNARY_Q4_0. - 1bit serve sets GGML_HRX_TERNARY_Q4_0=1 for files stamped onebit.ternary_q4_0 (tools/ternary_to_q4_0.py), unless the user set it. It prints that the packed copy costs about 30% of the file in extra GPU memory. - docs/hrx.md: the IQ1 numbers, and the packed ternary path with its memory cost (+4 GiB for Bonsai-2-27B) and decode speed (9.3 -> 15.4 tok/s). - registry/architectures.json regenerated for the pin. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
pushed a commit
that referenced
this pull request
Oct 1, 2026
…3.8-27B pp512 98 -> 335 llama.cpp fork #55 routes Q4_K/Q5_K/IQ4_XS prompt matmuls to the existing q8_1 x4 prefill kernel: - Q5_K and IQ4_XS get Q4_K's activation/GLU policy. - The generic fused SwiGLU (priority 290) leaves 256-2048-token chunks to it. On Qwen3.8-27B UD-Q4_K_XL those projections had gone to the generic F32 WMMA kernels, 94% of HRX prompt time. Measured on this pin (6e42b51, which also carries #53/#54): - llama-bench -b 512 -ub 512: pp512 98.5 -> 334.6, pp2048 98.1 -> 309.5 tok/s. - 1bit serve --device hrx, 14,435-token prompt: 90.7 -> 264.7 tok/s, same text. - KLD vs the BF16 logits (wikitext-2, 20 x 512): 0.00712, same top token 96.27%. - test-backend-ops MUL_MAT on HRX0: 287/287. GGML_HRX_Q8_PREFILL_RELAX=0 restores the old routing. Docs: a docs/hrx.md section and "Our patches" entry. The docs/serve.md auto-routing note ("97-99 tok/s until fork #55 is pinned") and the docs/laya.md policy note now carry the measured numbers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
added a commit
that referenced
this pull request
Oct 1, 2026
…3.8-27B pp512 98 -> 335 (#273) * Pin llama.cpp 6e42b51: HRX prompt matmuls on the q8_1 x4 kernel, Qwen3.8-27B pp512 98 -> 335 llama.cpp fork #55 routes Q4_K/Q5_K/IQ4_XS prompt matmuls to the existing q8_1 x4 prefill kernel: - Q5_K and IQ4_XS get Q4_K's activation/GLU policy. - The generic fused SwiGLU (priority 290) leaves 256-2048-token chunks to it. On Qwen3.8-27B UD-Q4_K_XL those projections had gone to the generic F32 WMMA kernels, 94% of HRX prompt time. Measured on this pin (6e42b51, which also carries #53/#54): - llama-bench -b 512 -ub 512: pp512 98.5 -> 334.6, pp2048 98.1 -> 309.5 tok/s. - 1bit serve --device hrx, 14,435-token prompt: 90.7 -> 264.7 tok/s, same text. - KLD vs the BF16 logits (wikitext-2, 20 x 512): 0.00712, same top token 96.27%. - test-backend-ops MUL_MAT on HRX0: 287/287. GGML_HRX_Q8_PREFILL_RELAX=0 restores the old routing. Docs: a docs/hrx.md section and "Our patches" entry. The docs/serve.md auto-routing note ("97-99 tok/s until fork #55 is pinned") and the docs/laya.md policy note now carry the measured numbers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * registry: record llama.cpp (hrx) pin 6e42b51 (tools/registry_build.py; no mapping changes) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What 1bit.gg was missing for search engines and link previews (robots.txt and sitemap.xml both returned 404):
https://1bit.gg/...), so www and old links fold into one address.og:url,og:type(article for docs and posts),og:imagewith a new 1200x630 card (site/assets/og-card.png), andtwitter:card=summary_large_imagewith title, description and image. Discord, X, Slack and Reddit show the card.WebSite+SoftwareSourceCode(repository, Apache-2.0, C++); each post is aBlogPostingwith its dates.noindex). Each page'slastmodis its source file's last commit, sopages.ymlnow checks out the full history.continue-on-error, so a failed ping never fails a deploy). The key is public by design.Built locally: every page carries its canonical URL and card; the home page and posts carry valid JSON-LD; the sitemap lists 23 URLs.
Google needs one step outside the repo: verify
1bit.ggin Search Console (a DNS TXT record at Namecheap) and submithttps://1bit.gg/sitemap.xml.🤖 Generated with Claude Code