Skip to content

Pin llama.cpp d5048ad: IQ2_XXS and IQ2_XS weights on HRX0 (llama.cpp #50) - #254

Merged
bong-water-water-bong merged 1 commit into
mainfrom
pin-llama-d5048ad
Sep 30, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
pin-llama-d5048ad

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves third_party/llama.cpp from d1747cb to d5048ad. The only change is llama.cpp #50, which gives IQ2_XXS and IQ2_XS weights kernels on HRX0.

Why it matters. IQ2_XXS and IQ2_XS had no matmul on HRX0. llama.cpp's load-time 512-token matmul test failed, so these weights went to CPU_REPACK and each use split the graph. Unsloth's UD GGUFs below 4 bits carry both.

What #50 does

  • Adds IQ2_XXS (format 24) and IQ2_XS (format 25) to the shared dequantizer, which covers prompt WMMA, generic decode and GET_ROWS.
  • Adds lane functions to the K-quant decode kernels, for 1 token and for 2-8 tokens. The grids are staged in workgroup memory, like IQ3_S.
  • The generators for the grid tables and decoders are in ggml/src/ggml-hrx/tools/generate_iq2_*.py, and their output is verbatim in the kernel files.

Measured on strixhalo (HRX0), Qwen3-4B requantized Q8_0 → IQ2_*

  • Speed, previous (IQ2 on the CPU) → this:
    • IQ2_XXS: pp512 59 → 715 tok/s, tg128 23.0 → 39.5 tok/s.
    • IQ2_XS: pp512 87 → 268 tok/s, tg128 28.0 → 30.2 tok/s.
    • tg128 is from 2 × 6 runs with the box lock held.
  • Accuracy, KLD against the Q8_0 model on the decode path (-b 1):
    • IQ2_XXS: 1.1306, against 1.141 on CPU.
    • IQ2_XS: 0.7744, against 0.786 on CPU.
  • Tests: test-backend-ops -b HRX0 passes 971/971. 64 batched IQ2_XXS matmul shapes are still declined and run on the CPU, as before.

Also in this PR:

  • registry/architectures.json regenerated for the pin. tools/registry_build.py --check-pins passes, and so does tools/check_pins.py origin/main (ahead).
  • docs/hrx.md: the sub-4-bit note now lists IQ1_S/IQ1_M as the remaining CPU fallback and gives the IQ2 numbers.

🤖 Generated with Claude Code

)

Moves third_party/llama.cpp from d1747cb to d5048ad. The only change is
llama.cpp #50: IQ2_XXS and IQ2_XS in the shared dequantizer and the K-quant
decode kernels, so these weights no longer go to the CPU at load time.

- registry/architectures.json regenerated for the pin.
- docs/hrx.md: the sub-4-bit note now lists IQ1_S/IQ1_M as the remaining CPU
  fallback and gives the IQ2 numbers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 30, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 33215a7

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Review (standing in for PR-Agent while coder-llm is paused)

Good to merge once build passes.

@bong-water-water-bong
bong-water-water-bong merged commit 7cf9363 into main Sep 30, 2026
10 of 11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin-llama-d5048ad branch September 30, 2026 17:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant