Skip to content

Pin llama.cpp d1747cb: Q2_K weights on HRX0 (llama.cpp #49) - #251

Merged
bong-water-water-bong merged 1 commit into
mainfrom
pin-q2k
Sep 30, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
pin-q2k

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves third_party/llama.cpp from b8d587e to d1747cb. The only change is llama.cpp #49, which gives Q2_K weights a kernel on HRX0.

Why it matters. Q2_K had no matmul on HRX0. At load time llama.cpp tests each weight with a 512-token matmul to decide where it lives; that test failed, so every Q2_K weight went to the CPU_REPACK buffer, and each use split the graph. Every Unsloth UD GGUF carries some Q2_K.

What #49 does

  • Adds Q2_K to the shared dequantizer, which covers prompt WMMA, generic decode and GET_ROWS.
  • Adds Q2_K lane functions to the K-quant decode kernels, for 1 token and for 2-8 tokens.

Measured on strixhalo (HRX0), Qwen3-4B requantized Q8_0 → Q2_K

  • Placement: 709 MiB in CPU_REPACK with 217 graph splits before; all weights on HRX0 with 73 splits after.
  • Speed: pp512 686 → 1287-1325 tok/s. tg128 47.0 → 74.4-74.8 tok/s; the spread of ±14 is likely first-run kernel compiles.
  • Accuracy: KLD against the Q8_0 model is 1.1045 on the prefill path and 1.1031 on decode (-b 1), against 1.1140 on CPU and 1.1224 for the previous HRX with Q2_K on the CPU.
  • Tests: test-backend-ops -b HRX0 passes 929/929; MUL_MAT q2_K 11/11.

Also in this PR:

  • registry/architectures.json regenerated for the pin. tools/registry_build.py --check-pins passes, and so does tools/check_pins.py origin/main (ahead).
  • docs/hrx.md: the sub-4-bit note now says what still falls back to the CPU (IQ2_XS/XXS, IQ1_S/M) and what context7.json: claim the Context7 library #49 changed.

🤖 Generated with Claude Code

b8d587e -> d1747cb is llama.cpp #49 only: Q2_K in the shared dequantizer
(prefill WMMA, generic decode, GET_ROWS) and in the K-quant decode kernels.
Qwen3-4B Q2_K on HRX0: Q2_K weights move from CPU_REPACK (709 MiB, 217 graph
splits) to HRX0 (73 splits); pp512 686 -> 1300 tok/s, tg128 47 -> 75; KLD
against the Q8_0 model equals the CPU's on both the prefill and decode paths.
registry/architectures.json regenerated for the pin (tools/registry_build.py).
docs/hrx.md: the sub-4-bit note.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 30, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit f71a758

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Review (standing in for PR-Agent while coder-llm is paused)

  • Pin: b8d587e..d1747cb is exactly fork context7.json: claim the Context7 library #49 (738460e, the review nit ef779b8, and the merge). I reviewed context7.json: claim the Context7 library #49 line by line against dequantize_row_q2_K; see my comment there.
  • Registry: llama.cpp (hrx) records d1747cb, so registry_pins passes. The counts are unchanged, as expected for a weight-format change.
  • docs/hrx.md: the sub-4-bit note is accurate. IQ2_XS, IQ2_XXS and IQ1_S/IQ1_M are still on CPU, which is why UD files split. Q2_K moved to HRX0 with the measured numbers.
  • No HRX re-check of other models needed: the change only adds a format branch (12) that existing formats never take. context7.json: claim the Context7 library #49's test-backend-ops 929/929 and its KLD numbers cover Q2_K itself.

Good to merge once build is green. pr_agent failing is infra while the model is paused.

@bong-water-water-bong
bong-water-water-bong merged commit acba045 into main Sep 30, 2026
10 of 11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin-q2k branch September 30, 2026 12:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant