Skip to content

Pin llama.cpp 764a256: IQ3_XXS and IQ2_S weights on HRX0 (llama.cpp #51) - #255

Merged
bong-water-water-bong merged 1 commit into
mainfrom
pin-llama-764a256
Oct 1, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
pin-llama-764a256

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves third_party/llama.cpp from d5048ad to 764a256. The only change is llama.cpp #51, which gives IQ3_XXS and IQ2_S weights kernels on HRX0.

Why it matters. Before #51:

  • IQ3_XXS had only a one-row-per-workitem matvec and no dequantizer path. llama.cpp's load-time 512-token matmul test failed, so IQ3_XXS weights went to CPU_REPACK.
  • IQ2_S decode took the generic dequantize path.

Unsloth's UD GGUFs below 4 bits carry both formats.

Measured on strixhalo (HRX0, balanced power mode), Qwen3-4B requantized from Q8_0

Format KLD vs the Q8_0 model, HRX decode path / CPU Prompt (pp512) Decode (tg)
IQ3_XXS, mixed quant 0.309 / 0.309 2.8 → 250 tok/s 1.5 → 8.9 tok/s
IQ3_XXS, every tensor IQ3_XXS — 510-560 tok/s 33.4 tok/s
IQ2_S 0.827 / 0.837 unchanged 1.5 → 11.3 tok/s
  • test-backend-ops -b HRX0: 972/972, plus three repeated GET_ROWS / MUL_MAT_ID / MUL_MAT runs on iq3_xxs and iq2_s with identical results each time.
  • Mixed files decode slower than pure ones. A SwiGLU gate/up pair whose two formats need different grids is declined, because both share one grid buffer. That's a follow-up.

Also in this PR:

  • registry/architectures.json regenerated for the pin. tools/registry_build.py --check-pins passes, and so does tools/check_pins.py origin/main (ahead).
  • docs/hrx.md: the sub-4-bit note adds the IQ3_XXS / IQ2_S numbers and the mixed-file caveat.

🤖 Generated with Claude Code

Moves third_party/llama.cpp from d5048ad to 764a256. The only change is
llama.cpp #51: IQ3_XXS in the shared dequantizer and the K-quant decode
kernels, and IQ2_S in the K-quant decode kernels.

- registry/architectures.json regenerated for the pin.
- docs/hrx.md: the sub-4-bit note gives the IQ3_XXS / IQ2_S numbers and the
  mixed-file caveat.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@bong-water-water-bong bong-water-water-bong left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review (PR-Agent duty). Looks good.

  • The submodule pin 764a256 is the tip of 1bit/hrx-vulkan-patched (checked with ls-remote) and is the merge of llama.cpp #51.
  • The registry diff changes only the HRX source line, as expected for a pin bump with no architecture changes.
  • The docs/hrx.md numbers match #51 (IQ3_XXS 33.4 pure / 8.9 mixed, IQ2_S 11.3, balanced power mode), and the mixed-grid limit is stated.
  • The repeated op runs before the merge were identical across three runs: MUL_MAT 11 OK, GET_ROWS 1 OK + 3 declined, MUL_MAT_ID 3 declined.

Merge when checks pass.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

51 - Partially compliant

Compliant requirements:

  • The documentation has been updated to remove the incorrect note about MiniMax-H3
  • The wiki is noted to be corrected in the ticket description

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Documentation Update

The documentation was updated to remove the incorrect note about MiniMax-H3 not running correctly on the llama.cpp build. This change aligns with the ticket's objective to correct the documentation and wiki.

- **Fast sub-4-bit kernels on HRX.** IQ1_S/IQ1_M matmuls still fall back to the CPU, and one
  such tensor splits every layer it is in, which is why Unsloth's smallest UD files are slow on
  HRX0. IQ3_S has a decode kernel (#43) but not a fast prefill one. Q2_K
  ([llama.cpp #49](https://github.com/1bit-MONSTER/llama.cpp/pull/49)), IQ2_XXS/IQ2_XS
  ([llama.cpp #50](https://github.com/1bit-MONSTER/llama.cpp/pull/50)) and IQ3_XXS/IQ2_S
  ([llama.cpp #51](https://github.com/1bit-MONSTER/llama.cpp/pull/51)) run on HRX0: in the shared
  dequantizer (prefill WMMA, generic decode, GET_ROWS; IQ2_S was already there) and the K-quant
  decode kernels. Before them, llama.cpp's load-time buffer check (a 512-token matmul) failed, and
  every such weight went to the CPU. On Qwen3-4B, with KLD against the Q8_0 model equal to the CPU's:
  - Q2_K: 709 MiB CPU_REPACK and 217 graph splits -> all on HRX0 and 73 splits; pp512
    686 -> 1300 tok/s, tg128 47 -> 75.
  - IQ2_XXS: pp512 59 -> 715 tok/s, tg128 23.0 -> 39.5.
  - IQ2_XS: pp512 87 -> 268 tok/s, tg128 28.0 -> 30.2.
  - IQ3_XXS: pp512 2.8 -> 250 tok/s (mixed quant; 510-560 with every tensor IQ3_XXS), tg 1.5 -> 8.9
    (mixed) and 33.4 (pure), balanced power mode.
  - IQ2_S: tg 1.5 -> 11.3 tok/s (pure); its prefill is unchanged.

  Some batched IQ2_XXS matmul shapes are still declined and run on the CPU. Mixed files decode
  slower than pure ones: a SwiGLU gate/up pair whose two formats need different grids shares one
  grid buffer, so it is declined and takes the generic path.
- **Q4NX on HRX.** Our Q4NX kernels lived in ggml-hrx2 and are not in ggml-hrx.

@context7

context7 Bot commented Oct 1, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ⚠️ Did not finish. View run

Commit 4cd3e07

@bong-water-water-bong
bong-water-water-bong merged commit 741fba9 into main Oct 1, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin-llama-764a256 branch October 1, 2026 07:43
bong-water-water-bong added a commit that referenced this pull request Oct 1, 2026
…255, #257) (#266)

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant