Repository navigation
Pin llama.cpp 764a256: IQ3_XXS and IQ2_S weights on HRX0 (llama.cpp #51) - #255
Merged
Merged
Conversation
Moves third_party/llama.cpp from d5048ad to 764a256. The only change is llama.cpp #51: IQ3_XXS in the shared dequantizer and the K-quant decode kernels, and IQ2_S in the K-quant decode kernels. - registry/architectures.json regenerated for the pin. - docs/hrx.md: the sub-4-bit note gives the IQ3_XXS / IQ2_S numbers and the mixed-file caveat. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
left a comment
Collaborator
Author
There was a problem hiding this comment.
Review (PR-Agent duty). Looks good.
- The submodule pin 764a256 is the tip of
1bit/hrx-vulkan-patched(checked with ls-remote) and is the merge of llama.cpp #51. - The registry diff changes only the HRX source line, as expected for a pin bump with no architecture changes.
- The docs/hrx.md numbers match #51 (IQ3_XXS 33.4 pure / 8.9 mixed, IQ2_S 11.3, balanced power mode), and the mixed-grid limit is stated.
- The repeated op runs before the merge were identical across three runs: MUL_MAT 11 OK, GET_ROWS 1 OK + 3 declined, MUL_MAT_ID 3 declined.
Merge when checks pass.
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Moves
third_party/llama.cppfromd5048adto764a256. The only change is llama.cpp #51, which gives IQ3_XXS and IQ2_S weights kernels on HRX0.Why it matters. Before #51:
Unsloth's UD GGUFs below 4 bits carry both formats.
Measured on strixhalo (HRX0, balanced power mode), Qwen3-4B requantized from Q8_0
test-backend-ops -b HRX0: 972/972, plus three repeated GET_ROWS / MUL_MAT_ID / MUL_MAT runs on iq3_xxs and iq2_s with identical results each time.Also in this PR:
registry/architectures.jsonregenerated for the pin.tools/registry_build.py --check-pinspasses, and so doestools/check_pins.py origin/main(ahead).docs/hrx.md: the sub-4-bit note adds the IQ3_XXS / IQ2_S numbers and the mixed-file caveat.🤖 Generated with Claude Code