Repository navigation
Pin llama.cpp 8c71384: IQ4_NL / IQ4_XS on HRX (Qwen3.8-27B UD: 58% -> 72% of Vulkan) - #234
Conversation
…1 graph splits) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
Moves
third_party/llama.cppc9283cf -> 8c71384 (fork #41). Regenerates the registry and updatesdocs/hrx.md.HRX declined IQ4_NL / IQ4_XS because AMD's nibble packers remapped every table index (XOR 12) for a lowering that doesn't apply on gfx1151. So every Unsloth UD GGUF ran its IQ4 matmuls on the CPU. Fork #41:
Qwen3.8-27B UD-Q4_K_XL on strixhalo, HRX0:
RFC #213: the 27B is the one model still under the decode gate (72%). The remaining gap is the Q5_K/IQ4_XS decode kernels' memory bandwidth.
🤖 Generated with Claude Code