Skip to content

Pin llama.cpp ade79d4: K-quant decode kernels on HRX (Qwen3.8-27B 72% -> 93% of Vulkan) - #235

Merged
bong-water-water-bong merged 1 commit into
mainfrom
pin-hrx-kquant
Sep 29, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
pin-hrx-kquant

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Pins third_party/llama.cpp to ade79d4 (llama.cpp #42). With it, HRX decode reads Q4_K, Q5_K, Q6_K, IQ4_NL, IQ4_XS and Q8_0 weights straight from their GGUF blocks:

  • The FFN gate/up pair is fused with SwiGLU.
  • Projections fold in the residual add.

Decode, tg128, HRX vs the Vulkan build on the same gfx1151 box (RFC #213 gate: within 10%):

Model HRX Vulkan Ratio
Qwen3.8-27B UD-Q4_K_XL 11.60 (was 9.0) 12.49 92.9%
ZAYA1-8B Q4_K_M 92.7 (was ~90) 93 ~100%
Qwen3-0.6B Q4_K_M 326.6 (was 320.8) 357 91.4%
Qwen3-Coder-30B-A3B Q4_K_M 90.6 93.7 96.6%

Accuracy and tests:

  • 27B decode-shaped wikitext PPL (-ub 1): 7.1550 vs 7.1547 on the old HRX path; CPU gives 7.1475.
  • test-backend-ops -o MUL_MAT -b HRX0 passes.
  • All six Loom check cases pass under the access sanitizer.

Thermals: the 27B peaks at 90–93 °C on HRX, against 67 °C on Vulkan.

Other changes:

  • registry/architectures.json is regenerated for the new pin; --check-pins passes.
  • docs/hrx.md gets a section on the new kernels.

🤖 Generated with Claude Code

… 72% -> 93% of Vulkan)

llama.cpp #42: one-token projections on Q4_K, Q5_K, Q6_K, IQ4_NL, IQ4_XS
and Q8_0 read the GGUF blocks directly (SwiGLU gate/up pair fused, residual
add folded into projections). Decode vs Vulkan on gfx1151: Qwen3.8-27B
UD-Q4_K_XL 11.60 / 12.49, ZAYA1-8B 92.7 / 93, Qwen3-0.6B 326.6 / 357,
Qwen3-Coder-30B-A3B 90.6 / 93.7. docs/hrx.md updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 29, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 4f784c9

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Performance Claim Verification

The PR includes performance claims such as "92.9% (was 9.0) tok/s" and "91.4% (was 320.8) tok/s" for Qwen3.8-27B and Qwen3-0.6B models. These numbers must be verified to ensure they are measured on the actual hardware (gfx1151 Strix Halo) and not just simulated or estimated values. Without concrete benchmarking data, these claims may be overstated or inaccurate.

| Qwen3.8-27B UD-Q4_K_XL | 11.60 | 12.49 | 92.9% (was 9.0) |
| ZAYA1-8B Q4_K_M | 92.7 | 93 | ~100% (was ~90) |
| Qwen3-0.6B Q4_K_M | 326.6 | 357 | 91.4% (was 320.8) |
| Qwen3-Coder-30B-A3B Q4_K_M | 90.6 | 93.7 | 96.6% (unchanged) |
Thermal Claim Verification

The PR states that "the 27B on HRX runs hot on Strix Halo (90-93 C peak vs 67 C for Vulkan)." This thermal performance claim needs to be substantiated with actual measured temperatures from the hardware to confirm the validity of the assertion.

~89 ms token). The 27B on HRX runs hot on Strix Halo (90-93 C peak vs 67 C for Vulkan).

@bong-water-water-bong
bong-water-water-bong merged commit 33cd8d3 into main Sep 29, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin-hrx-kquant branch September 29, 2026 20:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant