Skip to content

HRX: pin our patched branch (IQ3_XXS, honest op claims), rebased on every bump - #29

Merged
bong-water-water-bong merged 1 commit into
mainfrom
hrx-patched-pin
Sep 24, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
hrx-patched-pin

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves third_party/llama.cpp from AMD's commit to 1bit-MONSTER/llama.cpp 1bit/hrx-vulkan-patched (d2a9239f9) = AMD's f1a0aca + 3 commits:

  • IQ3_XXS matmul on HRX: kernel kernels/hrx/mul_mat_vec_iq3xxs_f32.loom and matcher common/dispatch-mul-mat-iq3-xxs.{h,cpp} (our files, full Apache notice), ported from the feat/hrx-port-ae91949 work onto AMD's current layout.
  • Honest op claims: ggml-hrx claimed every op of a declared type, so the scheduler handed it nodes it could not dispatch (unsupported HRX node, no CPU fallback). It now claims only what its dispatcher can execute (except nodes already in HRX memory, like KV views, which cannot move). AMD's IQ4_NL/IQ4_XS matmul and GET_ROWS for IQ4_XS / batched IQ3_S give wrong values, so those nodes are declined.

Verified on Strix Halo (Qwen3-0.6B, perplexity 8 x 512 wikitext tokens):

File Vulkan0 HRX0, AMD's commit HRX0, patched
Q4_K_M 22.53 22.52 22.52, same speed
UD-Q4_K_XL 22.43 fails 22.41
UD-Q2_K_XL 36.26 fails 36.46
UD-IQ2_M 55.44 fails 55.58
  • test-backend-ops -b HRX0: 791 OK, 0 failed (about 1150 fail on AMD's commit).
  • tests/serve_e2e.sh passes on hrx and vulkan; UD-Q2_K_XL serves on hrx.
  • Correct is not fast: the UD files run IQ4_XS / sub-4-bit layers on the CPU, so they belong on Vulkan.

bump-hrx.yml now rebases our commits onto each new AMD pair, tagging the previous tip patched-<sha> first so every pinned commit stays reachable; a conflict fails the run. I dry-ran that logic against a local copy of the fork with a fake new AMD commit: the 3 commits replayed with authorship kept, the branch moved, the old tip stayed reachable by tag.

🤖 Generated with Claude Code

… on every bump

- third_party/llama.cpp -> 1bit-MONSTER/llama.cpp 1bit/hrx-vulkan-patched
  (d2a9239f9): AMD's f1a0aca plus an IQ3_XXS matmul kernel and honest op claims,
  so ggml-hrx only takes nodes it can execute and the rest falls back.
- bump-hrx.yml: our commits are rebased onto each new AMD pair; the previous tip
  is tagged patched-<sha> first, so every pinned commit stays reachable. A
  conflict fails the run for a hand rebase.
- docs/hrx.md: "Our patches", with the Strix Halo results: test-backend-ops
  -b HRX0 791 OK / 0 failed; UD-Q4_K_XL, UD-Q2_K_XL and UD-IQ2_M perplexity on
  HRX0 now matches Vulkan0; Q4_K_M unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit cd865a6 into main Sep 24, 2026
1 check passed
@bong-water-water-bong
bong-water-water-bong deleted the hrx-patched-pin branch September 24, 2026 04:55
bong-water-water-bong pushed a commit that referenced this pull request Sep 27, 2026
…bal barrier (#123, #140)

Brings 1bit-MONSTER/llama.cpp#28 and #29: in every decode-split reduce_fused variant the barrier
between the reduce's global output stores and pack_completed_q8's output loads was
kernel.barrier<workgroup> (LDS only), so the next_q8 copy could be packed from stale output.
It is now kernel.barrier<global> followed by the LDS barrier (5 sites, ops and qwen3_moe
corpora; #29 restores the LDS fence #28 dropped and adds the missing global barrier to
the undispatched standalone reduce_f32 kernels). The pin moves
fa226f9 -> 00adc2b, fast-forward.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Sep 27, 2026
…bal barrier (#123, #140) (#176)

Brings 1bit-MONSTER/llama.cpp#28 and #29: in every decode-split reduce_fused variant the barrier
between the reduce's global output stores and pack_completed_q8's output loads was
kernel.barrier<workgroup> (LDS only), so the next_q8 copy could be packed from stale output.
It is now kernel.barrier<global> followed by the LDS barrier (5 sites, ops and qwen3_moe
corpora; #29 restores the LDS fence #28 dropped and adds the missing global barrier to
the undispatched standalone reduce_f32 kernels). The pin moves
fa226f9 -> 00adc2b, fast-forward.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant