Skip to content

Bump HRX: llama.cpp 02f2880c9f42 + hrx-system 4ba76c18eafe (AMD's tested pair) - #361

Merged
bong-water-water-bong merged 2 commits into
mainfrom
bump-hrx/02f2880c9f42-c0b135a778cc
Oct 8, 2026
Merged

bong-water-water-bong merged 2 commits into
mainfrom
bump-hrx/02f2880c9f42-c0b135a778cc

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves BOTH submodules onto the pair ROCm/ggml-staging-automation now pins, each merged with our own commits (the workflow performs the merge; a conflict fails the run and names the files).

from to
third_party/llama.cpp (1bit-MONSTER/llama.cpp 1bit/hrx-vulkan-patched) e44c9d01a4d5 e44c9d01a4d5 (on AMD 02f2880c9f42)
third_party/hrx-system (1bit-MONSTER/hrx-system 1bit/main) 98d05d94a9f9 4ba76c18eafe (on AMD c0b135a778cc)

Our commits, rebased onto AMD's pin:

CI here builds without HRX. Before merging, on Strix Halo (and test-backend-ops -b HRX0):
cmake -B build -G Ninja -DONEBIT_HRX=ON && cmake --build build --target onebit && for d in hrx cpu; do tests/serve_e2e.sh build/1bit <Qwen3-0.6B Q4_K_M .gguf> $d; done

@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit 42d31a9)

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

314 - Partially compliant

Compliant requirements:

  • The PR updates the hrx-system submodule to the fixed version (4ba76c18eafe) which includes the necessary changes to resolve the compilation issue
  • The PR includes the specific commits that address the flash-attention kernel compilation problem (e.g., shared-memory tile visibility, always-stage fallback)

Non-compliant requirements:

  • The PR does not include explicit verification or testing steps for the fix

Requires further human verification:

  • Need to verify that the actual compilation issue is resolved on the hardware (Strix Halo)
  • Need to confirm that the fix doesn't introduce regressions in other models or configurations

94 - Partially compliant

Compliant requirements:

  • The PR updates the hrx-system submodule which may affect the model registry and checked models

Non-compliant requirements:

  • No changes to registry/architectures.json, registry/checked.json, or census workflow are included in this diff

Requires further human verification:

  • Need to verify that the updated hrx-system doesn't break the model registry or checked models
  • Need to confirm that the census workflow still functions correctly after this update

93 - Partially compliant

Compliant requirements:

  • The PR updates the hrx-system submodule which may include performance improvements for Laya scorer

Non-compliant requirements:

  • No direct changes to Laya scorer implementation or documentation are included in this diff

Requires further human verification:

  • Need to verify that the performance improvements for Laya scorer are actually achieved
  • Need to confirm that the output remains byte-identical
⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Submodule Update

The PR updates the hrx-system submodule to a new commit (4ba76c18eafe) which includes fixes for the decode-split flash-attention kernel compilation issue. While this is the correct fix for the reported issue, it's important to verify that this update doesn't introduce any regressions or compatibility issues with other parts of the system. The update includes several commits that specifically address the flash-attention kernel compilation problem, but without explicit testing or verification steps in this PR, it's uncertain whether all aspects of the fix have been properly validated.

Subproject commit 4ba76c18eafecd1d3c5c3e022c4e097601c9af04

@bong-water-water-bong
bong-water-water-bong enabled auto-merge (squash) October 8, 2026 15:39
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

Persistent review updated to latest commit 42d31a9

@bong-water-water-bong
bong-water-water-bong merged commit c4d9a9b into main Oct 8, 2026
7 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the bump-hrx/02f2880c9f42-c0b135a778cc branch October 8, 2026 15:43
bong-water-water-bong added a commit that referenced this pull request Oct 9, 2026
#361 (c4d9a9b) is titled "Bump HRX: llama.cpp 02f2880c9f42 + hrx-system
4ba76c18eafe (AMD's tested pair)" but its diff changes only
third_party/hrx-system. The llama.cpp gitlink stayed at e44c9d01a4d5, whose
ggml/src/ggml-hrx/loom-jit.cpp still uses loomc_amdgpu_runtime_global_flags_t,
loomc_amdgpu_emit_options_t and LOOMC_AMDGPU_RUNTIME_GLOBAL_*, which
4ba76c18eafe removed. The pinned pair therefore does not compile:

  loom-jit.cpp:70:5: error: unknown type name 'loomc_amdgpu_runtime_global_flags_t'
  loom-jit.cpp:553:1: error: unknown type name 'loomc_amdgpu_runtime_global_flags_t'
  ... 10 errors, all in loom-jit.cpp

GitHub CI builds the pin without HRX (ONEBIT_HRX defaults OFF), so it cannot
catch this class of break; a from-scratch HRX build on strixhalo is what
surfaced it.

Restore third_party/hrx-system to 98d05d94a9f9, the revision #359 pinned and
which exports the API our loom-jit path needs. llama.cpp is unchanged.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant