Repository navigation
Bump HRX: llama.cpp e57beb9721af (HIP plumbing hook, HRX fixes, Flash-Next) - #311
Merged
Merged
Conversation
Moves third_party/llama.cpp (1bit-MONSTER/llama.cpp 1bit/hrx-vulkan-patched) 2e5fcf56a24d -> e57beb9721af; hrx-system stays at 51b1739ae5fd. Registry regenerated at the new pin: 290 -> 292 architectures (qwen4exp). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
registry/architectures.json was generated against the previous gitlink (registry_pins failed). tests/device_route.sh used qwen4exp as its "not built" example; at this pin Flash-Next runs on HRX, so the test now checks it takes the plain route and keeps zamba2 as the refused example. Docs that listed qwen4exp as unbuilt say it is back. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Collaborator
Author
|
CI fix (fc9502c): Verified on Strix Halo at 1b13c6b (performance mode, 120 W cap, max Tctl 66 °C), fresh clone
Release is held; not merging without the owner's word. |
This was referenced Oct 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Moves our HRX llama.cpp fork pin to the current branch tip; hrx-system is unchanged.
1bit/hrx-vulkan-patched)2e5fcf56a24de57beb9721af51b1739ae5fd51b1739ae5fdWhat the new pin brings (fork PRs #81, #82, #83, #86, #87, #88):
ggml-hrx/hip/: HIP-through-HRX plumbing and theGGML_HRX_HIP_ADDON_DIRhook that-DONEBIT_GPU_PRIVATE(Build option for private HIP kernels on HRX (-DONEBIT_GPU_PRIVATE) #307) requires;cmake/hrx.cmakecurrently fails on this pin check, so Build option for private HIP kernels on HRX (-DONEBIT_GPU_PRIVATE) #307 is only usable after this bump.qwen4exp) on HRX, with the NextN/MTP draft head and lazy tensor reads from upstream.Registry: regenerated at the new pin (
tools/registry_build.py,--check-pinspasses): 290 -> 292 mapped architectures (Qwen4ExpForCausalLM,Qwen4ExpForConditionalGeneration->qwen4exp).Issues: #284, #286, #290 — fixes verified in the fork PRs; the closing comments follow the engine-side e2e below.
CI here builds without HRX. Before merging, on Strix Halo:
cmake -B build -G Ninja -DONEBIT_HRX=ON && cmake --build build --target onebit && tests/hrx_lemonade_e2e.sh build/1bit build/hrx/llama/bin/llama-server— queued behind the current GPU sweep on strixhalo; evidence will be posted here. Release is held: merge only on the owner's word.
🤖 Generated with Claude Code