Skip to content

Step 2: HRX + Vulkan in one llama.cpp build, wired into Lemonade - #7

Merged
bong-water-water-bong merged 3 commits into
embed-lemonadefrom
hrx-vulkan
Sep 22, 2026
Merged

bong-water-water-bong merged 3 commits into
embed-lemonadefrom
hrx-vulkan

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Step 2 of the port (docs/PORTING.md). Stacked on #6, which is stacked on #5.

What

  • Pinned sources:
    • ROCm/hrx 0bc22fb: the runtime.
    • ROCm/hrx-system 6743075f: the last loom-link with --mode=selective, which the fork's kernel catalog needs.
    • Your llama.cpp fork on branch 1bit-engine/hrx-vulkan (3b33c8a9).
  • Patches: patches/hrx/0001–0003 replace edits that until now existed only in a working tree. 0001 and 0002 are build plumbing. 0003 is a real runtime fix: the gfx1151 PM4 probe.
  • Build: -DONEBIT_HRX=ON (off by default) builds all three from the build tree, leaving the submodules clean.
    • Only the targets ggml-hrx2 uses get built.
    • hrx_assemble_prefix.cmake reproduces the prefix that used to be hand-assembled and undocumented.
  • Lemonade wiring: 1bit lemonade runs llamacpp-hrx on HRX20 and llamacpp/Vulkan on Vulkan0, both from this one llama-server. User-set config keys win.
  • Lemonade delta: a new hrx_device option (default HRX0) replaces the hard-coded device name. It's documented in UPSTREAM.md.

Verified on Strix Halo (fresh clone, submodules at pins, 217 s build)

  • Reproducible kernels: all 48/48 HRX kernel artifacts are byte-identical to the working hand-assembled build's.
  • Speed parity: Vulkan0 tg 148.7 vs 151.4, HRX20 zaya1-8b Q4NX tg 18.3 vs 17.8.
  • Through Lemonade (tests/hrx_lemonade_e2e.sh): the same checkpoint via both recipes, greedy decoding. The test also checks the serving process.
recipe device answer decode
llamacpp Vulkan0 Paris 117.3 tok/s
llamacpp-hrx HRX20 Paris 33.5 tok/s
  • Default build: with ONEBIT_HRX=OFF (what CI runs) it's unchanged, and the smoke test passes.

Not yet (docs/hrx.md)

  • zaya Q4NX through Lemonade: that needs the model registry from step 5.
  • A relocatable package: the HRX libraries load via build-tree RPATHs.
  • No HRX in CI, since the runner has no AMD GPU.

🤖 Generated with Claude Code

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Pinned submodules: ROCm/hrx 0bc22fb (runtime), ROCm/hrx-system 6743075f
  (loom-link with --mode=selective), and the llama.cpp fork on branch
  1bit-engine/hrx-vulkan (3b33c8a9). patches/hrx 0001-0003 hold the local
  changes: build plumbing plus the gfx1151 PM4 probe fix.
- cmake/hrx.cmake (ONEBIT_HRX, off by default): clones each submodule into
  the build tree, patches it there, and builds only what ggml-hrx2 needs.
  hrx_assemble_prefix.cmake reproduces the previously hand-assembled prefix.
- 1bit lemonade uses the resulting llama-server for llamacpp-hrx (HRX20)
  and for llamacpp's Vulkan backend (Vulkan0), replacing only defaults.
- Lemonade delta: an hrx_device option (default HRX0) replaces the
  hard-coded device.
- Verified on Strix Halo from a fresh clone (217 s): 48/48 kernel artifacts
  byte-identical to the working build; decode parity; Lemonade serves the
  same checkpoint on Vulkan0 (117 tok/s) and HRX20 (34 tok/s)
  (tests/hrx_lemonade_e2e.sh).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit 23cc198 into embed-lemonade Sep 22, 2026
2 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the hrx-vulkan branch September 22, 2026 21:48
bong-water-water-bong added a commit that referenced this pull request Sep 22, 2026
…#6)

* Step 1: embed Lemonade's server core in the 1bit binary

- third_party/lemonade: Lemonade v11.9.0, vendored complete, plus the
  local deltas in its UPSTREAM.md (the onebit backend, /v1/registry,
  the hrx-b66 pin, the embed CMake patch). Copied from 1bit-MONSTER main
  (last changed 256e68dd7).
- app/main.cpp: `1bit lemonade [lemond options]`, ported from
  run_embedded_lemonade in tools/unified_server.cpp. FLM pinning is
  dropped (FLM is eliminated); the engine-registry injection returns
  with the NPU engine (step 3).
- tests/smoke_lemonade.sh (ctest + CI): health ok, catalog with
  llamacpp-hrx, system-info. Uses scratch dirs.
- docs/lemonade.md; porting map adds step 5 (every HF model: registry +
  daily census). third_party/ is exempt from our copyright notice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: install libdrm-dev (Lemonade's server core includes libdrm/drm.h)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* smoke: require HRX catalog models only where Lemonade reports HRX supported

CI runners have no AMD GPU, so Lemonade correctly hides llamacpp-hrx there.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Step 2: HRX + Vulkan in one llama.cpp build, wired into Lemonade (#7)

* wip: step 2 HRX + Vulkan build (testing from a fresh clone)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* wip: pin the HRX submodules (gitlinks)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Step 2: HRX + Vulkan in one llama.cpp build, wired into Lemonade

- Pinned submodules: ROCm/hrx 0bc22fb (runtime), ROCm/hrx-system 6743075f
  (loom-link with --mode=selective), and the llama.cpp fork on branch
  1bit-engine/hrx-vulkan (3b33c8a9). patches/hrx 0001-0003 hold the local
  changes: build plumbing plus the gfx1151 PM4 probe fix.
- cmake/hrx.cmake (ONEBIT_HRX, off by default): clones each submodule into
  the build tree, patches it there, and builds only what ggml-hrx2 needs.
  hrx_assemble_prefix.cmake reproduces the previously hand-assembled prefix.
- 1bit lemonade uses the resulting llama-server for llamacpp-hrx (HRX20)
  and for llamacpp's Vulkan backend (Vulkan0), replacing only defaults.
- Lemonade delta: an hrx_device option (default HRX0) replaces the
  hard-coded device.
- Verified on Strix Halo from a fresh clone (217 s): 48/48 kernel artifacts
  byte-identical to the working build; decode parity; Lemonade serves the
  same checkpoint on Vulkan0 (117 tok/s) and HRX20 (34 tok/s)
  (tests/hrx_lemonade_e2e.sh).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Sep 26, 2026
…es (#119)

Was 4 commits behind 1bit/hrx-vulkan-patched, missing every landed fix for
both open HRX issues:

- engine#108 (dense on HRX0, MoE experts on another device): route-ids
  binding-length fix (llama.cpp#7) plus the fused router launch-geometry fix
  (llama.cpp#12) that PR #7's description claimed was included but wasn't -
  verified separately this session against the exact repro
  (-dev HRX0,Vulkan0 -ts 1,0 -ot exps=Vulkan0, Qwen3-Coder-30B-A3B): correct
  and deterministic across repeated prompts, disabling the fused dispatch
  reproduces the issue's documented fallback error exactly.
- engine#115 (HRX flash-attention decode-split all_rejected above 2048 KV
  tokens): capacity-cap fix (llama.cpp#9) - falls through to the general
  wmma kernel above the cap instead of crashing.

Built onebit (-DONEBIT_HRX=ON) against the new pin and re-verified both
fixes end to end before this bump.

Co-authored-by: agent <agent@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
bong-water-water-bong pushed a commit that referenced this pull request Sep 26, 2026
…oints and three HRX fixes

Our llama.cpp #7 (HRX MoE router route-ids binding length), #9 (decode-split
flash-attention capacity bound), #11 (ZAYA sliding-window attention, ZAYA1-74B),
#12 (MoE router experts-per-wave launch geometry) and #14 (ZAYA legacy
checkpoints; --remote fetches chat_template.jinja).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant