Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions docs/hrx.md
Original file line number Diff line number Diff line change
Expand Up @@ -450,9 +450,10 @@ kernels). It is kept for a later port of that work to ggml-hrx. Its zaya
architecture has been ported to this branch (ZAYA1).

The engine's separate upstream llama.cpp pin for Vulkan (`third_party/llama.cpp-vulkan`) is gone
(RFC #213 stage 3). Architectures only that pin had (Qwen3.8-Flash-Next's `qwen4exp`, Zyphra
Zamba, Zamba2 and BlackMamba, and a few upstream-only ones) are refused by `1bit serve` until they
are ported here; Lemonade's own llamacpp backend serves them meanwhile. This build has no Vulkan
(RFC #213 stage 3). Architectures only that pin had (Zyphra Zamba, Zamba2 and BlackMamba, and a
few upstream-only ones) are refused by `1bit serve` until they are ported here; Lemonade's own
llamacpp backend serves them meanwhile. Qwen3.8-Flash-Next (`qwen4exp`, with its MTP draft head)
was ported at pin `e57beb97`. This build has no Vulkan
backend (`GGML_VULKAN=OFF`).

### The GPU's matrix units (WMMA)
Expand Down
6 changes: 3 additions & 3 deletions docs/lemonade.md
Original file line number Diff line number Diff line change
Expand Up @@ -122,9 +122,9 @@ Lemonade's to serve:
- `1bit serve --device vulkan`, `--device rocm` and `--device zinc` exit at once with that reason, as
do `--lean`, `--adaptive`, `--long-model`, `--prefill-device` and `--zinc`
([serve.md](serve.md#removed-devices-and-flags)).
- An architecture only the removed Vulkan build ran (Qwen3.8-Flash-Next's `qwen4exp`, Zyphra Zamba,
Zamba2 and BlackMamba, and a few upstream-only ones) is refused with a pointer to Lemonade's
`llamacpp` recipe.
- An architecture only the removed Vulkan build ran (Zyphra Zamba, Zamba2 and BlackMamba, and a
few upstream-only ones) is refused with a pointer to Lemonade's `llamacpp` recipe.
Qwen3.8-Flash-Next (`qwen4exp`) runs on HRX since llama.cpp pin `e57beb97`.

In the fork ([1bit-MONSTER/lemonade](https://github.com/1bit-MONSTER/lemonade)), branch
`1bit/onebit-no-vulkan` drops `vulkan` from the `onebit` backend (`onebit.h`,
Expand Down
2 changes: 1 addition & 1 deletion docs/registry.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ hand:
The `vulkan` column (the upstream llama.cpp release the engine built for Vulkan) left with that
build in RFC #213 stage 3. With it went the HF class names only the upstream
converter registers: 33 classes, 23 of them for architectures the HRX fork does not run (among
them `qwen4exp`, `zamba`, `zamba2`, `blackmamba`) and 10 aliases of architectures it does run
them `zamba`, `zamba2`, `blackmamba`; `qwen4exp` is back since pin `e57beb97`) and 10 aliases of architectures it does run
(DFlash drafters, `exaone-moe`, `nemotron_h_moe`), which a GGUF converted elsewhere still runs on.
The `zinc` column (ZINC's `parseArchitecture`; 63 HF architectures on 2026-10-03) left with ZINC in
the core strip (2026-10-04).
Expand Down
5 changes: 2 additions & 3 deletions docs/serve.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,7 +68,7 @@ fell back to Vulkan (engine #271); with no Vulkan or ROCm build left:

| Case | `auto` |
|---|---|
| an architecture our llama.cpp does not build: anything outside `gguf_architectures.hrx` in registry/architectures.json, which `tools/registry_build.py` reads from the pinned fork and the build compiles in (among them Qwen3.8-Flash-Next's `qwen4exp`; Zyphra Zamba, Zamba2, BlackMamba; Spark2.5, BailingMoeV3, HRM text, MuseGlimmer, Kimi K3, Maple, GraniteSwitch, Granite SWA, HY v4, MiniMax, Dots3 Note, PocketTTS, Qwen3-TTS) | refused on hrx and cpu, pointing at Lemonade's llamacpp backend; Zyphra and Flash-Next are to be ported to HRX |
| an architecture our llama.cpp does not build: anything outside `gguf_architectures.hrx` in registry/architectures.json, which `tools/registry_build.py` reads from the pinned fork and the build compiles in (among them Zyphra Zamba, Zamba2, BlackMamba; Spark2.5, BailingMoeV3, HRM text, MuseGlimmer, Kimi K3, Maple, GraniteSwitch, Granite SWA, HY v4, MiniMax, Dots3 Note, PocketTTS, Qwen3-TTS) | refused on hrx and cpu, pointing at Lemonade's llamacpp backend; Zyphra is to be ported to HRX (Flash-Next's `qwen4exp` was, at pin `e57beb97`) |
| `--moe-slots` | refused: not in this build ([below](#moe-models-larger-than-memory---moe-slots)) |
| `--parallel N` > 1 on a gated delta-net model (Qwen3.5, Qwen3.8, Qwen3-Next) | HRX0 with one slot, and a warning on stderr: HRX runs one such sequence at a time (no multi-sequence delta-net yet). `--device cpu` keeps `--parallel N` |
| `--mmproj` | HRX0. **Unverified:** vision on HRX is an open RFC #213 gate, to be checked on ZAYA1-VL; if it fails there, `auto` goes back to the CPU for `--mmproj` |
Expand Down Expand Up @@ -173,8 +173,7 @@ is the model's 644 MiB Q8_0 output matrix, read three times per decode step at d
A draft only proposes; the model checks every token, so the head can be cut down to the
tokens it is likely to propose. `tools/mtp_draft_vocab.py` writes such a head from Unsloth's
*shared* head file and the model's own output rows. Our llama.cpp fork (`qwen4exp`,
`src/models/qwen4exp-draft-vocab.cpp` on the removed Vulkan branch) ran it; it comes back with the
HRX port of Flash-Next:
`src/models/qwen4exp-draft-vocab.cpp`) runs it on HRX since pin `e57beb97`:

```sh
python3 tools/mtp_draft_vocab.py \
Expand Down
21 changes: 17 additions & 4 deletions registry/architectures.json
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
{
"about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.",
"sources": {
"llama.cpp (hrx)": "2e5fcf56a24d112733e11627041c5eb54a671bc8"
"llama.cpp (hrx)": "e57beb9721afd72066f541e218893c908cba8bf4"
},
"counts": {
"hrx": 290,
"hrx": 292,
"npu": 7,
"architectures": 290,
"mapped": 290
"architectures": 292,
"mapped": 292
},
"architectures": {
"AfmoeForCausalLM": {
Expand Down Expand Up @@ -1458,6 +1458,18 @@
"hrx"
]
},
"Qwen4ExpForCausalLM": {
"gguf": "qwen4exp",
"backends": [
"hrx"
]
},
"Qwen4ExpForConditionalGeneration": {
"gguf": "qwen4exp",
"backends": [
"hrx"
]
},
"QwenForCausalLM": {
"gguf": "qwen",
"backends": [
Expand Down Expand Up @@ -1890,6 +1902,7 @@
"qwen3next",
"qwen3vl",
"qwen3vlmoe",
"qwen4exp",
"refact",
"rnd1",
"rwkv6",
Expand Down
14 changes: 9 additions & 5 deletions tests/device_route.sh
Original file line number Diff line number Diff line change
Expand Up @@ -21,9 +21,11 @@
# - --device auto with --mtp: the same route (HRX drafts on HRX0 now; it used to keep Vulkan),
# - --device auto with --mmproj: the same route as a plain file (HRX0 is unverified for vision,
# an open RFC #213 gate to check on ZAYA1-VL),
# - Qwen3.8-Flash-Next's qwen4exp (ported to HRX at llama.cpp pin e57beb97) takes the same route
# as a plain file,
# - an architecture our llama.cpp does not build (not in registry/architectures.json's
# gguf_architectures.hrx: Qwen3.8-Flash-Next's qwen4exp, Zyphra Zamba2, a made-up one) is
# refused with a pointer to Lemonade's llamacpp backend, under auto and when asked for by name,
# gguf_architectures.hrx: Zyphra Zamba2, a made-up one) is refused with a pointer to Lemonade's
# llamacpp backend, under auto and when asked for by name,
# - --moe-slots is refused as not in this build (it moves to HRX with Flash-Next),
# - --parallel 4 on a gated delta-net model: in a build with HRX, auto serves it on HRX0 with one
# slot and says why; --device cpu keeps -np 4,
Expand Down Expand Up @@ -103,9 +105,11 @@ check "--device auto runs --mtp on $where" '[ "$(after "$scratch/mtp.json" --dev
run "$scratch/mmproj.json" plain --mmproj "$scratch/head.gguf"
check "--device auto runs --mmproj on $where" '[ "$(after "$scratch/mmproj.json" --device)" = "$expected" ] && [ "$(after "$scratch/mmproj.json" --mmproj)" = "$scratch/head.gguf" ]'

check "an architecture our llama.cpp does not map (qwen4exp) is refused under auto" 'refused "Lemonade" flashnext'
check " and on --device hrx" 'refused "architecture qwen4exp is not in" flashnext --device hrx'
check " and on --device cpu (zamba2)" 'refused "architecture zamba2 is not in" zamba --device cpu'
run "$scratch/flashnext.json" flashnext
check "--device auto runs Flash-Next (qwen4exp) on $where" '[ "$(after "$scratch/flashnext.json" --device)" = "$expected" ]'
check "an architecture our llama.cpp does not map (zamba2) is refused under auto" 'refused "Lemonade" zamba'
check " and on --device hrx" 'refused "architecture zamba2 is not in" zamba --device hrx'
check " and on --device cpu" 'refused "architecture zamba2 is not in" zamba --device cpu'
check " and any architecture the registry does not list" 'refused "architecture not-an-arch is not in" madeup'
check "an H32 file is refused by name" 'refused "H32 file" h32'
check " on --device hrx too" 'refused "dropped that format" h32 --device hrx'
Expand Down
2 changes: 1 addition & 1 deletion third_party/llama.cpp
Submodule llama.cpp updated 84 files
+13 −0 common/arg.cpp
+1 −0 common/common.cpp
+2 −0 common/common.h
+60 −4 common/speculative.cpp
+3 −0 conversion/__init__.py
+6 −2 conversion/base.py
+195 −0 conversion/qwen4exp.py
+10 −0 ggml/include/ggml.h
+42 −14 ggml/src/ggml-cpu/ops.cpp
+36 −14 ggml/src/ggml-cuda/dsv4-hc.cu
+1 −1 ggml/src/ggml-cuda/ggml-cuda.cu
+1 −0 ggml/src/ggml-hrx/CMakeLists.txt
+10 −0 ggml/src/ggml-hrx/dispatch_registration/common/dispatch-attention-sink.cpp
+2 −0 ggml/src/ggml-hrx/dispatch_registration/common/dispatch-common.cpp
+2 −1 ggml/src/ggml-hrx/ggml-hrx.cpp
+108 −0 ggml/src/ggml-hrx/hip/README.md
+108 −0 ggml/src/ggml-hrx/hip/dispatch-hip-scale.cpp
+57 −0 ggml/src/ggml-hrx/hip/embed_hip_code_objects.py
+134 −0 ggml/src/ggml-hrx/hip/ggml-hrx-hip.cmake
+51 −0 ggml/src/ggml-hrx/hip/hip-capabilities.cpp
+36 −0 ggml/src/ggml-hrx/hip/hip-capabilities.h
+65 −0 ggml/src/ggml-hrx/hip/hip-code-objects.cpp
+34 −0 ggml/src/ggml-hrx/hip/hip-code-objects.h
+40 −0 ggml/src/ggml-hrx/hip/hip-dispatches.cpp
+58 −0 ggml/src/ggml-hrx/hip/hip-dispatches.h
+69 −0 ggml/src/ggml-hrx/hip/hip-kernel-loader.cpp
+43 −0 ggml/src/ggml-hrx/hip/hip-kernel-loader.h
+103 −0 ggml/src/ggml-hrx/hip/hip-kernel-registry.cpp
+76 −0 ggml/src/ggml-hrx/hip/hip-kernel-registry.h
+126 −0 ggml/src/ggml-hrx/hip/hip-smoke.cpp
+32 −0 ggml/src/ggml-hrx/hip/kernels/hip_scale_f32.hip
+50 −0 ggml/src/ggml-hrx/hip/kernels/hip_smoke.hip
+4 −0 ggml/src/ggml-hrx/kernel-corpus/kernel-corpus.cpp
+31 −35 ggml/src/ggml-hrx/kernel-corpus/kernels/loom-libs/motifs/mul_mat_f32_f32_wmma_core.loom
+31 −61 ggml/src/ggml-hrx/kernel-corpus/kernels/loom-libs/motifs/mul_mat_id_f32_f32_wmma_core.loom
+45 −53 ggml/src/ggml-hrx/kernel-corpus/kernels/loom-libs/ops/mul_mat_id_swiglu_f32_f32_wmma.loom
+45 −53 ggml/src/ggml-hrx/kernel-corpus/kernels/loom-libs/ops/mul_mat_swiglu_f32_f32_wmma.loom
+19 −4 ggml/src/ggml-hrx/runtime/command-program-executor.cpp
+4 −0 ggml/src/ggml-hrx/runtime/command-program-executor.h
+33 −5 ggml/src/ggml-hrx/runtime/hrx-sleeping-wait.cpp
+4 −0 ggml/src/ggml-hrx/runtime/hrx-sleeping-wait.h
+9 −0 ggml/src/ggml-hrx/runtime/kernel-executable-cache.cpp
+1 −0 ggml/src/ggml-hrx/runtime/loom-kernel-jit.h
+38 −10 ggml/src/ggml.c
+115 −0 gguf-py/gguf/constants.py
+34 −0 gguf-py/gguf/gguf_writer.py
+61 −0 gguf-py/gguf/lazy.py
+68 −0 gguf-py/gguf/tensor_mapping.py
+10 −2 include/llama.h
+1 −0 src/CMakeLists.txt
+55 −0 src/llama-arch.cpp
+32 −0 src/llama-arch.h
+9 −4 src/llama-context.cpp
+22 −1 src/llama-hparams.cpp
+27 −0 src/llama-hparams.h
+192 −21 src/llama-kv-cache.cpp
+32 −2 src/llama-kv-cache.h
+48 −26 src/llama-kv-cells.h
+679 −0 src/llama-memory-hybrid-idx.cpp
+160 −0 src/llama-memory-hybrid-idx.h
+56 −4 src/llama-memory-recurrent.cpp
+4 −0 src/llama-memory-recurrent.h
+58 −12 src/llama-mmap.cpp
+6 −1 src/llama-mmap.h
+120 −60 src/llama-model-loader.cpp
+57 −4 src/llama-model-loader.h
+56 −0 src/llama-model-saver.cpp
+1 −0 src/llama-model-saver.h
+66 −7 src/llama-model.cpp
+27 −0 src/llama-model.h
+21 −1 src/llama-quant.cpp
+2 −0 src/llama.cpp
+1 −1 src/models/gemma4.cpp
+123 −0 src/models/models.h
+71 −0 src/models/qwen4exp-draft-vocab.cpp
+40 −0 src/models/qwen4exp-draft-vocab.h
+1,579 −0 src/models/qwen4exp.cpp
+34 −11 tests/test-backend-ops.cpp
+37 −3 tests/test-llama-archs.cpp
+2 −1 tools/cli/README.md
+2 −0 tools/completion/README.md
+1 −0 tools/llama-bench/README.md
+66 −3 tools/llama-bench/llama-bench.cpp
+2 −0 tools/server/README.md
Loading