diff --git a/docs/hrx.md b/docs/hrx.md index 5f170aa..24d315f 100644 --- a/docs/hrx.md +++ b/docs/hrx.md @@ -450,9 +450,10 @@ kernels). It is kept for a later port of that work to ggml-hrx. Its zaya architecture has been ported to this branch (ZAYA1). The engine's separate upstream llama.cpp pin for Vulkan (`third_party/llama.cpp-vulkan`) is gone -(RFC #213 stage 3). Architectures only that pin had (Qwen3.8-Flash-Next's `qwen4exp`, Zyphra -Zamba, Zamba2 and BlackMamba, and a few upstream-only ones) are refused by `1bit serve` until they -are ported here; Lemonade's own llamacpp backend serves them meanwhile. This build has no Vulkan +(RFC #213 stage 3). Architectures only that pin had (Zyphra Zamba, Zamba2 and BlackMamba, and a +few upstream-only ones) are refused by `1bit serve` until they are ported here; Lemonade's own +llamacpp backend serves them meanwhile. Qwen3.8-Flash-Next (`qwen4exp`, with its MTP draft head) +was ported at pin `e57beb97`. This build has no Vulkan backend (`GGML_VULKAN=OFF`). ### The GPU's matrix units (WMMA) diff --git a/docs/lemonade.md b/docs/lemonade.md index c76bb34..daca2f9 100644 --- a/docs/lemonade.md +++ b/docs/lemonade.md @@ -122,9 +122,9 @@ Lemonade's to serve: - `1bit serve --device vulkan`, `--device rocm` and `--device zinc` exit at once with that reason, as do `--lean`, `--adaptive`, `--long-model`, `--prefill-device` and `--zinc` ([serve.md](serve.md#removed-devices-and-flags)). -- An architecture only the removed Vulkan build ran (Qwen3.8-Flash-Next's `qwen4exp`, Zyphra Zamba, - Zamba2 and BlackMamba, and a few upstream-only ones) is refused with a pointer to Lemonade's - `llamacpp` recipe. +- An architecture only the removed Vulkan build ran (Zyphra Zamba, Zamba2 and BlackMamba, and a + few upstream-only ones) is refused with a pointer to Lemonade's `llamacpp` recipe. + Qwen3.8-Flash-Next (`qwen4exp`) runs on HRX since llama.cpp pin `e57beb97`. In the fork ([1bit-MONSTER/lemonade](https://github.com/1bit-MONSTER/lemonade)), branch `1bit/onebit-no-vulkan` drops `vulkan` from the `onebit` backend (`onebit.h`, diff --git a/docs/registry.md b/docs/registry.md index 8011b73..d02511f 100644 --- a/docs/registry.md +++ b/docs/registry.md @@ -39,7 +39,7 @@ hand: The `vulkan` column (the upstream llama.cpp release the engine built for Vulkan) left with that build in RFC #213 stage 3. With it went the HF class names only the upstream converter registers: 33 classes, 23 of them for architectures the HRX fork does not run (among -them `qwen4exp`, `zamba`, `zamba2`, `blackmamba`) and 10 aliases of architectures it does run +them `zamba`, `zamba2`, `blackmamba`; `qwen4exp` is back since pin `e57beb97`) and 10 aliases of architectures it does run (DFlash drafters, `exaone-moe`, `nemotron_h_moe`), which a GGUF converted elsewhere still runs on. The `zinc` column (ZINC's `parseArchitecture`; 63 HF architectures on 2026-10-03) left with ZINC in the core strip (2026-10-04). diff --git a/docs/serve.md b/docs/serve.md index 3121df7..83e8bca 100644 --- a/docs/serve.md +++ b/docs/serve.md @@ -68,7 +68,7 @@ fell back to Vulkan (engine #271); with no Vulkan or ROCm build left: | Case | `auto` | |---|---| -| an architecture our llama.cpp does not build: anything outside `gguf_architectures.hrx` in registry/architectures.json, which `tools/registry_build.py` reads from the pinned fork and the build compiles in (among them Qwen3.8-Flash-Next's `qwen4exp`; Zyphra Zamba, Zamba2, BlackMamba; Spark2.5, BailingMoeV3, HRM text, MuseGlimmer, Kimi K3, Maple, GraniteSwitch, Granite SWA, HY v4, MiniMax, Dots3 Note, PocketTTS, Qwen3-TTS) | refused on hrx and cpu, pointing at Lemonade's llamacpp backend; Zyphra and Flash-Next are to be ported to HRX | +| an architecture our llama.cpp does not build: anything outside `gguf_architectures.hrx` in registry/architectures.json, which `tools/registry_build.py` reads from the pinned fork and the build compiles in (among them Zyphra Zamba, Zamba2, BlackMamba; Spark2.5, BailingMoeV3, HRM text, MuseGlimmer, Kimi K3, Maple, GraniteSwitch, Granite SWA, HY v4, MiniMax, Dots3 Note, PocketTTS, Qwen3-TTS) | refused on hrx and cpu, pointing at Lemonade's llamacpp backend; Zyphra is to be ported to HRX (Flash-Next's `qwen4exp` was, at pin `e57beb97`) | | `--moe-slots` | refused: not in this build ([below](#moe-models-larger-than-memory---moe-slots)) | | `--parallel N` > 1 on a gated delta-net model (Qwen3.5, Qwen3.8, Qwen3-Next) | HRX0 with one slot, and a warning on stderr: HRX runs one such sequence at a time (no multi-sequence delta-net yet). `--device cpu` keeps `--parallel N` | | `--mmproj` | HRX0. **Unverified:** vision on HRX is an open RFC #213 gate, to be checked on ZAYA1-VL; if it fails there, `auto` goes back to the CPU for `--mmproj` | @@ -173,8 +173,7 @@ is the model's 644 MiB Q8_0 output matrix, read three times per decode step at d A draft only proposes; the model checks every token, so the head can be cut down to the tokens it is likely to propose. `tools/mtp_draft_vocab.py` writes such a head from Unsloth's *shared* head file and the model's own output rows. Our llama.cpp fork (`qwen4exp`, -`src/models/qwen4exp-draft-vocab.cpp` on the removed Vulkan branch) ran it; it comes back with the -HRX port of Flash-Next: +`src/models/qwen4exp-draft-vocab.cpp`) runs it on HRX since pin `e57beb97`: ```sh python3 tools/mtp_draft_vocab.py \ diff --git a/registry/architectures.json b/registry/architectures.json index 0be110f..cd57dc3 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -1,13 +1,13 @@ { "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { - "llama.cpp (hrx)": "2e5fcf56a24d112733e11627041c5eb54a671bc8" + "llama.cpp (hrx)": "e57beb9721afd72066f541e218893c908cba8bf4" }, "counts": { - "hrx": 290, + "hrx": 292, "npu": 7, - "architectures": 290, - "mapped": 290 + "architectures": 292, + "mapped": 292 }, "architectures": { "AfmoeForCausalLM": { @@ -1458,6 +1458,18 @@ "hrx" ] }, + "Qwen4ExpForCausalLM": { + "gguf": "qwen4exp", + "backends": [ + "hrx" + ] + }, + "Qwen4ExpForConditionalGeneration": { + "gguf": "qwen4exp", + "backends": [ + "hrx" + ] + }, "QwenForCausalLM": { "gguf": "qwen", "backends": [ @@ -1890,6 +1902,7 @@ "qwen3next", "qwen3vl", "qwen3vlmoe", + "qwen4exp", "refact", "rnd1", "rwkv6", diff --git a/tests/device_route.sh b/tests/device_route.sh index d6ff9f0..8e6d663 100755 --- a/tests/device_route.sh +++ b/tests/device_route.sh @@ -21,9 +21,11 @@ # - --device auto with --mtp: the same route (HRX drafts on HRX0 now; it used to keep Vulkan), # - --device auto with --mmproj: the same route as a plain file (HRX0 is unverified for vision, # an open RFC #213 gate to check on ZAYA1-VL), +# - Qwen3.8-Flash-Next's qwen4exp (ported to HRX at llama.cpp pin e57beb97) takes the same route +# as a plain file, # - an architecture our llama.cpp does not build (not in registry/architectures.json's -# gguf_architectures.hrx: Qwen3.8-Flash-Next's qwen4exp, Zyphra Zamba2, a made-up one) is -# refused with a pointer to Lemonade's llamacpp backend, under auto and when asked for by name, +# gguf_architectures.hrx: Zyphra Zamba2, a made-up one) is refused with a pointer to Lemonade's +# llamacpp backend, under auto and when asked for by name, # - --moe-slots is refused as not in this build (it moves to HRX with Flash-Next), # - --parallel 4 on a gated delta-net model: in a build with HRX, auto serves it on HRX0 with one # slot and says why; --device cpu keeps -np 4, @@ -103,9 +105,11 @@ check "--device auto runs --mtp on $where" '[ "$(after "$scratch/mtp.json" --dev run "$scratch/mmproj.json" plain --mmproj "$scratch/head.gguf" check "--device auto runs --mmproj on $where" '[ "$(after "$scratch/mmproj.json" --device)" = "$expected" ] && [ "$(after "$scratch/mmproj.json" --mmproj)" = "$scratch/head.gguf" ]' -check "an architecture our llama.cpp does not map (qwen4exp) is refused under auto" 'refused "Lemonade" flashnext' -check " and on --device hrx" 'refused "architecture qwen4exp is not in" flashnext --device hrx' -check " and on --device cpu (zamba2)" 'refused "architecture zamba2 is not in" zamba --device cpu' +run "$scratch/flashnext.json" flashnext +check "--device auto runs Flash-Next (qwen4exp) on $where" '[ "$(after "$scratch/flashnext.json" --device)" = "$expected" ]' +check "an architecture our llama.cpp does not map (zamba2) is refused under auto" 'refused "Lemonade" zamba' +check " and on --device hrx" 'refused "architecture zamba2 is not in" zamba --device hrx' +check " and on --device cpu" 'refused "architecture zamba2 is not in" zamba --device cpu' check " and any architecture the registry does not list" 'refused "architecture not-an-arch is not in" madeup' check "an H32 file is refused by name" 'refused "H32 file" h32' check " on --device hrx too" 'refused "dropped that format" h32 --device hrx' diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 2e5fcf5..e57beb9 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 2e5fcf56a24d112733e11627041c5eb54a671bc8 +Subproject commit e57beb9721afd72066f541e218893c908cba8bf4