diff --git a/README.md b/README.md index 61c7b026..b4419bd8 100644 --- a/README.md +++ b/README.md @@ -50,7 +50,8 @@ Windows and 1bit OS ([docs/releases.md](docs/releases.md)). > runs from our llama.cpp on Vulkan, HRX and ROCm, matching transformers, at 93 tok/s decode in > Q4_K_M on Vulkan, and ZAYA1-74B-preview at 35 tok/s > ([docs/vulkan.md](docs/vulkan.md#zaya1-zyphra-from-our-llamacpp)); the rest of Zyphra's -> family (Zamba, Zamba2, BlackMamba) runs on Vulkan too. Experimental, +> family (Zamba, Zamba2, BlackMamba) runs on Vulkan too, and so do its vision models, ZAYA1-VL-8B +> and Zamba2-VL, through `1bit serve --mmproj`. Experimental, > and closed source: Qwen3.6-35B-A3B on the NPU through a private add-on, parity against fp64 passes, > 16.3-16.5 tok/s decode ([docs/npu.md](docs/npu.md#private-routes)). Step 4, the Laya router, > has landed as an opt-in: `1bit serve --device auto --laya-model ` picks the device per diff --git a/app/serve.cpp b/app/serve.cpp index d21cd88d..397b8276 100644 --- a/app/serve.cpp +++ b/app/serve.cpp @@ -623,7 +623,9 @@ Launch launch_for(const Options& o, const std::string& device, int child_port) { } if (!o.mmproj.empty()) { if (device == "zinc" || device == "mlx" || device == "ds4") throw std::runtime_error("--mmproj works on the llama.cpp devices (vulkan, hrx, rocm)"); - argv.insert(argv.end(), {"--mmproj", o.mmproj}); + // an image is decoded as one ubatch: models that attend to it bidirectionally (ZAYA1-VL, + // Gemma 3) need all of it in one, and Qwen2.5-VL towers cap an image at 4096 tokens + argv.insert(argv.end(), {"--mmproj", o.mmproj, "-b", "4096", "-ub", "4096"}); } return l; } diff --git a/docs/serve.md b/docs/serve.md index e8a5edc4..75c8de7e 100644 --- a/docs/serve.md +++ b/docs/serve.md @@ -100,9 +100,10 @@ then `$ONEBIT_LLAMA_SERVER` / `$ONEBIT_ZINC` / `$ONEBIT_DS4`, then `llama-server `--mmproj ` hands llama-server a vision projector on the llama.cpp routes (vulkan, hrx, rocm). Chat messages may then carry `image_url` parts (a `data:` URL or a file URL), which llama.cpp's mtmd encodes and places in the prompt. The mmproj comes from -`convert_hf_to_gguf.py --mmproj` on the same checkpoint as the model. Zyphra's Zamba2-VL, for one, -is in [docs/vulkan.md](vulkan.md). The NPU, ZINC, DwarfStar, MLX and ONNX routes take no -`--mmproj`. +`convert_hf_to_gguf.py --mmproj` on the same checkpoint as the model. With it, llama-server runs +with `-b 4096 -ub 4096`: an image is decoded as one ubatch, which models that attend to an image +bidirectionally (ZAYA1-VL, Gemma 3) need. Zyphra's Zamba2-VL and ZAYA1-VL-8B are in +[docs/vulkan.md](vulkan.md). The NPU, ZINC, DwarfStar, MLX and ONNX routes take no `--mmproj`. ## Multi-token prediction (`--mtp`) diff --git a/docs/vulkan.md b/docs/vulkan.md index 2b4f219c..ccee72f4 100644 --- a/docs/vulkan.md +++ b/docs/vulkan.md @@ -124,6 +124,13 @@ circle and the text "HELLO 42": The 1.2B and 2.7B misread the text the same way on the CPU and on Vulkan, so the misreading comes from the models, not from the port. +GGUFs made before [llama.cpp #17](https://github.com/1bit-MONSTER/llama.cpp/pull/17) carry a chat +template that left each image inside the user turn instead of in front of it, as Zyphra's +template has it. llama-server marks images with `<__media___>`, an id random per server, and +that template knew only `<__media__>`. Reconvert those GGUFs, or serve them with +`--chat-template-file`. With the new template the image starts at the same prompt position as in +Zyphra's processor. + Large models need `--ctx-size`: without it llama-server allocates the KV cache for the model's full trained context (262,144 tokens for Qwen3.8), which does not fit next to Flash-Next's 104 GiB of weights. @@ -243,3 +250,39 @@ Until [llama.cpp #15](https://github.com/1bit-MONSTER/llama.cpp/pull/15), ZAYA r results whenever a ubatch held more than one sequence: `llama-perplexity` at its default batch, and `1bit serve` answering concurrent requests. At four sequences per ubatch, ZAYA1-8B F16's wikitext perplexity was 58,259 against 27.89 with one; now both are 27.89 on the CPU. + +### ZAYA1-VL-8B: images + +[ZAYA1-VL-8B](https://huggingface.co/Zyphra/ZAYA1-VL-8B) is ZAYA1 with Qwen2.5-VL's vision tower +and two additions ([llama.cpp #18](https://github.com/1bit-MONSTER/llama.cpp/pull/18)): + +- **Vision-only LoRA.** Rank-8 LoRA on CCA's q, k, both value projections and o_proj, and rank 32 + on every expert, used on image tokens only. mtmd decodes each image as its own ubatch of + embeddings, so that ubatch takes the LoRA and text ubatches run the ZAYA1 graph unchanged. +- **Bidirectional attention within an image.** The mmproj key `clip.vision.decode_non_causal` + makes mtmd decode the image non-causally. The whole image must fit one ubatch, so `1bit serve + --mmproj` passes `-b 4096 -ub 4096` (Qwen2.5-VL caps an image at 4,096 tokens). + +``` +python third_party/llama.cpp/convert_hf_to_gguf.py --outtype f16 --outfile zaya1-vl-8b-F16.gguf +python third_party/llama.cpp/convert_hf_to_gguf.py --mmproj --outtype f16 --outfile mmproj-zaya1-vl-8b-F16.gguf +1bit serve -m zaya1-vl-8b-F16.gguf --mmproj mmproj-zaya1-vl-8b-F16.gguf --device vulkan --ctx-size 8192 +``` + +Checked on Strix Halo against Zyphra's own code (their transformers branch `zaya1-vl`, FP32 on +the CPU), teacher-forced on llama.cpp's greedy answers to three image questions (101 tokens). +The prompts are identical: 208 tokens, with the image at the same position. + +| Image decode | Reference | Top-1 agreement | +|---|---|---| +| causal (the flag set false) | Zyphra's eager path | 100/101 | +| bidirectional (the default) | eager, with the image block bidirectional | 100/101 | + +Zyphra's flash-attention path, which its paper describes, makes everything up to two tokens past +the image bidirectional. llama.cpp decodes the image bidirectionally and the tokens around it +causally. + +On a synthetic image (a red square, a blue circle, "HELLO 42") the model reads the text exactly, +and on a photo it recognises the New York Times front page of the moon landing. F16 decodes at +38-51 tok/s on Vulkan0. ZAYA1-8B is unchanged: 95/96 teacher-forced, the same wikitext +perplexity. diff --git a/registry/architectures.json b/registry/architectures.json index f5158b6c..953977f3 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -1,17 +1,17 @@ { "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { - "llama.cpp (vulkan)": "2c9d3983d5c1cb68d331f0d8de0c3c738c7caf54", - "llama.cpp (hrx)": "83e1c41a4d39972662e2d1dac0669160bf45cad4", + "llama.cpp (vulkan)": "4a2c0656f14c3ee79f85c9c1936300a498dc8ad1", + "llama.cpp (hrx)": "f30cc43995600c69df4f969a92a99721c7709762", "zinc": "3a35e76d64ebb91e2d82e16ddda20ee865ce1d45" }, "counts": { - "hrx": 299, + "hrx": 300, "npu": 1, - "vulkan": 322, + "vulkan": 323, "zinc": 64, - "architectures": 322, - "mapped": 322 + "architectures": 323, + "mapped": 323 }, "architectures": { "AfmoeForCausalLM": { @@ -2282,6 +2282,13 @@ "vulkan" ] }, + "Zaya1VLForConditionalGeneration": { + "gguf": "zaya", + "backends": [ + "hrx", + "vulkan" + ] + }, "ZayaForCausalLM": { "gguf": "zaya", "backends": [ diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 83e1c41a..f30cc439 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 83e1c41a4d39972662e2d1dac0669160bf45cad4 +Subproject commit f30cc43995600c69df4f969a92a99721c7709762 diff --git a/third_party/llama.cpp-vulkan b/third_party/llama.cpp-vulkan index 2c9d3983..4a2c0656 160000 --- a/third_party/llama.cpp-vulkan +++ b/third_party/llama.cpp-vulkan @@ -1 +1 @@ -Subproject commit 2c9d3983d5c1cb68d331f0d8de0c3c738c7caf54 +Subproject commit 4a2c0656f14c3ee79f85c9c1936300a498dc8ad1