diff --git a/README.md b/README.md
index 61c7b026..b4419bd8 100644
--- a/README.md
+++ b/README.md
@@ -50,7 +50,8 @@ Windows and 1bit OS ([docs/releases.md](docs/releases.md)).
> runs from our llama.cpp on Vulkan, HRX and ROCm, matching transformers, at 93 tok/s decode in
> Q4_K_M on Vulkan, and ZAYA1-74B-preview at 35 tok/s
> ([docs/vulkan.md](docs/vulkan.md#zaya1-zyphra-from-our-llamacpp)); the rest of Zyphra's
-> family (Zamba, Zamba2, BlackMamba) runs on Vulkan too. Experimental,
+> family (Zamba, Zamba2, BlackMamba) runs on Vulkan too, and so do its vision models, ZAYA1-VL-8B
+> and Zamba2-VL, through `1bit serve --mmproj`. Experimental,
> and closed source: Qwen3.6-35B-A3B on the NPU through a private add-on, parity against fp64 passes,
> 16.3-16.5 tok/s decode ([docs/npu.md](docs/npu.md#private-routes)). Step 4, the Laya router,
> has landed as an opt-in: `1bit serve --device auto --laya-model
` picks the device per
diff --git a/app/serve.cpp b/app/serve.cpp
index d21cd88d..397b8276 100644
--- a/app/serve.cpp
+++ b/app/serve.cpp
@@ -623,7 +623,9 @@ Launch launch_for(const Options& o, const std::string& device, int child_port) {
}
if (!o.mmproj.empty()) {
if (device == "zinc" || device == "mlx" || device == "ds4") throw std::runtime_error("--mmproj works on the llama.cpp devices (vulkan, hrx, rocm)");
- argv.insert(argv.end(), {"--mmproj", o.mmproj});
+ // an image is decoded as one ubatch: models that attend to it bidirectionally (ZAYA1-VL,
+ // Gemma 3) need all of it in one, and Qwen2.5-VL towers cap an image at 4096 tokens
+ argv.insert(argv.end(), {"--mmproj", o.mmproj, "-b", "4096", "-ub", "4096"});
}
return l;
}
diff --git a/docs/serve.md b/docs/serve.md
index e8a5edc4..75c8de7e 100644
--- a/docs/serve.md
+++ b/docs/serve.md
@@ -100,9 +100,10 @@ then `$ONEBIT_LLAMA_SERVER` / `$ONEBIT_ZINC` / `$ONEBIT_DS4`, then `llama-server
`--mmproj ` hands llama-server a vision projector on the llama.cpp routes
(vulkan, hrx, rocm). Chat messages may then carry `image_url` parts (a `data:` URL or a file
URL), which llama.cpp's mtmd encodes and places in the prompt. The mmproj comes from
-`convert_hf_to_gguf.py --mmproj` on the same checkpoint as the model. Zyphra's Zamba2-VL, for one,
-is in [docs/vulkan.md](vulkan.md). The NPU, ZINC, DwarfStar, MLX and ONNX routes take no
-`--mmproj`.
+`convert_hf_to_gguf.py --mmproj` on the same checkpoint as the model. With it, llama-server runs
+with `-b 4096 -ub 4096`: an image is decoded as one ubatch, which models that attend to an image
+bidirectionally (ZAYA1-VL, Gemma 3) need. Zyphra's Zamba2-VL and ZAYA1-VL-8B are in
+[docs/vulkan.md](vulkan.md). The NPU, ZINC, DwarfStar, MLX and ONNX routes take no `--mmproj`.
## Multi-token prediction (`--mtp`)
diff --git a/docs/vulkan.md b/docs/vulkan.md
index 2b4f219c..ccee72f4 100644
--- a/docs/vulkan.md
+++ b/docs/vulkan.md
@@ -124,6 +124,13 @@ circle and the text "HELLO 42":
The 1.2B and 2.7B misread the text the same way on the CPU and on Vulkan, so the misreading comes
from the models, not from the port.
+GGUFs made before [llama.cpp #17](https://github.com/1bit-MONSTER/llama.cpp/pull/17) carry a chat
+template that left each image inside the user turn instead of in front of it, as Zyphra's
+template has it. llama-server marks images with `<__media___>`, an id random per server, and
+that template knew only `<__media__>`. Reconvert those GGUFs, or serve them with
+`--chat-template-file`. With the new template the image starts at the same prompt position as in
+Zyphra's processor.
+
Large models need `--ctx-size`: without it llama-server allocates the KV cache
for the model's full trained context (262,144 tokens for Qwen3.8), which does not
fit next to Flash-Next's 104 GiB of weights.
@@ -243,3 +250,39 @@ Until [llama.cpp #15](https://github.com/1bit-MONSTER/llama.cpp/pull/15), ZAYA r
results whenever a ubatch held more than one sequence: `llama-perplexity` at its default batch,
and `1bit serve` answering concurrent requests. At four sequences per ubatch, ZAYA1-8B F16's
wikitext perplexity was 58,259 against 27.89 with one; now both are 27.89 on the CPU.
+
+### ZAYA1-VL-8B: images
+
+[ZAYA1-VL-8B](https://huggingface.co/Zyphra/ZAYA1-VL-8B) is ZAYA1 with Qwen2.5-VL's vision tower
+and two additions ([llama.cpp #18](https://github.com/1bit-MONSTER/llama.cpp/pull/18)):
+
+- **Vision-only LoRA.** Rank-8 LoRA on CCA's q, k, both value projections and o_proj, and rank 32
+ on every expert, used on image tokens only. mtmd decodes each image as its own ubatch of
+ embeddings, so that ubatch takes the LoRA and text ubatches run the ZAYA1 graph unchanged.
+- **Bidirectional attention within an image.** The mmproj key `clip.vision.decode_non_causal`
+ makes mtmd decode the image non-causally. The whole image must fit one ubatch, so `1bit serve
+ --mmproj` passes `-b 4096 -ub 4096` (Qwen2.5-VL caps an image at 4,096 tokens).
+
+```
+python third_party/llama.cpp/convert_hf_to_gguf.py --outtype f16 --outfile zaya1-vl-8b-F16.gguf
+python third_party/llama.cpp/convert_hf_to_gguf.py --mmproj --outtype f16 --outfile mmproj-zaya1-vl-8b-F16.gguf
+1bit serve -m zaya1-vl-8b-F16.gguf --mmproj mmproj-zaya1-vl-8b-F16.gguf --device vulkan --ctx-size 8192
+```
+
+Checked on Strix Halo against Zyphra's own code (their transformers branch `zaya1-vl`, FP32 on
+the CPU), teacher-forced on llama.cpp's greedy answers to three image questions (101 tokens).
+The prompts are identical: 208 tokens, with the image at the same position.
+
+| Image decode | Reference | Top-1 agreement |
+|---|---|---|
+| causal (the flag set false) | Zyphra's eager path | 100/101 |
+| bidirectional (the default) | eager, with the image block bidirectional | 100/101 |
+
+Zyphra's flash-attention path, which its paper describes, makes everything up to two tokens past
+the image bidirectional. llama.cpp decodes the image bidirectionally and the tokens around it
+causally.
+
+On a synthetic image (a red square, a blue circle, "HELLO 42") the model reads the text exactly,
+and on a photo it recognises the New York Times front page of the moon landing. F16 decodes at
+38-51 tok/s on Vulkan0. ZAYA1-8B is unchanged: 95/96 teacher-forced, the same wikitext
+perplexity.
diff --git a/registry/architectures.json b/registry/architectures.json
index f5158b6c..953977f3 100644
--- a/registry/architectures.json
+++ b/registry/architectures.json
@@ -1,17 +1,17 @@
{
"about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.",
"sources": {
- "llama.cpp (vulkan)": "2c9d3983d5c1cb68d331f0d8de0c3c738c7caf54",
- "llama.cpp (hrx)": "83e1c41a4d39972662e2d1dac0669160bf45cad4",
+ "llama.cpp (vulkan)": "4a2c0656f14c3ee79f85c9c1936300a498dc8ad1",
+ "llama.cpp (hrx)": "f30cc43995600c69df4f969a92a99721c7709762",
"zinc": "3a35e76d64ebb91e2d82e16ddda20ee865ce1d45"
},
"counts": {
- "hrx": 299,
+ "hrx": 300,
"npu": 1,
- "vulkan": 322,
+ "vulkan": 323,
"zinc": 64,
- "architectures": 322,
- "mapped": 322
+ "architectures": 323,
+ "mapped": 323
},
"architectures": {
"AfmoeForCausalLM": {
@@ -2282,6 +2282,13 @@
"vulkan"
]
},
+ "Zaya1VLForConditionalGeneration": {
+ "gguf": "zaya",
+ "backends": [
+ "hrx",
+ "vulkan"
+ ]
+ },
"ZayaForCausalLM": {
"gguf": "zaya",
"backends": [
diff --git a/third_party/llama.cpp b/third_party/llama.cpp
index 83e1c41a..f30cc439 160000
--- a/third_party/llama.cpp
+++ b/third_party/llama.cpp
@@ -1 +1 @@
-Subproject commit 83e1c41a4d39972662e2d1dac0669160bf45cad4
+Subproject commit f30cc43995600c69df4f969a92a99721c7709762
diff --git a/third_party/llama.cpp-vulkan b/third_party/llama.cpp-vulkan
index 2c9d3983..4a2c0656 160000
--- a/third_party/llama.cpp-vulkan
+++ b/third_party/llama.cpp-vulkan
@@ -1 +1 @@
-Subproject commit 2c9d3983d5c1cb68d331f0d8de0c3c738c7caf54
+Subproject commit 4a2c0656f14c3ee79f85c9c1936300a498dc8ad1