Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,8 @@ Windows and 1bit OS ([docs/releases.md](docs/releases.md)).
> runs from our llama.cpp on Vulkan, HRX and ROCm, matching transformers, at 93 tok/s decode in
> Q4_K_M on Vulkan, and ZAYA1-74B-preview at 35 tok/s
> ([docs/vulkan.md](docs/vulkan.md#zaya1-zyphra-from-our-llamacpp)); the rest of Zyphra's
> family (Zamba, Zamba2, BlackMamba) runs on Vulkan too. Experimental,
> family (Zamba, Zamba2, BlackMamba) runs on Vulkan too, and so do its vision models, ZAYA1-VL-8B
> and Zamba2-VL, through `1bit serve --mmproj`. Experimental,
> and closed source: Qwen3.6-35B-A3B on the NPU through a private add-on, parity against fp64 passes,
> 16.3-16.5 tok/s decode ([docs/npu.md](docs/npu.md#private-routes)). Step 4, the Laya router,
> has landed as an opt-in: `1bit serve --device auto --laya-model <dir>` picks the device per
Expand Down
4 changes: 3 additions & 1 deletion app/serve.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -623,7 +623,9 @@ Launch launch_for(const Options& o, const std::string& device, int child_port) {
}
if (!o.mmproj.empty()) {
if (device == "zinc" || device == "mlx" || device == "ds4") throw std::runtime_error("--mmproj works on the llama.cpp devices (vulkan, hrx, rocm)");
argv.insert(argv.end(), {"--mmproj", o.mmproj});
// an image is decoded as one ubatch: models that attend to it bidirectionally (ZAYA1-VL,
// Gemma 3) need all of it in one, and Qwen2.5-VL towers cap an image at 4096 tokens
argv.insert(argv.end(), {"--mmproj", o.mmproj, "-b", "4096", "-ub", "4096"});
}
return l;
}
Expand Down
7 changes: 4 additions & 3 deletions docs/serve.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,9 +100,10 @@ then `$ONEBIT_LLAMA_SERVER` / `$ONEBIT_ZINC` / `$ONEBIT_DS4`, then `llama-server
`--mmproj <mmproj.gguf>` hands llama-server a vision projector on the llama.cpp routes
(vulkan, hrx, rocm). Chat messages may then carry `image_url` parts (a `data:` URL or a file
URL), which llama.cpp's mtmd encodes and places in the prompt. The mmproj comes from
`convert_hf_to_gguf.py --mmproj` on the same checkpoint as the model. Zyphra's Zamba2-VL, for one,
is in [docs/vulkan.md](vulkan.md). The NPU, ZINC, DwarfStar, MLX and ONNX routes take no
`--mmproj`.
`convert_hf_to_gguf.py --mmproj` on the same checkpoint as the model. With it, llama-server runs
with `-b 4096 -ub 4096`: an image is decoded as one ubatch, which models that attend to an image
bidirectionally (ZAYA1-VL, Gemma 3) need. Zyphra's Zamba2-VL and ZAYA1-VL-8B are in
[docs/vulkan.md](vulkan.md). The NPU, ZINC, DwarfStar, MLX and ONNX routes take no `--mmproj`.

## Multi-token prediction (`--mtp`)

Expand Down
43 changes: 43 additions & 0 deletions docs/vulkan.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,6 +124,13 @@ circle and the text "HELLO 42":
The 1.2B and 2.7B misread the text the same way on the CPU and on Vulkan, so the misreading comes
from the models, not from the port.

GGUFs made before [llama.cpp #17](https://github.com/1bit-MONSTER/llama.cpp/pull/17) carry a chat
template that left each image inside the user turn instead of in front of it, as Zyphra's
template has it. llama-server marks images with `<__media_<id>__>`, an id random per server, and
that template knew only `<__media__>`. Reconvert those GGUFs, or serve them with
`--chat-template-file`. With the new template the image starts at the same prompt position as in
Zyphra's processor.

Large models need `--ctx-size`: without it llama-server allocates the KV cache
for the model's full trained context (262,144 tokens for Qwen3.8), which does not
fit next to Flash-Next's 104 GiB of weights.
Expand Down Expand Up @@ -243,3 +250,39 @@ Until [llama.cpp #15](https://github.com/1bit-MONSTER/llama.cpp/pull/15), ZAYA r
results whenever a ubatch held more than one sequence: `llama-perplexity` at its default batch,
and `1bit serve` answering concurrent requests. At four sequences per ubatch, ZAYA1-8B F16's
wikitext perplexity was 58,259 against 27.89 with one; now both are 27.89 on the CPU.

### ZAYA1-VL-8B: images

[ZAYA1-VL-8B](https://huggingface.co/Zyphra/ZAYA1-VL-8B) is ZAYA1 with Qwen2.5-VL's vision tower
and two additions ([llama.cpp #18](https://github.com/1bit-MONSTER/llama.cpp/pull/18)):

- **Vision-only LoRA.** Rank-8 LoRA on CCA's q, k, both value projections and o_proj, and rank 32
on every expert, used on image tokens only. mtmd decodes each image as its own ubatch of
embeddings, so that ubatch takes the LoRA and text ubatches run the ZAYA1 graph unchanged.
- **Bidirectional attention within an image.** The mmproj key `clip.vision.decode_non_causal`
makes mtmd decode the image non-causally. The whole image must fit one ubatch, so `1bit serve
--mmproj` passes `-b 4096 -ub 4096` (Qwen2.5-VL caps an image at 4,096 tokens).

```
python third_party/llama.cpp/convert_hf_to_gguf.py <Zyphra/ZAYA1-VL-8B> --outtype f16 --outfile zaya1-vl-8b-F16.gguf
python third_party/llama.cpp/convert_hf_to_gguf.py <Zyphra/ZAYA1-VL-8B> --mmproj --outtype f16 --outfile mmproj-zaya1-vl-8b-F16.gguf
1bit serve -m zaya1-vl-8b-F16.gguf --mmproj mmproj-zaya1-vl-8b-F16.gguf --device vulkan --ctx-size 8192
```

Checked on Strix Halo against Zyphra's own code (their transformers branch `zaya1-vl`, FP32 on
the CPU), teacher-forced on llama.cpp's greedy answers to three image questions (101 tokens).
The prompts are identical: 208 tokens, with the image at the same position.

| Image decode | Reference | Top-1 agreement |
|---|---|---|
| causal (the flag set false) | Zyphra's eager path | 100/101 |
| bidirectional (the default) | eager, with the image block bidirectional | 100/101 |

Zyphra's flash-attention path, which its paper describes, makes everything up to two tokens past
the image bidirectional. llama.cpp decodes the image bidirectionally and the tokens around it
causally.

On a synthetic image (a red square, a blue circle, "HELLO 42") the model reads the text exactly,
and on a photo it recognises the New York Times front page of the moon landing. F16 decodes at
38-51 tok/s on Vulkan0. ZAYA1-8B is unchanged: 95/96 teacher-forced, the same wikitext
perplexity.
19 changes: 13 additions & 6 deletions registry/architectures.json
Original file line number Diff line number Diff line change
@@ -1,17 +1,17 @@
{
"about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.",
"sources": {
"llama.cpp (vulkan)": "2c9d3983d5c1cb68d331f0d8de0c3c738c7caf54",
"llama.cpp (hrx)": "83e1c41a4d39972662e2d1dac0669160bf45cad4",
"llama.cpp (vulkan)": "4a2c0656f14c3ee79f85c9c1936300a498dc8ad1",
"llama.cpp (hrx)": "f30cc43995600c69df4f969a92a99721c7709762",
"zinc": "3a35e76d64ebb91e2d82e16ddda20ee865ce1d45"
},
"counts": {
"hrx": 299,
"hrx": 300,
"npu": 1,
"vulkan": 322,
"vulkan": 323,
"zinc": 64,
"architectures": 322,
"mapped": 322
"architectures": 323,
"mapped": 323
},
"architectures": {
"AfmoeForCausalLM": {
Expand Down Expand Up @@ -2282,6 +2282,13 @@
"vulkan"
]
},
"Zaya1VLForConditionalGeneration": {
"gguf": "zaya",
"backends": [
"hrx",
"vulkan"
]
},
"ZayaForCausalLM": {
"gguf": "zaya",
"backends": [
Expand Down
2 changes: 1 addition & 1 deletion third_party/llama.cpp-vulkan
Loading