Skip to content

Zyphra ZAYA1-VL-8B (pin llama.cpp f30cc43), Zamba2-VL template fix (pin llama.cpp-vulkan 4a2c0656) - #149

Merged
bong-water-water-bong merged 3 commits into
mainfrom
zaya1-vl
Sep 26, 2026
Merged

bong-water-water-bong merged 3 commits into
mainfrom
zaya1-vl

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Zyphra ZAYA1-VL-8B through 1bit serve --mmproj, plus the Zamba2-VL chat-template fix.

Pins

serve: with --mmproj, llama-server gets -b 4096 -ub 4096, so each image decodes as a single ubatch. ZAYA1-VL's bidirectional attention needs that.

Checked on Strix Halo (F16 text, F16 mmproj) against Zyphra's own code (their transformers branch zaya1-vl, FP32, CPU), teacher-forced on llama.cpp's greedy answers to three image questions (101 tokens). The prompts are identical: 208 tokens, with the image at the same position.

Image decode Top-1 agreement
causal, vs Zyphra's eager path 100/101
bidirectional (the default), vs eager with the image block bidirectional 100/101
  • On the test image the model reads "HELLO 42" exactly, and on a photo it recognises the 1969 New York Times moon-landing front page.
  • F16 decodes at 38–51 tok/s on Vulkan0.
  • ZAYA1-8B is unchanged: 95/96 teacher-forced, and wikitext perplexity 32.5834 at one sequence per batch.
  • Zamba2-VL-2.7B with the new template: the image starts at prompt index 2, as in Zyphra's processor (it was 4), and the answers are still correct.

Docs: docs/vulkan.md gets a ZAYA1-VL-8B section and a note on Zamba2-VL GGUFs made before #17. docs/serve.md and the README are updated too. registry/architectures.json is regenerated and now maps Zaya1VLForConditionalGeneration.

🤖 Generated with Claude Code

bong-water-water-bong and others added 3 commits September 26, 2026 13:42
… image in one ubatch

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…Zamba2-VL template, #17)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…656 (Zaya1VLForConditionalGeneration mapped)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 26, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit cffa08f

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

18 - Partially compliant

Compliant requirements:

  • Pin third_party/laya to NandhaKishorM/laya main 1e28ac20 (Apache-2.0)
  • Pin config/laya.json to Hugging Face convaiinnovations/laya revision aa8c91ca
  • Add scripts/fetch-laya.sh that downloads exactly that revision and verifies every file
  • Add bump-laya.yml that runs daily and moves the source and the revision together
  • Update docs/laya.md and mark PORTING step 4 as pinned

Non-compliant requirements:

  • None

Requires further human verification:

  • None

17 - Partially compliant

Compliant requirements:

  • Pin third_party/linux to upstream torvalds/linux v7.3-rc4 (93f51579)
  • Pin config/kernel/strixhalo.config to Strix Halo's current kernel config
  • Add scripts/build-kernel.sh that builds with LLVM and out-of-tree
  • Add bump-linux.yml that moves the pin to the newest upstream tag
  • Update docs/kernel.md and docs/PORTING.md

Non-compliant requirements:

  • None

Requires further human verification:

  • None

19 - Partially compliant

Compliant requirements:

  • Pin third_party/tokenizers to huggingface/tokenizers v0.23.2 (88a4498a)
  • Add hf_tokenizers/ C ABI over it
  • Pin Rust release in rust-toolchain.toml and Cargo.lock
  • Add -DONEBIT_HF_TOKENIZERS=ON build option
  • Add tools/make_tokenizer_golden.py to make goldens with Python tokenizers==0.23.2
  • Add bump-tokenizers.yml to move the pin to each new release

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Incorrect mmproj argument handling

The --mmproj argument handling in app/serve.cpp introduces a potential issue where the -b 4096 -ub 4096 flags are always added when --mmproj is specified, regardless of whether the device supports it. This could lead to unexpected behavior or errors on devices that do not support these flags, such as ZINC, MLX, or DS4, even though the code already checks for these devices and throws an error. The check should be consistent with the error throwing logic.

if (!o.mmproj.empty()) {
    if (device == "zinc" || device == "mlx" || device == "ds4") throw std::runtime_error("--mmproj works on the llama.cpp devices (vulkan, hrx, rocm)");
    // an image is decoded as one ubatch: models that attend to it bidirectionally (ZAYA1-VL,
    // Gemma 3) need all of it in one, and Qwen2.5-VL towers cap an image at 4096 tokens
    argv.insert(argv.end(), {"--mmproj", o.mmproj, "-b", "4096", "-ub", "4096"});
}
Incomplete documentation for ZAYA1-VL-8B

The documentation for ZAYA1-VL-8B in docs/vulkan.md mentions the model's vision-only LoRA and bidirectional attention but does not clearly explain how to handle the chat template issue that arises with older GGUFs. It should provide clear instructions on how to reconvert GGUFs or use --chat-template-file to ensure compatibility with the correct template.

GGUFs made before [llama.cpp #17](https://github.com/1bit-MONSTER/llama.cpp/pull/17) carry a chat
template that left each image inside the user turn instead of in front of it, as Zyphra's
template has it. llama-server marks images with `<__media_<id>__>`, an id random per server, and
that template knew only `<__media__>`. Reconvert those GGUFs, or serve them with
`--chat-template-file`. With the new template the image starts at the same prompt position as in
Zyphra's processor.

@bong-water-water-bong
bong-water-water-bong merged commit 8cbca67 into main Sep 26, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the zaya1-vl branch September 26, 2026 16:46
bong-water-water-bong added a commit that referenced this pull request Sep 26, 2026
…x (engine#115) (#151)

* third_party: bump llama.cpp pin for the HRX decode-split multipass fix (engine#115)

Was 1e775cd, needed to move to include both the ZAYA1-VL work already on
main (f30cc43, via #149) and the multi-pass decode-split reduction from
llama.cpp#13 that #121 was trying to add - #121 conflicted because its
target commit predated #149's pin advance. This points at the current
1bit/hrx-vulkan-patched tip (8dd75eb), a strict superset of both.

Verified: onebit builds; the exact >2048-token needle-retrieval repro from
engine#115 is correct and deterministic (4 back-to-back runs at 2124 and
3599 tokens, cache_prompt:false). The known residual fault (engine#123) is
real and reproduces more readily here than llama.cpp#13's own report
(0/5 there; 4/4 here across two separate runs, isolated to depths >2048,
Vulkan0 and short-context HRX0 unaffected) - noted on #123 for whoever
picks up the follow-up fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* registry: regenerate for llama.cpp 8dd75eb and llama.cpp-vulkan a29be7f

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: agent <agent@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant