…x (engine#115) (#151)
* third_party: bump llama.cpp pin for the HRX decode-split multipass fix (engine#115)
Was 1e775cd, needed to move to include both the ZAYA1-VL work already on
main (f30cc43, via #149) and the multi-pass decode-split reduction from
llama.cpp#13 that #121 was trying to add - #121 conflicted because its
target commit predated #149's pin advance. This points at the current
1bit/hrx-vulkan-patched tip (8dd75eb), a strict superset of both.
Verified: onebit builds; the exact >2048-token needle-retrieval repro from
engine#115 is correct and deterministic (4 back-to-back runs at 2124 and
3599 tokens, cache_prompt:false). The known residual fault (engine#123) is
real and reproduces more readily here than llama.cpp#13's own report
(0/5 there; 4/4 here across two separate runs, isolated to depths >2048,
Vulkan0 and short-context HRX0 unaffected) - noted on #123 for whoever
picks up the follow-up fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* registry: regenerate for llama.cpp 8dd75eb and llama.cpp-vulkan a29be7f
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: agent <agent@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Zyphra ZAYA1-VL-8B through
1bit serve --mmproj, plus the Zamba2-VL chat-template fix.Pins
third_party/llama.cpp83e1c41→f30cc43: adds llama.cpp#18, ZAYA1-VL.clip.vision.decode_non_causal;third_party/llama.cpp-vulkan2c9d3983→4a2c0656: adds llama.cpp#17. The Zamba2-VL template now recognises llama-server's random media markers and puts the image in front of the user turn, as Zyphra's does. The branch tip has since moved on to Tokenizers: Hugging Face tokenizers v0.23.2 behind our C ABI, exact on 18 models #19 (MoE streaming); this PR pins Linux: pin upstream v7.3-rc4, build it with amdxdna in-tree, keep it current #17's merge and leaves Tokenizers: Hugging Face tokenizers v0.23.2 behind our C ABI, exact on 18 models #19 out.serve: with
--mmproj, llama-server gets-b 4096 -ub 4096, so each image decodes as a single ubatch. ZAYA1-VL's bidirectional attention needs that.Checked on Strix Halo (F16 text, F16 mmproj) against Zyphra's own code (their transformers branch
zaya1-vl, FP32, CPU), teacher-forced on llama.cpp's greedy answers to three image questions (101 tokens). The prompts are identical: 208 tokens, with the image at the same position.Docs:
docs/vulkan.mdgets a ZAYA1-VL-8B section and a note on Zamba2-VL GGUFs made before #17.docs/serve.mdand the README are updated too.registry/architectures.jsonis regenerated and now mapsZaya1VLForConditionalGeneration.🤖 Generated with Claude Code