Repository navigation
ZAYA1 (Zyphra) on the GPU: pin llama.cpp 5556bf2, serve routes fork-only architectures - #107
Conversation
|
Docs7 for 1bit-monster/engine
Commit |
… llama.cpp `1bit serve --device vulkan` reads the GGUF's general.architecture (app/gguf_meta.h) and runs an architecture only our llama.cpp implements (zaya) on the HRX build's Vulkan0, even when the upstream Vulkan build is present. registry_build counts such architectures for vulkan the same way, and check_models.tsv checks ZAYA1-8B Q4_K_M on vulkan. docs/vulkan.md: converting and serving ZAYA1, with the checks against transformers and the measured speeds on Strix Halo. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…egistry Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ama.cpp #2, #5, #6) The pin adds the zaya architecture and its converter (#2, #5) on top of #104's #95 fixes, and carries #106's scalar-scatter SET_ROWS onto the tracked branch (#6): #106 pinned c358334, which is on no branch of the fork. registry/architectures.json, regenerated: ZayaForCausalLM maps to hrx and vulkan (266 architectures, all mapped). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…1-8B included) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
d1dc85c to
5638df3
Compare
PR Reviewer Guide 🔍(Review updated until commit 6a4d800)Here are some key observations to aid the review process:
|
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Persistent review updated to latest commit 6a4d800 |
ZAYA1-8B (Zyphra) joins Qwen in the showcase, served by
1bit serveon the Strix Halo GPU.What changes
third_party/llama.cpp→5556bf2. That's 1bit-MONSTER/llama.cpp Tokenizer: byte-level BPE from GGUF, exact against HF tokenizers #2 and Remove the CPU reference and tokenizer; port the existing engine instead #5 (ZAYA) on top of HRX: pin llama.cpp 96f6b89, fixing #95 (GLM-4.7-Flash, Qwen3-Coder at full context) #104's HRX: chat on Qwen3-Coder-30B-A3B and GLM-4.7-Flash fails with 500 "Compute error." #95 fixes, plus Steps 1-2: embedded Lemonade, and HRX + Vulkan in one llama.cpp build #6, which carries hrx: scalar-scatter SET_ROWS for the non-FA path (-fa off) (#95) #106's SET_ROWS commit onto the tracked branch. hrx: scalar-scatter SET_ROWS for the non-FA path (-fa off) (#95) #106 pinnedc358334, which is on no fork branch. It adds thezayaarchitecture, a converter for the transformers checkpoint, and support on Vulkan, HRX and ROCm (HIP).1bit serve --device vulkanreads the GGUF'sgeneral.architecture(app/gguf_meta.h, a header-only reader). An architecture that only our llama.cpp implements (zaya) runs on the HRX build'sVulkan0, even when the upstream Vulkan build is present.--device hrxruns it onHRX0.tools/registry_build.pycounts such architectures for vulkan the same way.ZayaForCausalLMis now mapped (hrx, vulkan; 266 architectures).check_models.tsvchecks ZAYA1-8B on vulkan and hrx.docs/vulkan.md(converting, serving, checks, speeds by device), README status, PORTING (row 5z),docs/hrx.md.Checks (strixhalo, Radeon 8060S)
Unit test:
gguf_meta(new, hosted CI), plussmoke_serveandthink_split.Routing: with the upstream Vulkan server set to
/bin/false, ZAYA still passesserve_e2e(routed to our build) and Qwen3-0.6B fails (routed to upstream). The routing works both ways.registry_check, vulkan and hrx rows: 16/16 PASS. This ran at engineb9812e0with thee2e6ca2HRX build; #6 changes only the-fa offpath, and ZAYA and Qwen3-0.6B pass-fa offon HRX0 with it. That's Qwen3-0.6B, Qwen2.5-7B, Qwen3-Coder-30B-A3B, Qwen3.6-35B-A3B, GLM-4.7-Flash, MiniCPM4-8B, MiniCPM5-1B and ZAYA1-8B, each on both devices.Accuracy against transformers FP32, teacher-forced (96 positions): the ZAYA1 F16 GGUF agrees on the top-1 token at 95/96. The miss is a 0.07-nat tie; transformers' own BF16 run matches FP32 at 91/96.
Speed, ZAYA1-8B Q4_K_M:
The engine's
--device rocmserver is built from ROCmFPX's tree, which has no ZAYA. The ROCm numbers above come from our llama.cpp built withGGML_HIP=ON.🤖 Generated with Claude Code