Repository navigation
zaya: ZAYA1-74B sliding-window attention, legacy checkpoints (ZAYA1-base, reasoning-base), --remote chat templates - #11
Conversation
ZAYA1-74B alternates hybrid_sliding and hybrid layers: sliding layers attend to the last sliding_window (4097) positions and use rope theta 1e4, the others 1e7. The converter now writes sliding_window, the per-layer pattern from layer_types and rope.freq_base_swa; the loader marks those layers SWA (LLAMA_SWA_TYPE_STANDARD: masked when q - k >= n_swa, as HF's create_sliding_window_causal_mask), so the hybrid memory becomes the iSWA one, and the graph takes each layer's rope base from get_rope_freq_base. Models without the keys (ZAYA1-8B) take the unchanged graph<false> path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review from the commit poll. The generic fixture writes
The rest looks right to me:
On "70/96": I measured the same with my harness on the Q4_K_M. The F16 GGUF scores 95/96, so the gap is quantisation. |
8c59770 to
8490f96
Compare
test-llama-archs writes the pattern as a scalar period and its save/reload drops it (llama_model_saver does not write the key). Reading a scalar through the per-layer array read stores the raw value in every layer (all layers slide), and without the key the load failed. Read it the way cohere2moe does: the scalar first as a period, then an array per layer (what our converter writes), and without the key the period is 2, which is ZAYA1-74B's pattern (even layers slide). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
8490f96 to
d8c84b2
Compare
|
Re-checked after d8c84b2: test-llama-archs passes on ubuntu x64/arm64 (the remaining Windows x64-openblas failure is test-thread-safety crashing at exit with 0xc0000374, which doesn't involve zaya). On strixhalo with this PR's zaya.cpp: |
4380daf
into
1bit/hrx-vulkan-patched
ZAYA1-74B-preview support plus the rest of the ZAYA1 family's checkpoints.
1. Sliding-window attention (ZAYA1-74B-preview). Its layers alternate between sliding-window attention (window 4097, rope theta 1e4) and full attention (theta 1e7). ZAYA1-8B has no sliding layers.
zaya.attention.sliding_window,zaya.attention.sliding_window_patternandzaya.rope.freq_base_swawhen the config has themgraph<iswa>and picks the rope base per layersliding_window_patternthe way cohere2moe does: a scalar first as a period, then an array per layer, and without the key a period of 2 (llama_model_saverdoes not write the key). The first version read a scalar through the array path, which failstest-llama-archs; fixed.2. Legacy (Megatron-style) checkpoints. ZAYA1-base, ZAYA1-reasoning-base and the
*-legacyrepos keep Zyphra's original layout: separate attention and MoE half-layers, per-layer config lists, and experts stored one by one. transformers 5 cannot load them. The converter normalizes the config and tensors to the transformers layout.3.
--remotefetcheschat_template.jinja. Without it, remote conversions had no chat template.Verified on Strix Halo:
test-llama-archs -a zaya: Vulkan and CPU OK, save/reload bit-exact (the synthetic model now takes the SWA path)--remote Zyphra/ZAYA1-8B-legacyconverts to a GGUF whose 1283 tensors are byte-identical to the one fromZyphra/ZAYA1-8B, with identical model metadata. Normalized configs equal the new ones for both the 8B and 74B pairs, including the 74B's window (legacy 4096 → 4097) and per-layer rope bases.Rebased onto #9.
🤖 Generated with Claude Code