Skip to content

zaya: ZAYA1-74B sliding-window attention, legacy checkpoints (ZAYA1-base, reasoning-base), --remote chat templates - #11

Merged
bong-water-water-bong merged 2 commits into
1bit/hrx-vulkan-patchedfrom
1bit/zaya-swa
Sep 26, 2026
Merged

bong-water-water-bong merged 2 commits into
1bit/hrx-vulkan-patchedfrom
1bit/zaya-swa

Conversation

@bong-water-water-bong

@bong-water-water-bong bong-water-water-bong commented Sep 26, 2026 •

Copy link
Copy Markdown

ZAYA1-74B-preview support plus the rest of the ZAYA1 family's checkpoints.

1. Sliding-window attention (ZAYA1-74B-preview). Its layers alternate between sliding-window attention (window 4097, rope theta 1e4) and full attention (theta 1e7). ZAYA1-8B has no sliding layers.

  • conversion/zaya.py writes zaya.attention.sliding_window, zaya.attention.sliding_window_pattern and zaya.rope.freq_base_swa when the config has them
  • src/models/zaya.cpp builds graph<iswa> and picks the rope base per layer
  • the loader reads sliding_window_pattern the way cohere2moe does: a scalar first as a period, then an array per layer, and without the key a period of 2 (llama_model_saver does not write the key). The first version read a scalar through the array path, which fails test-llama-archs; fixed.

2. Legacy (Megatron-style) checkpoints. ZAYA1-base, ZAYA1-reasoning-base and the *-legacy repos keep Zyphra's original layout: separate attention and MoE half-layers, per-layer config lists, and experts stored one by one. transformers 5 cannot load them. The converter normalizes the config and tensors to the transformers layout.

3. --remote fetches chat_template.jinja. Without it, remote conversions had no chat template.

Verified on Strix Halo:

  • test-llama-archs -a zaya: Vulkan and CPU OK, save/reload bit-exact (the synthetic model now takes the SWA path)
  • ZAYA1-8B: F16 reconverted after the rebase picks the same top token as transformers FP32 at 95/96 teacher-forced positions, the same as before
  • Legacy mapping: --remote Zyphra/ZAYA1-8B-legacy converts to a GGUF whose 1283 tensors are byte-identical to the one from Zyphra/ZAYA1-8B, with identical model metadata. Normalized configs equal the new ones for both the 8B and 74B pairs, including the 74B's window (legacy 4096 → 4097) and per-layer rope bases.
  • ZAYA1-74B-preview (HF revision bc8eb466): Q8_0 79.7 GB with the SWA keys; Q4_K_M 42.6 GiB, llama-bench Vulkan0 pp512 505.8 tok/s, tg128 35.4 tok/s; chat coherent

Rebased onto #9.

🤖 Generated with Claude Code

ZAYA1-74B alternates hybrid_sliding and hybrid layers: sliding layers attend
to the last sliding_window (4097) positions and use rope theta 1e4, the
others 1e7. The converter now writes sliding_window, the per-layer pattern
from layer_types and rope.freq_base_swa; the loader marks those layers SWA
(LLAMA_SWA_TYPE_STANDARD: masked when q - k >= n_swa, as HF's
create_sliding_window_causal_mask), so the hybrid memory becomes the iSWA
one, and the graph takes each layer's rope base from get_rope_freq_base.
Models without the keys (ZAYA1-8B) take the unchanged graph<false> path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong

Copy link
Copy Markdown
Author

Review from the commit poll. test-llama-archs fails with this change (fork CI: ubuntu x64/arm64, and the Windows jobs that run it):

26: | zaya | AMD EPYC 9V45 | MoE | E llama_model_load: error loading model: error loading model hyperparameters: key not found in model: zaya.attention.sliding_wi…

The generic fixture writes LLM_KV_ATTENTION_SLIDING_WINDOW for every arch, so the new get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, …, false) finds one and turns on SWA. A companion key that is read as required is then missing. Two ways to fix it:

  • key SWA off the pattern (or rope.freq_base_swa) being present, not the window alone. Only the converter writes the pattern, and only for ZAYA1-74B;
  • or set the ZAYA fixture in tests/test-llama-archs.cpp explicitly (arch == LLM_ARCH_ZAYA, like its SSM_CONV_KERNEL = 2), and cover the SWA path with a fixture that carries all three keys.

The rest looks right to me:

  • ZAYA1-8B is unchanged;
  • LLAMA_SWA_TYPE_STANDARD (masked when q − k ≥ n_swa) matches HF's sliding-window mask;
  • --rope-freq-base overrides still reach the non-SWA layers through get_rope_freq_base.

On "70/96": I measured the same with my harness on the Q4_K_M. The F16 GGUF scores 95/96, so the gap is quantisation.

test-llama-archs writes the pattern as a scalar period and its save/reload
drops it (llama_model_saver does not write the key). Reading a scalar through
the per-layer array read stores the raw value in every layer (all layers
slide), and without the key the load failed. Read it the way cohere2moe does:
the scalar first as a period, then an array per layer (what our converter
writes), and without the key the period is 2, which is ZAYA1-74B's pattern
(even layers slide).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong

Copy link
Copy Markdown
Author

Re-checked after d8c84b2: test-llama-archs passes on ubuntu x64/arm64 (the remaining Windows x64-openblas failure is test-thread-safety crashing at exit with 0xc0000374, which doesn't involve zaya). On strixhalo with this PR's zaya.cpp: test-llama-archs -a zaya OK on Vulkan and CPU, and ZAYA1-8B Q4_K_M perplexity on Vulkan is 21.5731, identical to before, so the 8B is unchanged. Merging.

@bong-water-water-bong
bong-water-water-bong merged commit 4380daf into 1bit/hrx-vulkan-patched Sep 26, 2026
7 of 26 checks passed
@bong-water-water-bong bong-water-water-bong changed the title zaya: sliding-window attention and a second rope base (ZAYA1-74B) zaya: ZAYA1-74B sliding-window attention, legacy checkpoints (ZAYA1-base, reasoning-base), --remote chat templates Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant