Skip to content

convert: ZAYA legacy checkpoints (ZAYA1-base, reasoning-base); --remote fetches chat_template.jinja - #14

Merged
bong-water-water-bong merged 2 commits into
1bit/hrx-vulkan-patchedfrom
1bit/zaya-legacy
Sep 26, 2026
Merged

bong-water-water-bong merged 2 commits into
1bit/hrx-vulkan-patchedfrom
1bit/zaya-legacy

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

Follow-up to #11: these two commits were pushed to 1bit/zaya-swa after #11 merged.

ZAYA legacy (Megatron-style) checkpoints. ZAYA1-base, ZAYA1-reasoning-base and Zyphra's *-legacy repos keep the original layout:

  • attention and MoE are separate half-layers (zaya_layers)
  • sizes live in per-layer lists
  • experts are stored one by one
  • each half-layer's res_scale merges the residual in front of it

transformers 5 cannot load them, because its ZayaConfig rejects the config. The converter normalizes the config and tensors to the transformers layout, and loads the tokenizer from its files alone.

--remote fetches chat_template.jinja. *.jinja was not in the download patterns, so remote conversions of repos that ship chat_template.jinja had no tokenizer.chat_template. ZAYA1-74B-preview's GGUF from #11's testing is one of them. Upstream has the same list.

Verified on Strix Halo:

  • --remote Zyphra/ZAYA1-8B-legacy converts to a GGUF whose 1283 tensors are byte-identical to the GGUF from Zyphra/ZAYA1-8B, with identical model metadata
  • Normalized configs equal the new-format ones for the 8B and 74B pairs, including the 74B's sliding window (legacy 4096 → 4097) and per-layer rope bases
  • With the fix, the remote snapshot of Zyphra/ZAYA1-base includes chat_template.jinja
  • flake8 settings from .flake8 (pyflakes + pycodestyle) clean; ty check conversion/zaya.py reports only the unresolved torch import of a torch-less environment

🤖 Generated with Claude Code

bong-water-water-bong and others added 2 commits September 26, 2026 07:34
ZAYA1-base, ZAYA1-reasoning-base and Zyphra's *-legacy repos keep the
original checkpoint layout: attention and MoE are separate half-layers
(zaya_layers), sizes live in per-layer lists (cca_num_q_heads,
num_query_groups_list, ffn_hidden_size_list, zaya_mlp_expansion), experts are
stored one by one, and each half-layer's res_scale merges the residual in
front of it. transformers 5 cannot load these (its ZayaConfig rejects the
config), so the converter normalizes them itself:

- the config to the transformers keys (40 blocks from 80 half-layers, heads
  and groups from the per-layer lists, swa_layers to layer_types and a
  sliding_window one larger, as Zyphra's own conversion of ZAYA1-74B-preview)
- the tensors to transformers names, stacking the experts, with the MoE
  half-layer's res_scale as the block's post-attention scale and the next
  attention half-layer's (the final one for the last block) as its post-MLP
  scale
- the tokenizer is loaded from its files alone

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
--remote downloads the config and tokenizer files by pattern, and *.jinja was
not among them, so models that ship their chat template as
chat_template.jinja (ZAYA1-74B-preview, ZAYA1-base and others) were converted
without tokenizer.chat_template.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit 567d2dc into 1bit/hrx-vulkan-patched Sep 26, 2026
1 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant