Repository navigation
convert: ZAYA legacy checkpoints (ZAYA1-base, reasoning-base); --remote fetches chat_template.jinja - #14
Merged
Conversation
ZAYA1-base, ZAYA1-reasoning-base and Zyphra's *-legacy repos keep the original checkpoint layout: attention and MoE are separate half-layers (zaya_layers), sizes live in per-layer lists (cca_num_q_heads, num_query_groups_list, ffn_hidden_size_list, zaya_mlp_expansion), experts are stored one by one, and each half-layer's res_scale merges the residual in front of it. transformers 5 cannot load these (its ZayaConfig rejects the config), so the converter normalizes them itself: - the config to the transformers keys (40 blocks from 80 half-layers, heads and groups from the per-layer lists, swa_layers to layer_types and a sliding_window one larger, as Zyphra's own conversion of ZAYA1-74B-preview) - the tensors to transformers names, stacking the experts, with the MoE half-layer's res_scale as the block's post-attention scale and the next attention half-layer's (the final one for the last block) as its post-MLP scale - the tokenizer is loaded from its files alone Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
--remote downloads the config and tokenizer files by pattern, and *.jinja was not among them, so models that ship their chat template as chat_template.jinja (ZAYA1-74B-preview, ZAYA1-base and others) were converted without tokenizer.chat_template. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
merged commit Sep 26, 2026
567d2dc
into
1bit/hrx-vulkan-patched
1 of 6 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #11: these two commits were pushed to
1bit/zaya-swaafter #11 merged.ZAYA legacy (Megatron-style) checkpoints. ZAYA1-base, ZAYA1-reasoning-base and Zyphra's
*-legacyrepos keep the original layout:zaya_layers)res_scalemerges the residual in front of ittransformers 5 cannot load them, because its
ZayaConfigrejects the config. The converter normalizes the config and tensors to the transformers layout, and loads the tokenizer from its files alone.--remotefetcheschat_template.jinja.*.jinjawas not in the download patterns, so remote conversions of repos that shipchat_template.jinjahad notokenizer.chat_template. ZAYA1-74B-preview's GGUF from #11's testing is one of them. Upstream has the same list.Verified on Strix Halo:
--remote Zyphra/ZAYA1-8B-legacyconverts to a GGUF whose 1283 tensors are byte-identical to the GGUF fromZyphra/ZAYA1-8B, with identical model metadataZyphra/ZAYA1-baseincludeschat_template.jinja.flake8(pyflakes + pycodestyle) clean;ty check conversion/zaya.pyreports only the unresolvedtorchimport of a torch-less environment🤖 Generated with Claude Code