Skip to content

Model-Saver: Write the SWA pattern, 15 more architectures roundtrip - #29042

Merged
ServeurpersoCom merged 2 commits into
ggml-org:masterfrom
ServeurpersoCom:model-saver-swa-pattern
Sep 18, 2026
Merged

ServeurpersoCom merged 2 commits into
ggml-org:masterfrom
ServeurpersoCom:model-saver-swa-pattern

Conversation

@ServeurpersoCom

@ServeurpersoCom ServeurpersoCom commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Improves CI coverage: the model saver now handles the SWA pattern, so 15 more architectures get the dummy models and the roundtrip, save/load state and fusion tests that come with them.

The diff touches 21 files but 18 of them are the same one-line substitution in a loader; the whole change to read is load_swa_pattern() and the saver hunk. The CI runs the roundtrip on every architecture, and a full local run of test-llama-archs on CPU and CUDA gives 254 roundtrips green with no failure.

Additional information

The missing sliding_window_pattern key is the usual reason an architecture ends up excluded from the saver (#27000 review), it got the line dropped from #25505 in review, and #24423 and #25731 are about to grow the exclusion list for the same reason.

The saver writes the pattern as one flag per layer and never collapses it to a scalar, since the loaders read a scalar as a period. The loaders that only understood a period now go through a shared load_swa_pattern() helper that accepts either form, which also fixes them silently dropping the arrays the olmo2, gemma3n and exaone4 converters already write (the published GGUFs match the default period, so no output changes). Three loaders lose their duplicated scalar-then-array block.

Enables plamo3, gemma3, cohere2, cohere2moe, olmo2, exaone-moe, afmoe, mimo2, spark2_5, muse-glimmer, mellum, laguna, granite_swa, dots3note and maple, all bit-exact on the test-llama-archs roundtrip. The 5 still excluded (gemma3n, bitnet, t5, apertus, step35) fail for unrelated reasons, each a small follow-up.

Requirements

Add llama_model_base::load_swa_pattern(), which reads
sliding_window_pattern either as one flag per layer or as a period
expanded by set_swa_pattern(), and use it in every loader that reads
the key as a period.

These loaders silently ignored an array and applied their default
period, although the converters of olmo2, gemma3n and exaone4 write
arrays. The published GGUFs match the defaults, so their outputs do
not change. The loaders that already accepted both forms lose their
duplicated scalar-then-array block, and use their declared default
period when the key is absent.
Write sliding_window_pattern as one flag per layer, nextn layers
included, for every model using SWA. The array is never collapsed to
a scalar, since the loaders read a scalar as a period.

Also write the MLA key/value lengths and KV LoRA rank of the SWA
layers, required by dots3note.

This enables the saver for plamo3, gemma3, cohere2, cohere2moe,
olmo2, exaone-moe, afmoe, mimo2, spark2_5, muse-glimmer, mellum,
laguna, granite_swa, dots3note and maple, all passing the bit-exact
roundtrip of test-llama-archs.
@github-actions github-actions Bot added the model Model specific label Sep 17, 2026
This was referenced Sep 17, 2026
@ServeurpersoCom

Copy link
Copy Markdown
Contributor Author

I already have the follow-up for four of the five remaining architectures ready, small saver omissions only; step35 needs a bit more thought, so it gets its own PR. That splits the review into three small pieces instead of one big one.

ServeurpersoCom added a commit to ServeurpersoCom/llama.cpp that referenced this pull request Sep 17, 2026
Write decoder_start_token_id as uint32 like the converter and the
loader, write the xielu arrays one value per layer, save tok_embd
when the model has no output tensor, and save the altup and per
layer projection tensors of gemma3n.

Follow-up of ggml-org#29042, step35 is now the last one left to enable.
ServeurpersoCom added a commit to ServeurpersoCom/llama.cpp that referenced this pull request Sep 18, 2026
Write decoder_start_token_id as uint32 like the converter and the
loader, write the xielu arrays one value per layer, save tok_embd
when the model has no output tensor, and save the altup and per
layer projection tensors of gemma3n.

Follow-up of ggml-org#29042, step35 is now the last one left to enable.
ServeurpersoCom added a commit to ServeurpersoCom/llama.cpp that referenced this pull request Sep 18, 2026
Write decoder_start_token_id as uint32 like the converter and the
loader, write the xielu arrays one value per layer, save tok_embd
when the model has no output tensor, and save the altup and per
layer projection tensors of gemma3n.

Follow-up of ggml-org#29042, step35 is now the last one left to enable.
ServeurpersoCom added a commit to ServeurpersoCom/llama.cpp that referenced this pull request Sep 18, 2026
Write decoder_start_token_id as uint32 like the converter and the
loader, write the xielu arrays one value per layer, save tok_embd
when the model has no output tensor, and save the altup and per
layer projection tensors of gemma3n.

Follow-up of ggml-org#29042, step35 is now the last one left to enable.
@ServeurpersoCom
ServeurpersoCom merged commit 542348a into ggml-org:master Sep 18, 2026
26 of 27 checks passed
fencerJP pushed a commit to fencerJP/llama-apu that referenced this pull request Sep 19, 2026
…gml-org#29042)

* llama: read the SWA pattern as a period or a per-layer array

Add llama_model_base::load_swa_pattern(), which reads
sliding_window_pattern either as one flag per layer or as a period
expanded by set_swa_pattern(), and use it in every loader that reads
the key as a period.

These loaders silently ignored an array and applied their default
period, although the converters of olmo2, gemma3n and exaone4 write
arrays. The published GGUFs match the defaults, so their outputs do
not change. The loaders that already accepted both forms lose their
duplicated scalar-then-array block, and use their declared default
period when the key is absent.

* model-saver: write the SWA pattern and the MLA SWA geometry

Write sliding_window_pattern as one flag per layer, nextn layers
included, for every model using SWA. The array is never collapsed to
a scalar, since the loaders read a scalar as a period.

Also write the MLA key/value lengths and KV LoRA rank of the SWA
layers, required by dots3note.

This enables the saver for plamo3, gemma3, cohere2, cohere2moe,
olmo2, exaone-moe, afmoe, mimo2, spark2_5, muse-glimmer, mellum,
laguna, granite_swa, dots3note and maple, all passing the bit-exact
roundtrip of test-llama-archs.
troycorbinz added a commit to eugene-plexus/library that referenced this pull request Oct 3, 2026
Upstream drift audit, 2026-10-03 (specs
docs/maintenance/upstream-drift-2026-10-03.md, llama.cpp minor items).
Each made the KV estimate too large: safe, but a model that fits could
be told it does not. The semantics are llama.cpp's own at tag b11375,
read in its source, and each is applied only where that source applies
it; any other architecture keeps the plain reading.

- A scalar attention.sliding_window_pattern is a period
  (load_swa_pattern -> set_swa_pattern, ggml-org/llama.cpp#29042): every
  n-th layer full, 0 = all sliding, 1 = none, prediction layers full,
  with each architecture's own dense_first. Thirteen architectures
  whose loaders read it that way with the file's own window; llama4,
  smallthinker, exaone4 and the two encoders are left out and say why.
- attention.shared_kv_layers: gemma4 caches only its first
  block_count - shared layers (has_kv, the reuse callback); gemma3n
  does the same with 20 written into its loader and never reads the
  key. attention_layers reports the layers that hold a cache.
- MLA: with key_length_mla and value_length_mla set (is_mla), the cache
  is K alone (has_v = !is_mla). The converter already writes key_length
  as kv_lora_rank + rope (576 on DeepSeek-V3) and head_count_kv 1, so
  K + V counted 1,088 elements a token a layer where llama.cpp stores
  576.

All three ride on the per-layer form: a LayerKV with 0 heads holds no
cache of its own, one with value_length 0 caches K alone, so the
starter file's layerRuns can carry them unchanged.

Checked against two real headers (anonymous ranged reads): the shipped
8B starter, gemma-4-E4B, declares shared_kv_layers 18, so 24 of 42
layers hold a cache (0.52 GiB at 32k where the plain reading said
0.91); the 14B, gemma-4-12b, declares 0 and reads exactly as its
starter entry already does.

Not done: a listed architecture's DEFAULT period when the file has no
pattern key (gemma3 6, gpt-oss 2, which their converters do not write);
the SWA cache's n_ubatch cells (llama.cpp sizes a sliding layer at
window + n_ubatch, padded to 256; the per-layer form counts the window).

Tests: tests/test_kv_conventions.py, 20 cases on GGUFs written with each
converter's keys; 11 red before the fix (per-layer form not built,
shared layers counted, MLA at 4,349,493,248 bytes for 2,302,672,896),
all green after; a sabotage pass of 15 breakages caught all 15.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Troy Corbin <troy.corbin@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model Model specific

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants