Skip to content

kv-cache: fix restoring mismatched KV cache rotation by saving exact rotation metadata - #28498

Merged
ggerganov merged 3 commits into
ggml-org:masterfrom
eapache:ehuus/save-restore-rotation
Oct 5, 2026
Merged

ggerganov merged 3 commits into
ggml-org:masterfrom
eapache:ehuus/save-restore-rotation

Conversation

@eapache

@eapache eapache commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Tighten up the save/restore of rotated KV caches. This is technically a bug already today, but a pretty edge-case-y one. However, it came up as a prereq to my work on lazy KV cache quantization (#28267, still WIP) which would hit this issue in a much more common path.

Additional information

Previously, llama_kv_cache only captured whether rotation was used, not the size of the rotation matrix, and no data on rotation was included in the save/restore path. If a rotated cache was restored into a process expecting an unrotated cache (or vice-versa), the data was silently misinterpreted and resulted in garbage. To reproduce, you can save a (default rotated) q8 K/V, then restart with LLAMA_ATTN_ROT_DISABLE=1 and restore that save.

Now, llama_kv_cache captures the size of the rotation matrix, and that metadata is also included in the save/restore path. Attempting to restore a cache with mismatching rotation is rejected with an error. This does involve bumping the version numbers and breaking compat with previously saved caches.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, Sol and Astra extracted this from the lazy-quant PR above

@github-actions github-actions Bot added the testing Everything test related label Sep 6, 2026
@eapache
eapache marked this pull request as ready for review September 6, 2026 15:10
@eapache
eapache force-pushed the ehuus/save-restore-rotation branch 2 times, most recently from e750ff6 to eb6c8b0 Compare September 19, 2026 14:08
@eapache eapache changed the title kv-cache: save exact KV rotation metadata, reject restoring mismatched rotation kv-cache: fix restoring mismatch KV cache rotation by saving exact rotation metadata Sep 19, 2026
@eapache eapache changed the title kv-cache: fix restoring mismatch KV cache rotation by saving exact rotation metadata kv-cache: fix restoring mismatched KV cache rotation by saving exact rotation metadata Sep 19, 2026
@eapache

eapache commented Sep 20, 2026

Copy link
Copy Markdown
Contributor Author

@ggerganov this is a relatively straightforward fix to the kv cache rotation logic which will unblock my work on lazy cache quanting (I am biased but of course I am fairly optimistic on how much it will help).

@eapache
eapache force-pushed the ehuus/save-restore-rotation branch from eb6c8b0 to 334a1d9 Compare October 3, 2026 13:54
@ggerganov ggerganov added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Oct 3, 2026
the test is now part of the save/load test matrix and runs against
every model under test, like the rest of the suite

it probes the KV cache type combinations supported by the model and
treats models that do not use attention rotation as passing vacuously

Assisted-by: pi:llama.cpp/Qwen3.8-27B
@ggerganov
ggerganov merged commit 2107910 into ggml-org:master Oct 5, 2026
12 checks passed
@eapache
eapache deleted the ehuus/save-restore-rotation branch October 5, 2026 11:39
Wizard815 pushed a commit to Wizard815/mx-llama.cpp-Rocm10 that referenced this pull request Oct 6, 2026
…rotation metadata (ggml-org#28498)

* kv-cache: save exact KV rotation metadata, reject restoring mismatched rotation

* tests : move the state rotation test to test-save-load-state

the test is now part of the save/load test matrix and runs against
every model under test, like the rest of the suite

it probes the KV cache type combinations supported by the model and
treats models that do not use attention rotation as passing vacuously

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* cont : skip unsupported KV caches

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
(cherry picked from commit 2107910)
@okigan

okigan commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

@ggerganov Thanks for merging I’d also like to get #26004 merged as it can make resume ~20–30× faster — e.g. Qwen3.8-27B goes from ~71s to ~3s, and Qwen3.8-Flash-Next from ~3.0s to ~0.1s.

It fixes a significant performance issue: after /slots save → restart → restore, hybrid/recurrent/SWA models can lose their context checkpoints and re-prefill the entire context. With #26004, the checkpoints survive the restore, so long contexts can resume by processing only the new tokens.

I think this complements the recent restore correctness fixes rather than overlapping with them. Would appreciate a look at getting it merged.

edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
…rotation metadata (ggml-org#28498)

* kv-cache: save exact KV rotation metadata, reject restoring mismatched rotation

* tests : move the state rotation test to test-save-load-state

the test is now part of the save/load test matrix and runs against
every model under test, like the rest of the suite

it probes the KV cache type combinations supported by the model and
treats models that do not use attention rotation as passing vacuously

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* cont : skip unsupported KV caches

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
(cherry picked from commit 2107910)
alainnothere added a commit to alainnothere/llama.cpp that referenced this pull request Oct 9, 2026
Two tailors in Lyon had been making the same coat for a year without meeting. One of them had sewn a hidden pocket for dice, because the customer gambled and liked to have them on him, and had sized it, lined it and tested it on his own hip. The other, who sold more coats, mailed over a pattern that spring with the same pocket drawn in, in the same place, in the same cloth, with a note that said the house of Vernet now considered the dice pocket standard. The first tailor laid the two patterns on the table and found that the seams crossed at fifteen points and that there was no honest way to have both pockets in one coat. He cut his out. He kept the stitch he'd invented for the lining, because the other pattern's lining tore on the Rhone wind and his didn't, and nobody from Vernet had ever stood on that quay.

The customer asked which pocket he'd got. The tailor said: theirs, with my thread, and the dice fit the same, and you will not find the seam because I moved it under the arm.

---

607 upstream commits (6d9c82e..06cad0b, b11480).

Dropped the fork's --spec-accept stochastic (f8a6903 / e9c788d /
60bc680) for upstream ggml-org#27694 --spec-draft-sampling probabilistic: same
rejection sampling, same three hook points, plus spec_retune() so the draft
samples at the target temperature. Running both side by side meant fifteen
conflicting hunks in common/speculative.cpp for one switch. test-spec-accept,
the stochastic/exact/fallback counters in server-common.h and
llama_sampler_grammar_is_active go with it. Server verify order is now synth,
replay, rejection, exact; the draft params block carries result_q/temp/seed
next to the tuner's n_draft_cur.

Kept: the keep vector on common_speculative_process, --spec-draft-auto,
--spec-draft-ctx-step, -lcd write-back, disk cache v4/v5 and checkpoint
spill, perf instrumentation, idempotent synchronize. server-http registers
upstream's callback lambda under path_prefix + path.

Upstream ggml-org#28498 writes n_rot_k/n_rot_v into the KV state blob, so disk-cache
files from the old binary fail with "incompatible key rotation" once, get
reprocessed and rewritten. Not a bug.

test-recurrent-state-rollback: upstream's dummy models (ggml-org#29133) are noisier
and Vulkan drifts ~1e-8 nmse between two contexts even without a restore, so
the composed/partial rollback tests bound nmse at 1e-4 and test_rollback at
1e-7 instead of asking for bit-exact logits. CPU stays bit-exact. 130/130
models pass on Vulkan.

Server pytests: 24 pass with the uncached upstream presets (tinylaya,
tinyopenjev) skipped; they cannot be fetched behind the proxy.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants