Repository navigation
kv-cache: fix restoring mismatched KV cache rotation by saving exact rotation metadata - #28498
Conversation
e750ff6 to
eb6c8b0
Compare
|
@ggerganov this is a relatively straightforward fix to the kv cache rotation logic which will unblock my work on lazy cache quanting (I am biased but of course I am fairly optimistic on how much it will help). |
eb6c8b0 to
334a1d9
Compare
the test is now part of the save/load test matrix and runs against every model under test, like the rest of the suite it probes the KV cache type combinations supported by the model and treats models that do not use attention rotation as passing vacuously Assisted-by: pi:llama.cpp/Qwen3.8-27B
…rotation metadata (ggml-org#28498) * kv-cache: save exact KV rotation metadata, reject restoring mismatched rotation * tests : move the state rotation test to test-save-load-state the test is now part of the save/load test matrix and runs against every model under test, like the rest of the suite it probes the KV cache type combinations supported by the model and treats models that do not use attention rotation as passing vacuously Assisted-by: pi:llama.cpp/Qwen3.8-27B * cont : skip unsupported KV caches --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> (cherry picked from commit 2107910)
|
@ggerganov Thanks for merging I’d also like to get #26004 merged as it can make resume ~20–30× faster — e.g. Qwen3.8-27B goes from ~71s to ~3s, and Qwen3.8-Flash-Next from ~3.0s to ~0.1s. It fixes a significant performance issue: after /slots save → restart → restore, hybrid/recurrent/SWA models can lose their context checkpoints and re-prefill the entire context. With #26004, the checkpoints survive the restore, so long contexts can resume by processing only the new tokens. I think this complements the recent restore correctness fixes rather than overlapping with them. Would appreciate a look at getting it merged. |
…rotation metadata (ggml-org#28498) * kv-cache: save exact KV rotation metadata, reject restoring mismatched rotation * tests : move the state rotation test to test-save-load-state the test is now part of the save/load test matrix and runs against every model under test, like the rest of the suite it probes the KV cache type combinations supported by the model and treats models that do not use attention rotation as passing vacuously Assisted-by: pi:llama.cpp/Qwen3.8-27B * cont : skip unsupported KV caches --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> (cherry picked from commit 2107910)
Two tailors in Lyon had been making the same coat for a year without meeting. One of them had sewn a hidden pocket for dice, because the customer gambled and liked to have them on him, and had sized it, lined it and tested it on his own hip. The other, who sold more coats, mailed over a pattern that spring with the same pocket drawn in, in the same place, in the same cloth, with a note that said the house of Vernet now considered the dice pocket standard. The first tailor laid the two patterns on the table and found that the seams crossed at fifteen points and that there was no honest way to have both pockets in one coat. He cut his out. He kept the stitch he'd invented for the lining, because the other pattern's lining tore on the Rhone wind and his didn't, and nobody from Vernet had ever stood on that quay. The customer asked which pocket he'd got. The tailor said: theirs, with my thread, and the dice fit the same, and you will not find the seam because I moved it under the arm. --- 607 upstream commits (6d9c82e..06cad0b, b11480). Dropped the fork's --spec-accept stochastic (f8a6903 / e9c788d / 60bc680) for upstream ggml-org#27694 --spec-draft-sampling probabilistic: same rejection sampling, same three hook points, plus spec_retune() so the draft samples at the target temperature. Running both side by side meant fifteen conflicting hunks in common/speculative.cpp for one switch. test-spec-accept, the stochastic/exact/fallback counters in server-common.h and llama_sampler_grammar_is_active go with it. Server verify order is now synth, replay, rejection, exact; the draft params block carries result_q/temp/seed next to the tuner's n_draft_cur. Kept: the keep vector on common_speculative_process, --spec-draft-auto, --spec-draft-ctx-step, -lcd write-back, disk cache v4/v5 and checkpoint spill, perf instrumentation, idempotent synchronize. server-http registers upstream's callback lambda under path_prefix + path. Upstream ggml-org#28498 writes n_rot_k/n_rot_v into the KV state blob, so disk-cache files from the old binary fail with "incompatible key rotation" once, get reprocessed and rewritten. Not a bug. test-recurrent-state-rollback: upstream's dummy models (ggml-org#29133) are noisier and Vulkan drifts ~1e-8 nmse between two contexts even without a restore, so the composed/partial rollback tests bound nmse at 1e-4 and test_rollback at 1e-7 instead of asking for bit-exact logits. CPU stays bit-exact. 130/130 models pass on Vulkan. Server pytests: 24 pass with the uncached upstream presets (tinylaya, tinyopenjev) skipped; they cannot be fetched behind the proxy.
Overview
Tighten up the save/restore of rotated KV caches. This is technically a bug already today, but a pretty edge-case-y one. However, it came up as a prereq to my work on lazy KV cache quantization (#28267, still WIP) which would hit this issue in a much more common path.
Additional information
Previously,
llama_kv_cacheonly captured whether rotation was used, not the size of the rotation matrix, and no data on rotation was included in the save/restore path. If a rotated cache was restored into a process expecting an unrotated cache (or vice-versa), the data was silently misinterpreted and resulted in garbage. To reproduce, you can save a (default rotated) q8 K/V, then restart withLLAMA_ATTN_ROT_DISABLE=1and restore that save.Now,
llama_kv_cachecaptures the size of the rotation matrix, and that metadata is also included in the save/restore path. Attempting to restore a cache with mismatching rotation is rejected with an error. This does involve bumping the version numbers and breaking compat with previously saved caches.Requirements