Repository navigation
server : use the serialized position range for context checkpoints - #6
Merged
Merged
Conversation
SWA caches only serialize the cells inside the sliding window, so a saved sequence state covers a narrower position range than the memory reports. Stamp checkpoints with the serialized range instead of the full memory range, so checkpoint reuse does not overclaim coverage and reuse closer checkpoints correctly. Add llama_memory_state_pos_min/max to query the saved coverage for a given set of save flags. The default is the memory range; caches that serialize less data (SWA masking, PARTIAL_ONLY saves, joint caches) narrow it accordingly. Refs: ggml-org#24411 (comment) Assisted-by: Devin Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Author
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Fix the context-checkpoint position range for SWA models in
llama-server.The server stamps each saved checkpoint with the
[pos_min, pos_max]range of the KV cache memory. For sliding-window attention (SWA) models,state_writeonly serializes the cells inside the window, so the saved state covers a narrower range than the memory claims. The overclaimed range lets the server reuse a checkpoint for tokens that were never actually saved, producing corrupted responses (see the TODO marked[TAG_CHECKPOINTS_FIX_POS_MIN]inserver-context.cppand the upstream discussion in ggml-org#24411).This adds
llama_memory_state_pos_min()/llama_memory_state_pos_max()next to the existingllama_memory_seq_pos_min/max()wrappers. They return the position range that a state save actually covers for a givenllama_state_seq_flags:llama_kv_cache: forn_swa > 0, scans the sequence's cells and skips masked cells - mirrors thestate_writefilter exactly (works for allswa_typevalues)llama_kv_cache_iswa/llama_kv_cache_dsa_iswa:PARTIAL_ONLYsaves only the SWA tier; full saves cover the intersection of both tiersllama_kv_cache_msa/llama_kv_cache_dsa: both sub-caches are always serialized, so coverage is the intersectionllama_memory_hybrid/llama_memory_hybrid_iswa(andhybrid_idxvia inheritance): coverage is the intersection of both memories;PARTIAL_ONLYhybrid saves only contain the recurrent partcreate_checkpointnow stamps checkpoints with the serialized range instead of the raw memory range.Verified locally with
tests/test-state-range.cpp(new unit test, gated on a GGUF path):gpt-oss-20b(n_swa = 128): memory reports[0, 255]; the state save covers[128, 255]— the previous code would have claimed the full[0, 255]Qwen3-0.6B(n_swa = 0): state range equals the memory range[0, 511]LFM2-1.2B(hybrid): state range equals the memory rangeAdditional information
Upstream reference: ggml-org#24411 (comment) - maintainer asked for a fix that reflects the correct position range plus unit tests and a more robust way to determine the saved range (this generalizes the
pos_max - n_swasketch to everyswa_typeand every composite memory type).The
pos_max > pos_nextcheck in the checkpoint matcher is kept as a guard against stale checkpoints, but the range stamps are now correct on their own.Requirements
Link to Devin session: https://app.devin.ai/sessions/593da2a391e54fa691b41b61a25e3179
Open in Devin Desktop: https://app.devin.ai/desktop/session/593da2a391e54fa691b41b61a25e3179?variant=devin
Requested by: @Jackson57279
Summary by cubic
Fixes the context-checkpoint position range in
llama-serverso SWA checkpoints no longer claim coverage of tokens that were never serialized, preventing corrupted responses when reusing them.Changes
llama_memory_state_pos_min/maxto report the actual position range covered by a state save (default: memory range, narrowed for SWA masking, partial saves, and joint caches).create_checkpointnow stamps checkpoints with the serialized range instead of the raw memory range.tests/test-state-range.cppto verify the saved range for SWA and non-SWA models.Written for commit 3c74a8f. Summary will update on new commits.