Skip to content

server : use the serialized position range for context checkpoints - #6

Merged
Jackson57279 merged 1 commit into
masterfrom
devin/swa-checkpoint-pos-range
Oct 8, 2026
Merged

Jackson57279 merged 1 commit into
masterfrom
devin/swa-checkpoint-pos-range

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Overview

Fix the context-checkpoint position range for SWA models in llama-server.

The server stamps each saved checkpoint with the [pos_min, pos_max] range of the KV cache memory. For sliding-window attention (SWA) models, state_write only serializes the cells inside the window, so the saved state covers a narrower range than the memory claims. The overclaimed range lets the server reuse a checkpoint for tokens that were never actually saved, producing corrupted responses (see the TODO marked [TAG_CHECKPOINTS_FIX_POS_MIN] in server-context.cpp and the upstream discussion in ggml-org#24411).

This adds llama_memory_state_pos_min() / llama_memory_state_pos_max() next to the existing llama_memory_seq_pos_min/max() wrappers. They return the position range that a state save actually covers for a given llama_state_seq_flags:

  • llama_kv_cache: for n_swa > 0, scans the sequence's cells and skips masked cells - mirrors the state_write filter exactly (works for all swa_type values)
  • llama_kv_cache_iswa / llama_kv_cache_dsa_iswa: PARTIAL_ONLY saves only the SWA tier; full saves cover the intersection of both tiers
  • llama_kv_cache_msa / llama_kv_cache_dsa: both sub-caches are always serialized, so coverage is the intersection
  • llama_memory_hybrid / llama_memory_hybrid_iswa (and hybrid_idx via inheritance): coverage is the intersection of both memories; PARTIAL_ONLY hybrid saves only contain the recurrent part
  • every other memory type keeps the memory range (default implementation)

create_checkpoint now stamps checkpoints with the serialized range instead of the raw memory range.

Verified locally with tests/test-state-range.cpp (new unit test, gated on a GGUF path):

  • gpt-oss-20b (n_swa = 128): memory reports [0, 255]; the state save covers [128, 255] — the previous code would have claimed the full [0, 255]
  • Qwen3-0.6B (n_swa = 0): state range equals the memory range [0, 511]
  • LFM2-1.2B (hybrid): state range equals the memory range

Additional information

Upstream reference: ggml-org#24411 (comment) - maintainer asked for a fix that reflects the correct position range plus unit tests and a more robust way to determine the saved range (this generalizes the pos_max - n_swa sketch to every swa_type and every composite memory type).

The pos_max > pos_next check in the checkpoint matcher is kept as a guard against stale checkpoints, but the range stamps are now correct on their own.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - implemented by Devin (Cognition AI) as part of an agentic research-and-implement session; the change, test, and verification were produced and checked by the agent for the fork

Link to Devin session: https://app.devin.ai/sessions/593da2a391e54fa691b41b61a25e3179
Open in Devin Desktop: https://app.devin.ai/desktop/session/593da2a391e54fa691b41b61a25e3179?variant=devin
Requested by: @Jackson57279


Summary by cubic

Fixes the context-checkpoint position range in llama-server so SWA checkpoints no longer claim coverage of tokens that were never serialized, preventing corrupted responses when reusing them.

Changes

  • Adds llama_memory_state_pos_min/max to report the actual position range covered by a state save (default: memory range, narrowed for SWA masking, partial saves, and joint caches).
  • create_checkpoint now stamps checkpoints with the serialized range instead of the raw memory range.
  • Adds tests/test-state-range.cpp to verify the saved range for SWA and non-SWA models.

Written for commit 3c74a8f. Summary will update on new commits.

View guided diff Turn on auto-fix

SWA caches only serialize the cells inside the sliding window, so a saved
sequence state covers a narrower position range than the memory reports.
Stamp checkpoints with the serialized range instead of the full memory
range, so checkpoint reuse does not overclaim coverage and reuse closer
checkpoints correctly.

Add llama_memory_state_pos_min/max to query the saved coverage for a
given set of save flags. The default is the memory range; caches that
serialize less data (SWA masking, PARTIAL_ONLY saves, joint caches)
narrow it accordingly.

Refs: ggml-org#24411 (comment)

Assisted-by: Devin
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@Jackson57279
Jackson57279 merged commit 0edfd2b into master Oct 8, 2026
11 of 26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant