Repository navigation
Conversation
Share target and draft checkpoint backing storage to avoid large deep copies when saving prompts to the in-memory cache. Also handle `std::bad_alloc` during complete cache entry construction, including `states.push_back()`. Assisted-by: Qwen3.8-27B
|
Hi @Hundsbuah, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
Hundsbuah
marked this pull request as ready for review
August 20, 2026 17:47
benjigill
added a commit
to benjigill/llama.cpp
that referenced
this pull request
Sep 23, 2026
…cal)
bench_qwen38.py cache suite, branch vs master build:
master branch
cold ttft s 19.72 19.89 +0.9%
warm prompt_n 18/16/17 18/16/17
multiturn prompt_n 26/22 26/22
multiturn ttft s 0.52/0.52 0.54/0.52
No regression; the suite does not reach the long-prefill lockup or the
--cache-ram limit these patches fix.
Assisted-by: Claude Opus 5.5
|
Related behavior we hit on a recurrent hybrid served over RPC: when the draft KV is full in the middle of a speculative pass, we purge idle slots and retry the draft pass once; if that still fails we clear the out-of-sync slots instead of failing the request. Under 3 concurrent sessions on a 64K unified pool this stopped speculative decoding from degrading the whole session on KV exhaustion. |
2 of 6 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Share target and draft checkpoint backing storage to avoid large deep copies when saving prompts to the in-memory cache.
Also handle
std::bad_allocduring complete cache entry construction, includingstates.push_back().Assisted-by: Qwen3.8-27B
Summary
This PR prevents excessive transient host-memory usage when
llama-serversaves prompts containing large context checkpoints into the in-memory
prompt cache.
The current implementation deep-copies the complete checkpoint payload
when a prompt is inserted into the cache. For long-context
hybrid/recurrent workloads, checkpoint state can become large enough for
this temporary duplication to cause significant memory pressure or
allocation failure even when the resulting cache entry itself satisfies
the configured
--cache-ramlimit.This PR:
reference-counted backing storage;
server_promptcopies;std::bad_allochandling to cover complete prompt-cache entryconstruction, including
states.push_back(...).Problem
server_promptstores checkpoints by value:Each checkpoint contains serialized state buffers such as:
When a prompt is saved to the in-memory prompt cache,
server_prompt_cache::alloc()constructs a new cache entry using theequivalent of:
states.push_back({ { prompt.tokens.clone(), prompt.checkpoints, }, ... });Copying
prompt.checkpointstherefore performs a deep copy of everydata_tgtanddata_dftbuffer.For long contexts or models requiring substantial recurrent checkpoint
state, this can make prompt-cache insertion temporarily require close to
an additional copy of the complete checkpoint payload.
Why
--cache-ramdoes not prevent the transient allocationThe cache admission logic includes checkpoint sizes when calculating the
logical size of a new cache entry:
This correctly limits the logical residency of the resulting cache entry.
However, during cache-entry construction the active prompt still owns its
original checkpoint buffers while the cache creates another copy.
The transient memory requirement is therefore approximately:
As a result, satisfying the logical cache-size limit does not guarantee
that enough host memory is available for the temporary deep copy required
to construct the cache entry.
Why checkpoints should not simply be removed
Clearing checkpoints before saving the prompt would avoid the duplication,
but would change behavior for hybrid/recurrent models.
These models may require captured checkpoints to restore recurrent state
when a later prompt shares only part of a cached prefix.
Discarding those checkpoints can prevent efficient partial-prefix reuse
and require substantially more prompt reprocessing.
This PR therefore preserves the checkpoints and removes only the
unnecessary duplication of their large serialized backing storage.
Implementation
Shared checkpoint payload
The large serialized target and draft checkpoint buffers now use
reference-counted backing storage.
Conceptually, instead of:
copies now behave as:
Only the lightweight ownership handle is copied.
Snapshot semantics are preserved
Checkpoint payloads are treated as immutable after capture.
When a checkpoint is updated, a new backing buffer is allocated and the
new serialized state is written into that buffer rather than modifying an
existing shared snapshot.
Conceptually:
This preserves independent checkpoint snapshot semantics while avoiding
unnecessary deep copies.
Target and draft state
Both large serialized state buffers use shared backing storage:
This preserves the behavior required for target and draft model state,
including speculative/MTP workloads.
The smaller speculative checkpoint state remains value-owned to keep the
change focused and avoid unnecessary modifications outside the primary
memory-pressure path.
OOM handling
The existing implementation catches
std::bad_allocaround:However, subsequent operations may allocate memory as well:
These operations were previously outside the protected section.
This PR extends the
tryblock to cover the complete cache-entryconstruction:
If any allocation involved in constructing the cache entry fails, the
existing cache-recovery path is used and the function returns without
leaving a partially constructed entry.
Temporary allocations are released automatically during stack unwinding.
Cache accounting
This PR intentionally does not change the existing logical cache-size
accounting.
Checkpoint payload sizes continue to count toward
--cache-rameven whentheir physical backing storage is shared.
This can make cache accounting conservative when multiple prompts refer
to the same backing allocation, potentially causing earlier eviction than
strict physical-byte accounting would require.
Keeping the existing accounting is intentional because it:
--cache-ramsemantics;Unique physical backing-store accounting or a global checkpoint pool can
be considered separately.
Expected memory behavior
Before this change, prompt-cache insertion may temporarily require:
With shared checkpoint backing storage, the same operation requires
approximately:
The large checkpoint payload is therefore retained only once instead of
being duplicated during cache insertion.
This does not make
--cache-rama total process-memory limit. Overallmemory usage still depends on active slots, model state, compute buffers,
allocator behavior and other process allocations.
Scope
The patch is deliberately limited to:
It does not change:
--cache-ramaccounting.Validation
The patch has been applied and checked against the target source tree.
The affected translation units compile successfully:
The recurrent-state rollback test build also proceeds through the affected
code without errors.
Additional runtime validation is appropriate for:
std::bad_allochandling;Motivation
The goal is to eliminate excessive transient checkpoint duplication
without sacrificing the state required for correct and efficient
hybrid/recurrent prompt reuse.
The core ownership model becomes:
This removes the problematic deep copy while retaining checkpoint
information needed for recurrent-state restoration and prompt reuse.