Skip to content

memory : remove KV cache size padding - #16812

Merged
ggerganov merged 4 commits into
masterfrom
gg/context-no-pad
Oct 28, 2025
Merged

ggerganov merged 4 commits into
masterfrom
gg/context-no-pad

Conversation

@ggerganov

@ggerganov ggerganov commented Oct 28, 2025 •

Copy link
Copy Markdown
Member

ref #16736 (comment)

Simplify the logic during memory module creation. Before, for KV caches, we used to pad the buffer size up to 256 cells since flash attention implementations did not support arbitrary K, V sizes. After the improvements in #16148 and related, we no longer need to explicitly do this padding.

For now, keeping support for llama_kv_cache size padding via the constructor's n_pad argument, although it is not currently used anymore.

Note that we continue to pad n_kv - this is the tensor shape for the K and V tensors for each graph:

https://github.com/ggml-org/llama.cpp/blob/1473d59a7bb0716aee4ab104ab09ecf7d978bd40/src/llama-kv-cache.cpp#L957-L972

We need to do this in order to reuse most of the compute graphs during the text generation phase. Additionally, this helps the performance with some of the backends.

Also, llama_model::create_memory() no longer mutates cparams.

Next, will rebase #16736 on top of this change and finish it.

@ggerganov
ggerganov requested a review from CISC as a code owner October 28, 2025 07:32
@JohannesGaessler

Copy link
Copy Markdown
Contributor

FYI in the CUDA backend, while non-padded inputs will produce correct results, the performance will be worse. One reason is that the vector and mma kernels don't have support for it (in principle also the wmma kernel but I want to remove that one soon). Another reason is that like with e.g. ALiBi I've added support for a non-padded KV cache only to the template specialization that doesn't have GQA-specific optimizations (to keep compilation times low and to mitigate the few % performance penalty from the OOB checks).

@ggerganov

Copy link
Copy Markdown
Member Author

What is the optimal padding for CUDA - 128 or 256?

@JohannesGaessler

Copy link
Copy Markdown
Contributor

As of right now the required padding is still 256. The vector and tile kernels I've already adapted to be able to only need a padding of 128 (without OOB checks). The mma kernel, wmma kernel, and some utility kernels still need a padding of 256. My plan is to implement support for Volta, AMD WMMA instructions, and AMD MFMA instructions directly in the mma kernel, then I can just remove the wmma kernel without having to make any changes to it. The hardware I need for development are a V100 (which should arrive today), a MI100 (which arrived yesterday), and an RDNA3+ GPU (already in hand). The mma kernel and the utility kernels can then be made to work with a padding of 128 without issue.

Caveat: the MI100 doesn't seem to work with the motherboard that I intended to use with it so I may need to shuffle around my hardware a bit.

@ggerganov

Copy link
Copy Markdown
Member Author

@JohannesGaessler The updated PR now continues to pad the n_kv to 256. This should basically keep the performance the same as on master. The only change is that the total size of the KV cache is no longer padded.

For example, we can now allocate a KV cache with 1000 cells (although not recommended), while on master this would have been automatically padded to 1024. Note that even with a total size of 1000 cells, most of the compute graphs will continue to have tensor shapes (n_kv) padded to 256 - f.ex. 256, 512, 768.

The main goal here is to simplify the logic around context size allocation per sequence and localize it during the context creation.

Comment thread tools/server/server.cpp Outdated
const auto n_ctx_train = llama_model_n_ctx_train(model);

if (slot.task->params.n_predict < 1 && slot.n_prompt_tokens() + slot.n_decoded >= n_ctx_train) {
if (slot.task->params.n_predict < 1 && slot.n_prompt_tokens() + slot.n_decoded >= slot.n_ctx) {

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The old logic of using the training context as a generation limit seems dubious - the context size of the slot should impose the limit.

After #16736, the slot.n_ctx will be capped to the training context either way, so this change should not make a really big difference either way.

@ngxson ngxson Oct 28, 2025 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In this case, do you think this code branch can be removed altogether? The n_ctx cap is already imposed on the slot.n_past >= slot.n_ctx condition above (L2933)

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, good point.

@github-actions github-actions Bot added examples python python script changes server labels Oct 28, 2025
Comment thread tools/server/tests/unit/test_ctx_shift.py
Comment thread tools/server/server.cpp
@@ -2866,10 +2866,12 @@ struct server_context {

// if context shifting is disabled, make sure that we don't run out of context
if (!params_base.ctx_shift && slot.n_past + 1 >= slot.n_ctx) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just note that slot.n_past won't reflect the correct number of tokens in KV cache in case the model uses M-RoPE. We should fix this later but I'm not entirely sure how. I'm thinking of these 2 solutions:

  • Rely on the server_tokens::size()
  • Add an API like llama_memory_seq_is_full which returns true if the memory is full

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, I'll open a PR to fix this.

@ggerganov
ggerganov merged commit 85a7d86 into master Oct 28, 2025
71 checks passed
@ggerganov
ggerganov deleted the gg/context-no-pad branch October 28, 2025 18:19
blime4 referenced this pull request in blime4/llama.cpp Feb 5, 2026
* memory : remove KV cache size padding

* cont : restore padding for n_kv tensor shape

* server : use slot context size instead of training context size

* server : simplify context limit logic
Seunghhon pushed a commit to Seunghhon/llama.cpp that referenced this pull request Apr 26, 2026
* memory : remove KV cache size padding

* cont : restore padding for n_kv tensor shape

* server : use slot context size instead of training context size

* server : simplify context limit logic
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
* memory : remove KV cache size padding

* cont : restore padding for n_kv tensor shape

* server : use slot context size instead of training context size

* server : simplify context limit logic
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
* memory : remove KV cache size padding

* cont : restore padding for n_kv tensor shape

* server : use slot context size instead of training context size

* server : simplify context limit logic
MrLordCat referenced this pull request in MrLordCat/llama.cpp-rdna-lab Jul 16, 2026
* memory : remove KV cache size padding

* cont : restore padding for n_kv tensor shape

* server : use slot context size instead of training context size

* server : simplify context limit logic
zhsy12345689 pushed a commit to zhsy12345689/llama.cpp-omni that referenced this pull request Jul 30, 2026
Back-port of upstream llama.cpp ggml-org/llama.cpp@85a7d867 (PR ggml-org#16812,
"memory : remove KV cache size padding"). This fork was sync'd to a
newer test_ctx_shift.py expecting predicted_n=120 against old padded-KV
server behavior, but server.cpp's no-shift early-stop path already runs
long enough to hit n_past+1 >= n_ctx (predicted_n=248), so CI has been
red on every master run for ~3 months.

Two surgical changes, both verbatim from upstream 85a7d86:
- tools/server/server.cpp: set slot.truncated=true in the
  !ctx_shift && n_past+1 >= n_ctx early-stop block (otherwise the
  flag is never set when context fills without ctx_shift).
- tools/server/tests/unit/test_ctx_shift.py: update the expected
  (n_predict=-1) parametrization from (120, True) to (248, True)
  with comment "8 tokens prompt + 248 tokens generated = 256 tokens
  total", matching upstream.

Verified locally by rebuilding llama-server on this branch and running
the full tools/server/tests/unit/test_ctx_shift.py (5 passed) plus the
adjacent test_completion.py::test_completion{,_stream} (4 passed).

Made-with: Cursor
zhsy12345689 pushed a commit to zhsy12345689/llama.cpp-omni that referenced this pull request Jul 30, 2026
Back-port of upstream llama.cpp ggml-org/llama.cpp@85a7d867 (PR ggml-org#16812,
"memory : remove KV cache size padding"). This fork was sync'd to a
newer test_ctx_shift.py expecting predicted_n=120 against old padded-KV
server behavior, but server.cpp's no-shift early-stop path already runs
long enough to hit n_past+1 >= n_ctx (predicted_n=248), so CI has been
red on every master run for ~3 months.

Two surgical changes, both verbatim from upstream 85a7d86:
- tools/server/server.cpp: set slot.truncated=true in the
  !ctx_shift && n_past+1 >= n_ctx early-stop block (otherwise the
  flag is never set when context fills without ctx_shift).
- tools/server/tests/unit/test_ctx_shift.py: update the expected
  (n_predict=-1) parametrization from (120, True) to (248, True)
  with comment "8 tokens prompt + 248 tokens generated = 256 tokens
  total", matching upstream.

Verified locally by rebuilding llama-server on this branch and running
the full tools/server/tests/unit/test_ctx_shift.py (5 passed) plus the
adjacent test_completion.py::test_completion{,_stream} (4 passed).

Signed-off-by: Halcao <nyz1500@gmail.com>
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
* memory : remove KV cache size padding

* cont : restore padding for n_kv tensor shape

* server : use slot context size instead of training context size

* server : simplify context limit logic
kainlan added a commit to kainlan/llama.cpp-intel-optimizations that referenced this pull request Sep 12, 2026
…rd (llama.cpp-oyfl)

What: extends ggml_backend_sycl_set_runtime_context()/_for_model() with a
seventh parameter, reserved_compute_buffer_bytes -- ggml_backend_sched_get
_buffer_size() for the SYCL backend, read after llama_context::sched_reserve()'s
graph-reserve passes complete. llama_context::sycl_resync_runtime_context_flash
_attn() (the shared helper introduced for the AUTO flash_attn re-check) gains
a query_reserved_compute_buffer bool: false (the default) forwards 0 for the
constructor's and resolve_fused_ops()'s calls, both of which necessarily run
before sched_reserve() has allocated anything; a new call at the end of
sched_reserve() itself passes true, querying the real per-backend size right
there (correct per-device even with multiple SYCL backends). Updates the
three call sites the API extension touches again (both nvfp4 device tests,
the lifecycle-wrapper's null-arg negative test).

Why: spec-review finding F11. llama_kv_cache::get_n_kv() sizes graph_reserve()'s
compute buffer from currently-used cells (0 at construction), padded to
n_pad_cur -- not the real n_ctx (a deliberate upstream design for graph-shape
stability, cited PR ggml-org#16812 in the code comment; not a fork bug, and not
something this ticket changes). So the buffer WILL regrow, outside the fixed
SYCL arena, once a real ubatch reaches the actual n_kv -- and the previous
guard (SCRATCH-zone capacity alone) had no way to account for that regrowth,
so it never fired at the exact repro shape (384/768 MiB demand under either
factor, against a 512 MiB zone) even though the real failure -- confirmed by
reading unified_alloc()'s "if zone is full, fall through to raw device
malloc below" fallthrough -- was outside-arena headroom exhaustion, not a
SCRATCH zone overflow in isolation. Getting the compute buffer's real size
(not guessing it) is what the next commit's redesigned predicate needs.

llama_memory_context_i (src/llama-memory.h) does not expose get_n_kv() --
only its ~7 concrete implementations do -- so the exact reserve-time n_kv
is not reachable without a new virtual method, which is out of proportion
here; the redesigned predicate (next commit) uses the full new kq/kqv size
without subtracting what the smaller reserve-time kq already contributed,
a deliberate, small (tens of MB) over-count in the safe direction.

Verification: `ctest --test-dir build -R test-sycl-nonfa-attn-scratch-demand`
unaffected (2/2 passed) -- this commit does not touch the demand formula.
Full pytest/format/bench-guard verification is reported with the next
commit, which is the one that actually consumes reserved_compute_buffer_bytes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TnqF7SEQf7GLk6nJTUuHWk
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

examples python python script changes server

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants