memory : remove KV cache size padding - #16812
Conversation
|
FYI in the CUDA backend, while non-padded inputs will produce correct results, the performance will be worse. One reason is that the vector and mma kernels don't have support for it (in principle also the wmma kernel but I want to remove that one soon). Another reason is that like with e.g. ALiBi I've added support for a non-padded KV cache only to the template specialization that doesn't have GQA-specific optimizations (to keep compilation times low and to mitigate the few % performance penalty from the OOB checks). |
|
What is the optimal padding for CUDA - 128 or 256? |
|
As of right now the required padding is still 256. The vector and tile kernels I've already adapted to be able to only need a padding of 128 (without OOB checks). The mma kernel, wmma kernel, and some utility kernels still need a padding of 256. My plan is to implement support for Volta, AMD WMMA instructions, and AMD MFMA instructions directly in the mma kernel, then I can just remove the wmma kernel without having to make any changes to it. The hardware I need for development are a V100 (which should arrive today), a MI100 (which arrived yesterday), and an RDNA3+ GPU (already in hand). The mma kernel and the utility kernels can then be made to work with a padding of 128 without issue. Caveat: the MI100 doesn't seem to work with the motherboard that I intended to use with it so I may need to shuffle around my hardware a bit. |
|
@JohannesGaessler The updated PR now continues to pad the For example, we can now allocate a KV cache with The main goal here is to simplify the logic around context size allocation per sequence and localize it during the context creation. |
1473d59 to
7ebe7f7
Compare
| const auto n_ctx_train = llama_model_n_ctx_train(model); | ||
|
|
||
| if (slot.task->params.n_predict < 1 && slot.n_prompt_tokens() + slot.n_decoded >= n_ctx_train) { | ||
| if (slot.task->params.n_predict < 1 && slot.n_prompt_tokens() + slot.n_decoded >= slot.n_ctx) { |
There was a problem hiding this comment.
The old logic of using the training context as a generation limit seems dubious - the context size of the slot should impose the limit.
After #16736, the slot.n_ctx will be capped to the training context either way, so this change should not make a really big difference either way.
There was a problem hiding this comment.
In this case, do you think this code branch can be removed altogether? The n_ctx cap is already imposed on the slot.n_past >= slot.n_ctx condition above (L2933)
| @@ -2866,10 +2866,12 @@ struct server_context { | |||
|
|
|||
| // if context shifting is disabled, make sure that we don't run out of context | |||
| if (!params_base.ctx_shift && slot.n_past + 1 >= slot.n_ctx) { | |||
There was a problem hiding this comment.
Just note that slot.n_past won't reflect the correct number of tokens in KV cache in case the model uses M-RoPE. We should fix this later but I'm not entirely sure how. I'm thinking of these 2 solutions:
- Rely on the
server_tokens::size() - Add an API like
llama_memory_seq_is_fullwhich returns true if the memory is full
There was a problem hiding this comment.
Yes, I'll open a PR to fix this.
* memory : remove KV cache size padding * cont : restore padding for n_kv tensor shape * server : use slot context size instead of training context size * server : simplify context limit logic
* memory : remove KV cache size padding * cont : restore padding for n_kv tensor shape * server : use slot context size instead of training context size * server : simplify context limit logic
* memory : remove KV cache size padding * cont : restore padding for n_kv tensor shape * server : use slot context size instead of training context size * server : simplify context limit logic
* memory : remove KV cache size padding * cont : restore padding for n_kv tensor shape * server : use slot context size instead of training context size * server : simplify context limit logic
* memory : remove KV cache size padding * cont : restore padding for n_kv tensor shape * server : use slot context size instead of training context size * server : simplify context limit logic
Back-port of upstream llama.cpp ggml-org/llama.cpp@85a7d867 (PR ggml-org#16812, "memory : remove KV cache size padding"). This fork was sync'd to a newer test_ctx_shift.py expecting predicted_n=120 against old padded-KV server behavior, but server.cpp's no-shift early-stop path already runs long enough to hit n_past+1 >= n_ctx (predicted_n=248), so CI has been red on every master run for ~3 months. Two surgical changes, both verbatim from upstream 85a7d86: - tools/server/server.cpp: set slot.truncated=true in the !ctx_shift && n_past+1 >= n_ctx early-stop block (otherwise the flag is never set when context fills without ctx_shift). - tools/server/tests/unit/test_ctx_shift.py: update the expected (n_predict=-1) parametrization from (120, True) to (248, True) with comment "8 tokens prompt + 248 tokens generated = 256 tokens total", matching upstream. Verified locally by rebuilding llama-server on this branch and running the full tools/server/tests/unit/test_ctx_shift.py (5 passed) plus the adjacent test_completion.py::test_completion{,_stream} (4 passed). Made-with: Cursor
Back-port of upstream llama.cpp ggml-org/llama.cpp@85a7d867 (PR ggml-org#16812, "memory : remove KV cache size padding"). This fork was sync'd to a newer test_ctx_shift.py expecting predicted_n=120 against old padded-KV server behavior, but server.cpp's no-shift early-stop path already runs long enough to hit n_past+1 >= n_ctx (predicted_n=248), so CI has been red on every master run for ~3 months. Two surgical changes, both verbatim from upstream 85a7d86: - tools/server/server.cpp: set slot.truncated=true in the !ctx_shift && n_past+1 >= n_ctx early-stop block (otherwise the flag is never set when context fills without ctx_shift). - tools/server/tests/unit/test_ctx_shift.py: update the expected (n_predict=-1) parametrization from (120, True) to (248, True) with comment "8 tokens prompt + 248 tokens generated = 256 tokens total", matching upstream. Verified locally by rebuilding llama-server on this branch and running the full tools/server/tests/unit/test_ctx_shift.py (5 passed) plus the adjacent test_completion.py::test_completion{,_stream} (4 passed). Signed-off-by: Halcao <nyz1500@gmail.com>
* memory : remove KV cache size padding * cont : restore padding for n_kv tensor shape * server : use slot context size instead of training context size * server : simplify context limit logic
…rd (llama.cpp-oyfl) What: extends ggml_backend_sycl_set_runtime_context()/_for_model() with a seventh parameter, reserved_compute_buffer_bytes -- ggml_backend_sched_get _buffer_size() for the SYCL backend, read after llama_context::sched_reserve()'s graph-reserve passes complete. llama_context::sycl_resync_runtime_context_flash _attn() (the shared helper introduced for the AUTO flash_attn re-check) gains a query_reserved_compute_buffer bool: false (the default) forwards 0 for the constructor's and resolve_fused_ops()'s calls, both of which necessarily run before sched_reserve() has allocated anything; a new call at the end of sched_reserve() itself passes true, querying the real per-backend size right there (correct per-device even with multiple SYCL backends). Updates the three call sites the API extension touches again (both nvfp4 device tests, the lifecycle-wrapper's null-arg negative test). Why: spec-review finding F11. llama_kv_cache::get_n_kv() sizes graph_reserve()'s compute buffer from currently-used cells (0 at construction), padded to n_pad_cur -- not the real n_ctx (a deliberate upstream design for graph-shape stability, cited PR ggml-org#16812 in the code comment; not a fork bug, and not something this ticket changes). So the buffer WILL regrow, outside the fixed SYCL arena, once a real ubatch reaches the actual n_kv -- and the previous guard (SCRATCH-zone capacity alone) had no way to account for that regrowth, so it never fired at the exact repro shape (384/768 MiB demand under either factor, against a 512 MiB zone) even though the real failure -- confirmed by reading unified_alloc()'s "if zone is full, fall through to raw device malloc below" fallthrough -- was outside-arena headroom exhaustion, not a SCRATCH zone overflow in isolation. Getting the compute buffer's real size (not guessing it) is what the next commit's redesigned predicate needs. llama_memory_context_i (src/llama-memory.h) does not expose get_n_kv() -- only its ~7 concrete implementations do -- so the exact reserve-time n_kv is not reachable without a new virtual method, which is out of proportion here; the redesigned predicate (next commit) uses the full new kq/kqv size without subtracting what the smaller reserve-time kq already contributed, a deliberate, small (tens of MB) over-count in the safe direction. Verification: `ctest --test-dir build -R test-sycl-nonfa-attn-scratch-demand` unaffected (2/2 passed) -- this commit does not touch the demand formula. Full pytest/format/bench-guard verification is reported with the next commit, which is the one that actually consumes reserved_compute_buffer_bytes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TnqF7SEQf7GLk6nJTUuHWk
ref #16736 (comment)
Simplify the logic during memory module creation. Before, for KV caches, we used to pad the buffer size up to 256 cells since flash attention implementations did not support arbitrary K, V sizes. After the improvements in #16148 and related, we no longer need to explicitly do this padding.
For now, keeping support for
llama_kv_cachesize padding via the constructor'sn_padargument, although it is not currently used anymore.Note that we continue to pad
n_kv- this is the tensor shape for the K and V tensors for each graph:https://github.com/ggml-org/llama.cpp/blob/1473d59a7bb0716aee4ab104ab09ecf7d978bd40/src/llama-kv-cache.cpp#L957-L972
We need to do this in order to reuse most of the compute graphs during the text generation phase. Additionally, this helps the performance with some of the backends.
Also,
llama_model::create_memory()no longer mutatescparams.Next, will rebase #16736 on top of this change and finish it.