Skip to content

cuda: reserve space for quantize kv-cache at startup - #23907

Merged
am17an merged 4 commits into
ggml-org:masterfrom
am17an:fattn-static-kv-cache
Jun 3, 2026
Merged

am17an merged 4 commits into
ggml-org:masterfrom
am17an:fattn-static-kv-cache

Conversation

@am17an

@am17an am17an commented May 30, 2026

Copy link
Copy Markdown
Contributor

Overview

ref #23646 (comment). Quantized kv-cache can lead to OOM even when using --fit since it does not know about these backend allocations. There are some other quantization buffers in FA and MMQ which should also be removed, but this one seems it takes the most space as it scales with ctx size.

Additional information

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, codex wrote this on my direction. I tested it on a few devices

@am17an
am17an requested a review from a team as a code owner May 30, 2026 10:25
@github-actions github-actions Bot added Nvidia GPU Issues specific to Nvidia GPUs ggml changes relating to the ggml tensor library for machine learning labels May 30, 2026
Comment thread ggml/src/ggml-cuda/fattn-common.cuh Outdated
Comment thread ggml/src/ggml-cuda/fattn.cu
Comment thread ggml/src/ggml-cuda/fattn.cu Outdated
Comment thread ggml/src/ggml-cuda/common.cuh Outdated
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Comment thread ggml/src/ggml-cuda/fattn-common.cuh Outdated
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
@ggerganov

Copy link
Copy Markdown
Member

Did you test -sm tensor setups? I think it should work, but might be worth double-checking.

@JohannesGaessler

Copy link
Copy Markdown
Contributor

I was going to say that -sm tensor is implicitly being tested via test-llama-archs but there only the FP16/FP16 configuration is being tested. More generally: do we already have automated tests for different KV cache types? If not it may make sense to add some in test-llama-archs.

@am17an

am17an commented May 31, 2026

Copy link
Copy Markdown
Contributor Author

I did not check yet, we can wait for #23792 to get merged and check again because currently it will throw.

@AbdulrahmanHashem

Copy link
Copy Markdown

Did you test -sm tensor setups? I think it should work, but might be worth double-checking.

i have after also merging #23792, and as far as i see there is not more memory creeping, on my system (5060ti + 2060 super)

in all the following cases there is a 125 mp allocation during first prompt processing

without kv quant
without kv quant + MTP
without kv quant + ngram-mod
with kv quant
with kv quant + MTP
with kv quant + ngram-mod

with or without kv quant + ngram-mod
there is an additional variable amount of allocation under 150mp during tg

on a different issue with ngram-mod with or without kv quant
i tried ngram-mod + MTP it causes a crash with no error during tg after thinking is done.
it lags the llama ui very hard and just stops tg with no logs before it crashes

@coder543

Copy link
Copy Markdown
Contributor

A couple of other things that I believe the fit algorithm is not reserving space for: cache-ram and ctx-checkpoints. On unified memory systems like the DGX Spark, this makes it hard to rely on the fit algorithm without specifying an arbitrarily large fit-target.

@am17an

am17an commented Jun 2, 2026

Copy link
Copy Markdown
Contributor Author

I think this should be ok to merge.

@am17an

am17an commented Jun 3, 2026 •

Copy link
Copy Markdown
Contributor Author

@ggml-org/ggml-cuda can I get another approval?

@am17an

am17an commented Jun 3, 2026

Copy link
Copy Markdown
Contributor Author

cc @ggml-org/maintainers, need another approval

@am17an
am17an merged commit f8f0a47 into ggml-org:master Jun 3, 2026
31 checks passed
@am17an
am17an deleted the fattn-static-kv-cache branch June 3, 2026 10:40
@TomTheWise

Copy link
Copy Markdown

Thank you! This commit solves the issue #23978

@thomasbergersen

Copy link
Copy Markdown

Hello @am17an, this change has resulted in an additional increase in GPU memory usage.

b9488 -> GPU:0 13.8/16.0 GB GPU:1 13.8/16.0 GB
b9489 -> GPU:0 14.9/16.0 GB GPU:1 14.2/16.0 GB

@JohannesGaessler

Copy link
Copy Markdown
Contributor

An increase in VRAM consumption is expected since llama.cpp is now pre-allocating the VRAM as part of the compute graphs. On master the initial VRAM consumption would be lower but eventually end up higher as the context fills up because the VRAM for converting the KV cache cannot be recycled for other operations. And previously the crash from OOMing would only happen after the program has already been running for some time which is undesirable if you're not babysitting it.

In any case, the VRAM consumption for KV cache conversion can be reduced by setting -ub to a lower value than 512.

@thomasbergersen

Copy link
Copy Markdown

After the test, after the pre-allocation adjustment was made, and after conducting multiple tests with long conversations, the GPU memory usage of my computer only increased by approximately 300 MB. Then is it absolutely safe from OOM?

m0nk111 added a commit to m0nk111/llama.cpp that referenced this pull request Jun 13, 2026
adrianhoehne pushed a commit to adrianhoehne/llama.cpp that referenced this pull request Jul 5, 2026
* cuda: reserve space for quantize kv-cache at startup

* address review comments

* remove forward decl

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* remove assert in ggml-cuda.cu

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
* cuda: reserve space for quantize kv-cache at startup

* address review comments

* remove forward decl

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* remove assert in ggml-cuda.cu

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
nicois added a commit to nicois/llama.cpp that referenced this pull request Aug 20, 2026
Investigation reference for a reproducible OOM on Arc Pro B70 (SYCL): the
--fit fitter sizes n_ctx against mb.total() = model + context + compute, but
a large amount of device memory is allocated outside that accounting, so it
over-commits. Measured error is 2265 MiB (f16 KV) to 3562 MiB (q8_0 KV) at
full context depth, against an effective margin of 2160 MiB.

The unaccounted term is ~1.2 GiB at load plus a depth-dependent part that
grows at 3.56 KiB/token with f16 KV and 7.64 KiB/token with q8_0. Since
quantizing the KV cache both frees budget that --fit spends on a deeper
context and roughly doubles the per-token unaccounted cost, quantized-KV
configurations are more likely to OOM than the f16 ones they replace --
even though at a pinned --ctx-size quantized KV behaves correctly (a
control measurement is included).

The doc identifies three classes of unaccounted allocation, notes that CUDA
already solved the flash-attention part by reserving it via get_alloc_size
(f8f0a47, upstream ggml-org#23907) and that SYCL never got the port, and sets out
staged follow-up work with an acceptance test.

scripts/server-vram-probe.py implements the three measurements (compare,
ladder, replicate) against a running server, reading real device memory from
llamacpp:vram_* rather than tallying ggml buffers. It guards the two traps
that silently corrupt results: free == total when ZES_ENABLE_SYSMAN is
unset, and unsettled readings after a router model swap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
nicois added a commit to nicois/llama.cpp that referenced this pull request Aug 20, 2026
…gml-org#23907)

The flash-attention paths stage F16 copies of K and V -- and for oneDNN SDPA also
a dense Q, a scale scalar and an F16 output -- and took them from the scratch
pool or, for K/V, from a grow-only side buffer. None of that is visible to
llama.cpp's memory breakdown, so --fit cannot budget it, and the requirement
scales with n_kv: a context the fitter accepted could still OOM part-way through
a long prefill. Measured on an Arc Pro B70, device usage grew +606 MiB (f16 KV)
to +1871 MiB (q8_0 KV) over a full-context prefill, against a fixed margin, and
the OOM landed in ggml_sycl_pool_vmm::alloc under
ggml_sycl_flash_attn_ext_onednn.

Reserve the space as padding on the FA output tensor instead, as CUDA does
(ggml_cuda_flash_attn_ext_get_alloc_size, upstream ggml-org#23907). The graph allocator
then accounts for it, and because llama_context reserves the worst-case graph
with the full KV cache (memory->init_full()), the reservation is sized for the
deepest context up front -- so the growth disappears from runtime rather than
merely being budgeted.

ggml_sycl_fattn_get_extra() is the single source of truth for the layout, used
both by get_alloc_size to size the reservation and by the kernels to obtain the
pointers, so the two cannot disagree. Per kernel: ONEDNN reserves Q/K/V/scale/
out; TILE reserves K/V when they are not already F16 (aliasing V to K when V is
a view of K, as before); VEC needs nothing since it reads quantized K/V
natively; MKL needs nothing since it dequantizes one KV-head chunk at a time.
Safe to call before dst->data is assigned, since ggml buffers are 128-aligned
like SYCL_BUFFER_ALIGNMENT, so the padding is base-independent.

GGML_SYCL_FA_RESERVE=0 restores the previous pool/side-buffer behaviour, as an
escape hatch if the reservation misbehaves on hardware not covered below.

Verified: builds clean with icpx 2026.1.1, and on an Intel iGPU with
--cache-type-k/v q8_0 (which puts prefill on TILE, so the K/V staging is
exercised) output is byte-identical with GGML_SYCL_FA_RESERVE=1 and 0.

NOT verified: the oneDNN SDPA path. oneDNN flash attention is Battlemage-only
(fattn-onednn.cpp gates on intel_gpu_bmg_g21/g31), so it cannot be exercised on
the Xe-LPG iGPU available here. That path needs output-correctness checking on
Battlemage, not just absence of an OOM -- a layout mismatch would corrupt
silently.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
shizhx added a commit to shizhx/llama-cpp-turboquant that referenced this pull request Sep 3, 2026
Move the f16 KV-dequant temps of flash attention into the FLASH_ATTN_EXT
compute buffer (sized via ggml_cuda_flash_attn_ext_get_alloc_size) instead
of allocating them from the memory pool per launch. The pool keeps every
historical peak-sized allocation, so on growing-context sessions each graph
re-capture stacked another ~2x-KV temp and OOMed at long context. The
compute buffer has a stable address (capture-safe by construction) and the
allocator grows it in place while dropping the old storage, so physical
memory stops climbing with each context increase. Remove the HIP raw-malloc
branch and the fa_f16_use_pool toggle, which only worked around these two
problems; this matches the upstream design (ggml-org/llama.cpp ggml-org#23907).
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* cuda: reserve space for quantize kv-cache at startup

* address review comments

* remove forward decl

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* remove assert in ggml-cuda.cu

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
TheTom pushed a commit to TheTom/llama-cpp-turboquant that referenced this pull request Sep 19, 2026
Move the f16 KV-dequant temps of flash attention into the FLASH_ATTN_EXT
compute buffer (sized via ggml_cuda_flash_attn_ext_get_alloc_size) instead
of allocating them from the memory pool per launch. The pool keeps every
historical peak-sized allocation, so on growing-context sessions each graph
re-capture stacked another ~2x-KV temp and OOMed at long context. The
compute buffer has a stable address (capture-safe by construction) and the
allocator grows it in place while dropping the old storage, so physical
memory stops climbing with each context increase. Remove the HIP raw-malloc
branch and the fa_f16_use_pool toggle, which only worked around these two
problems; this matches the upstream design (ggml-org/llama.cpp ggml-org#23907).
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* cuda: reserve space for quantize kv-cache at startup

* address review comments

* remove forward decl

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* remove assert in ggml-cuda.cu

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Nvidia GPU Issues specific to Nvidia GPUs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants