Repository navigation
cuda: reserve space for quantize kv-cache at startup - #23907
Conversation
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
|
Did you test |
|
I was going to say that |
|
I did not check yet, we can wait for #23792 to get merged and check again because currently it will throw. |
i have after also merging #23792, and as far as i see there is not more memory creeping, on my system (5060ti + 2060 super) in all the following cases there is a 125 mp allocation during first prompt processing without kv quant with or without kv quant + ngram-mod on a different issue with ngram-mod with or without kv quant |
|
A couple of other things that I believe the fit algorithm is not reserving space for: |
|
I think this should be ok to merge. |
|
@ggml-org/ggml-cuda can I get another approval? |
|
cc @ggml-org/maintainers, need another approval |
|
Thank you! This commit solves the issue #23978 |
|
Hello @am17an, this change has resulted in an additional increase in GPU memory usage. b9488 -> GPU:0 13.8/16.0 GB GPU:1 13.8/16.0 GB |
|
An increase in VRAM consumption is expected since llama.cpp is now pre-allocating the VRAM as part of the compute graphs. On master the initial VRAM consumption would be lower but eventually end up higher as the context fills up because the VRAM for converting the KV cache cannot be recycled for other operations. And previously the crash from OOMing would only happen after the program has already been running for some time which is undesirable if you're not babysitting it. In any case, the VRAM consumption for KV cache conversion can be reduced by setting |
|
After the test, after the pre-allocation adjustment was made, and after conducting multiple tests with long conversations, the GPU memory usage of my computer only increased by approximately 300 MB. Then is it absolutely safe from OOM? |
* cuda: reserve space for quantize kv-cache at startup * address review comments * remove forward decl Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * remove assert in ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* cuda: reserve space for quantize kv-cache at startup * address review comments * remove forward decl Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * remove assert in ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Investigation reference for a reproducible OOM on Arc Pro B70 (SYCL): the --fit fitter sizes n_ctx against mb.total() = model + context + compute, but a large amount of device memory is allocated outside that accounting, so it over-commits. Measured error is 2265 MiB (f16 KV) to 3562 MiB (q8_0 KV) at full context depth, against an effective margin of 2160 MiB. The unaccounted term is ~1.2 GiB at load plus a depth-dependent part that grows at 3.56 KiB/token with f16 KV and 7.64 KiB/token with q8_0. Since quantizing the KV cache both frees budget that --fit spends on a deeper context and roughly doubles the per-token unaccounted cost, quantized-KV configurations are more likely to OOM than the f16 ones they replace -- even though at a pinned --ctx-size quantized KV behaves correctly (a control measurement is included). The doc identifies three classes of unaccounted allocation, notes that CUDA already solved the flash-attention part by reserving it via get_alloc_size (f8f0a47, upstream ggml-org#23907) and that SYCL never got the port, and sets out staged follow-up work with an acceptance test. scripts/server-vram-probe.py implements the three measurements (compare, ladder, replicate) against a running server, reading real device memory from llamacpp:vram_* rather than tallying ggml buffers. It guards the two traps that silently corrupt results: free == total when ZES_ENABLE_SYSMAN is unset, and unsettled readings after a router model swap. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…gml-org#23907) The flash-attention paths stage F16 copies of K and V -- and for oneDNN SDPA also a dense Q, a scale scalar and an F16 output -- and took them from the scratch pool or, for K/V, from a grow-only side buffer. None of that is visible to llama.cpp's memory breakdown, so --fit cannot budget it, and the requirement scales with n_kv: a context the fitter accepted could still OOM part-way through a long prefill. Measured on an Arc Pro B70, device usage grew +606 MiB (f16 KV) to +1871 MiB (q8_0 KV) over a full-context prefill, against a fixed margin, and the OOM landed in ggml_sycl_pool_vmm::alloc under ggml_sycl_flash_attn_ext_onednn. Reserve the space as padding on the FA output tensor instead, as CUDA does (ggml_cuda_flash_attn_ext_get_alloc_size, upstream ggml-org#23907). The graph allocator then accounts for it, and because llama_context reserves the worst-case graph with the full KV cache (memory->init_full()), the reservation is sized for the deepest context up front -- so the growth disappears from runtime rather than merely being budgeted. ggml_sycl_fattn_get_extra() is the single source of truth for the layout, used both by get_alloc_size to size the reservation and by the kernels to obtain the pointers, so the two cannot disagree. Per kernel: ONEDNN reserves Q/K/V/scale/ out; TILE reserves K/V when they are not already F16 (aliasing V to K when V is a view of K, as before); VEC needs nothing since it reads quantized K/V natively; MKL needs nothing since it dequantizes one KV-head chunk at a time. Safe to call before dst->data is assigned, since ggml buffers are 128-aligned like SYCL_BUFFER_ALIGNMENT, so the padding is base-independent. GGML_SYCL_FA_RESERVE=0 restores the previous pool/side-buffer behaviour, as an escape hatch if the reservation misbehaves on hardware not covered below. Verified: builds clean with icpx 2026.1.1, and on an Intel iGPU with --cache-type-k/v q8_0 (which puts prefill on TILE, so the K/V staging is exercised) output is byte-identical with GGML_SYCL_FA_RESERVE=1 and 0. NOT verified: the oneDNN SDPA path. oneDNN flash attention is Battlemage-only (fattn-onednn.cpp gates on intel_gpu_bmg_g21/g31), so it cannot be exercised on the Xe-LPG iGPU available here. That path needs output-correctness checking on Battlemage, not just absence of an OOM -- a layout mismatch would corrupt silently. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Move the f16 KV-dequant temps of flash attention into the FLASH_ATTN_EXT compute buffer (sized via ggml_cuda_flash_attn_ext_get_alloc_size) instead of allocating them from the memory pool per launch. The pool keeps every historical peak-sized allocation, so on growing-context sessions each graph re-capture stacked another ~2x-KV temp and OOMed at long context. The compute buffer has a stable address (capture-safe by construction) and the allocator grows it in place while dropping the old storage, so physical memory stops climbing with each context increase. Remove the HIP raw-malloc branch and the fa_f16_use_pool toggle, which only worked around these two problems; this matches the upstream design (ggml-org/llama.cpp ggml-org#23907).
* cuda: reserve space for quantize kv-cache at startup * address review comments * remove forward decl Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * remove assert in ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Move the f16 KV-dequant temps of flash attention into the FLASH_ATTN_EXT compute buffer (sized via ggml_cuda_flash_attn_ext_get_alloc_size) instead of allocating them from the memory pool per launch. The pool keeps every historical peak-sized allocation, so on growing-context sessions each graph re-capture stacked another ~2x-KV temp and OOMed at long context. The compute buffer has a stable address (capture-safe by construction) and the allocator grows it in place while dropping the old storage, so physical memory stops climbing with each context increase. Remove the HIP raw-malloc branch and the fa_f16_use_pool toggle, which only worked around these two problems; this matches the upstream design (ggml-org/llama.cpp ggml-org#23907).
* cuda: reserve space for quantize kv-cache at startup * address review comments * remove forward decl Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * remove assert in ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Overview
ref #23646 (comment). Quantized kv-cache can lead to OOM even when using
--fitsince it does not know about these backend allocations. There are some other quantization buffers in FA and MMQ which should also be removed, but this one seems it takes the most space as it scales with ctx size.Additional information
Requirements