Skip to content

server: add ctx-per-slot (--kv-unified-per-slot) - #24124

Merged
ngxson merged 8 commits into
ggml-org:masterfrom
bartowski1182:auto_parallel
Aug 27, 2026
Merged

ngxson merged 8 commits into
ggml-org:masterfrom
bartowski1182:auto_parallel

Conversation

@bartowski1182

@bartowski1182 bartowski1182 commented Jun 4, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Adds two new arguments for llama-server context provisioning

kv-unified-per-slot

First argument, --kv-unified-per-slot

This specifies how much context each slot is limited to in the kvu mode, so for example if a user wants to have N slots with each having 4096 context, they can specify --kv-unified-per-slot 4096, it will provision 16384 total context:

./build/bin/llama-server -m Qwen3.5-4B-Q4_K_M.gguf --kv-unified-per-slot 4096
...
0.02.635.267 W llama_context: n_ctx_seq (16384) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
...
0.03.025.317 I srv    load_model: capping per-slot context (16384) to --kv-unified-per-slot (4096)
...
0.03.084.366 I slot   load_model: id  0 | task -1 | new slot, n_ctx = 4096
0.03.084.370 I slot   load_model: id  1 | task -1 | new slot, n_ctx = 4096
0.03.084.370 I slot   load_model: id  2 | task -1 | new slot, n_ctx = 4096
0.03.084.370 I slot   load_model: id  3 | task -1 | new slot, n_ctx = 4096
...

If specified with -c, you will then limit the total context:

./build-cpu/bin/llama-server -m Qwen3.5-4B-Q4_K_M.gguf -c 8192 --kv-unified-per-slot 4096
...
0.02.650.929 W llama_context: n_ctx_seq (8192) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
...
0.02.921.467 I srv    load_model: capping per-slot context (8192) to --kv-unified-per-slot (4096)
...
0.02.973.263 I slot   load_model: id  0 | task -1 | new slot, n_ctx = 4096
0.02.973.267 I slot   load_model: id  1 | task -1 | new slot, n_ctx = 4096
0.02.973.267 I slot   load_model: id  2 | task -1 | new slot, n_ctx = 4096
0.02.973.267 I slot   load_model: id  3 | task -1 | new slot, n_ctx = 4096
...

If you set -c to a value lower than --kv-unified-per-slot, it will warn and cap to -c's value:

load_model: --kv-unified-per-slot (4096) exceeds the per-slot pool capacity (2048) - cap has no effect, slots are limited to 2048 (raise the KV pool with -c, or unset -c to size it to n_parallel * kv_unified_per_slot)

removed ctx-pool-slots
Details

ctx-pool-slots

Also adds another new flag, --ctx-pool-slots, which specifies how many slots worth of context should be allocated, this is similar to setting -c to an explicit value but doesn't require the user to manually do the math ahead of time:

./build-cpu/bin/llama-server -m Qwen3.5-4B-Q4_K_M.gguf --kv-unified-per-slot 4096 --ctx-pool-slots 2
...
0.02.650.929 W llama_context: n_ctx_seq (8192) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
...
0.02.921.467 I srv    load_model: capping per-slot context (8192) to --kv-unified-per-slot (4096)
...
0.02.973.263 I slot   load_model: id  0 | task -1 | new slot, n_ctx = 4096
0.02.973.267 I slot   load_model: id  1 | task -1 | new slot, n_ctx = 4096
0.02.973.267 I slot   load_model: id  2 | task -1 | new slot, n_ctx = 4096
0.02.973.267 I slot   load_model: id  3 | task -1 | new slot, n_ctx = 4096
...

If you specify more pool slots than np, it will warn and clamp to np:

srv llama_server: --ctx-pool-slots (6) exceeds n_parallel (4), clamping

Since this is purely a helper argument, I wouldn't mind dropping it from the PR, I think it would be nice to have but I've been told the server arguments are already getting a bit bloated...

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, used for brainstorming and code changes/testing

@bartowski1182 bartowski1182 changed the title Add ctx-per-slot argument for unifid KV cache Add ctx-per-slot argument for unified KV cache Jun 4, 2026
@bartowski1182
bartowski1182 marked this pull request as ready for review June 4, 2026 13:56
@bartowski1182
bartowski1182 requested review from a team as code owners June 4, 2026 13:56
@bartowski1182

Copy link
Copy Markdown
Contributor Author

@ngxson can you take a look?

@ngxson

ngxson commented Jun 10, 2026 •

Copy link
Copy Markdown
Collaborator

IMO this change is quite excessive:

  • --ctx-per-slot make sense, but can be implemented as a simple n_ctx = n_ctx_per_slot * n_parallel, there is no need to implement any changes on server
  • --ctx-pool-slots is not very intuitive to understand. I don't think it worth adding because the use case is quite too small. you can simply do -c $(( 4096*2 )) in bash, which looks even cleaner

@bartowski1182

Copy link
Copy Markdown
Contributor Author

What do you mean by your first point?

Right now I don't believe there is any way to limit how much context an individual slot can use when in unified mode, the only limit is the training context

I don't think setting n_ctx will affect anything in unified mode specifically, that would work for non-unified

As for the second, yeah I just liked it as a minor convenience, happy to remove it if it's too excessive

@bartowski1182

Copy link
Copy Markdown
Contributor Author

Removed the ctx-pool-slots and put ctx-per-slot to an int (was intending to later add a max argument but will put it back to stoi if and when that happens)

I double checked and I'm 95% sure there is no way to limit the context per slot without adding the server code, there's probably a different way to do it in the server code, but something needs to happen there regardless

@bartowski1182

Copy link
Copy Markdown
Contributor Author

Any other thoughts @ngxson ?

@adderek

adderek commented Jun 25, 2026 •

Copy link
Copy Markdown

I am using

llama-server -m bartowski-Qwen_Qwen3.6-27B-IQ4_NL.gguf \
--port 8081 --cache-type-k turbo4 --cache-type-v turbo4 \
-fa on -ngl 99 -b 4096 -ub 4096 -t 12 --offline --models-max 1 --cont-batching --tools all --jinja \
--cache-ram 8192 --temperature 0.10 --ctx-checkpoints 4 --reasoning-budget 256 \
--spec-type ngram-cache --spec-draft-n-max 16 \
--kv-offload \
-c 200000 -np 4

and the last part is the key: context 200k num-processes 4
each gets 50k
Isn't that the same effect as you would get with -ctx-per-slot 50000 ?

What I would instead like is to have 1 slot with 100k and 2 slots with 50k each. Currently I can have non-unified with each slot divided equally or unified with context limit for total... and this is probably what I should be using for my own agent :) like --kv-unified -c 200000 -np3 allowing any to take up to 200k but crash (?) when 200k total is reached.

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Jul 13, 2026
@ngxson

ngxson commented Jul 24, 2026

Copy link
Copy Markdown
Collaborator

Right now I don't believe there is any way to limit how much context an individual slot can use when in unified mode, the only limit is the training context

hmm right, I was thinking about per-request max_tokens but that only limit the number of generated tokens, not whole context size

since this feature is KV-unified-only, I'd suggest renaming it to --kv-unified-per-slot

Will be ok to merge after the renaming

@ngxson ngxson self-assigned this Aug 27, 2026
@ngxson ngxson changed the title Add ctx-per-slot argument for unified KV cache server: add ctx-per-slot (--kv-unified-per-slot) Aug 27, 2026
@ngxson
ngxson merged commit 1844325 into ggml-org:master Aug 27, 2026
23 of 26 checks passed
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
* Add ctx-per-slot argument for unifid KV cache

* Swap out ctx fractions for ctx pool slots

* Formatting cleanup

* Remove ctx-pool-slots, make ctx-per-slot an int

* refactor it

---------

Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* Add ctx-per-slot argument for unifid KV cache

* Swap out ctx fractions for ctx pool slots

* Formatting cleanup

* Remove ctx-pool-slots, make ctx-per-slot an int

* refactor it

---------

Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* Add ctx-per-slot argument for unifid KV cache

* Swap out ctx fractions for ctx pool slots

* Formatting cleanup

* Remove ctx-pool-slots, make ctx-per-slot an int

* refactor it

---------

Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
* Add ctx-per-slot argument for unifid KV cache

* Swap out ctx fractions for ctx pool slots

* Formatting cleanup

* Remove ctx-pool-slots, make ctx-per-slot an int

* refactor it

---------

Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
ucchino pushed a commit to ucchino/autonomous-grid that referenced this pull request Oct 2, 2026
llama.cpp's `--ctx-size` is not a per-request cap: it is the entire KV pool, and
`--parallel N` statically carves it into N slots holding `ctx/N` tokens each. The
unified KV cache that would share one buffer instead is only enabled when the
slot count is `auto`, and `start_llm` always passes `--parallel` explicitly — so
the division always applied.

Forwarding the operator's number raw meant `grid join --max-concurrency 4
--ctx-size 32000` gave each request 8000 tokens, while `remote/serve.py`
advertised the undivided 32000 to the grid as `context_window`. The auto-router
then ranked against a window four times larger than any single request could use.

Scaling `-c` by the slot count makes the flag mean what every other runtime Grid
fronts already means by it: vLLM's `--max-model-len` and SGLang's
`--context-length` are both per-sequence caps, with the KV pool sized separately
by a memory fraction. `context_window` needs no change — it sends the raw
per-request number, which is now the truth rather than an N-fold over-promise.

What does NOT carry over is the cost model. Those two page the KV cache and treat
the number as a ceiling; llama.cpp reserves every slot's share up front. So
`--max-concurrency 8 --ctx-size 32000` really does allocate 256k of KV and, per
this launcher's existing fail-loud policy, refuses to boot rather than shrinking.
`LlamaProcess` now carries the launch parameters so `wait_for_models` can say so:
llama.cpp's own log reports only that it could not allocate 128000 tokens, and
nothing in it connects that back to the 32000 the operator typed.

BEHAVIOR CHANGE: anyone who already compensated by hand — `--ctx-size 128000
--max-concurrency 4` to get 32k per request — now asks for 512k and fails to start.

Refs: ggml-org/llama.cpp#11681 (the division), ggml-org/llama.cpp#24124
(`--kv-unified-per-slot`, the fuller alignment — needs a build-floor check
against MIN_LLAMA_SERVER_BUILD before it can be adopted).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WgR4BpPKeCLuunxjLKZogh
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation examples server

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants