server: add ctx-per-slot (--kv-unified-per-slot) - #24124
Conversation
|
@ngxson can you take a look? |
|
IMO this change is quite excessive:
|
|
What do you mean by your first point? Right now I don't believe there is any way to limit how much context an individual slot can use when in unified mode, the only limit is the training context I don't think setting n_ctx will affect anything in unified mode specifically, that would work for non-unified As for the second, yeah I just liked it as a minor convenience, happy to remove it if it's too excessive |
|
Removed the I double checked and I'm 95% sure there is no way to limit the context per slot without adding the server code, there's probably a different way to do it in the server code, but something needs to happen there regardless |
|
Any other thoughts @ngxson ? |
|
I am using and the last part is the key: context 200k num-processes 4 What I would instead like is to have 1 slot with 100k and 2 slots with 50k each. Currently I can have non-unified with each slot divided equally or unified with context limit for total... and this is probably what I should be using for my own agent :) like |
hmm right, I was thinking about per-request since this feature is KV-unified-only, I'd suggest renaming it to Will be ok to merge after the renaming |
* Add ctx-per-slot argument for unifid KV cache * Swap out ctx fractions for ctx pool slots * Formatting cleanup * Remove ctx-pool-slots, make ctx-per-slot an int * refactor it --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
* Add ctx-per-slot argument for unifid KV cache * Swap out ctx fractions for ctx pool slots * Formatting cleanup * Remove ctx-pool-slots, make ctx-per-slot an int * refactor it --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
* Add ctx-per-slot argument for unifid KV cache * Swap out ctx fractions for ctx pool slots * Formatting cleanup * Remove ctx-pool-slots, make ctx-per-slot an int * refactor it --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
* Add ctx-per-slot argument for unifid KV cache * Swap out ctx fractions for ctx pool slots * Formatting cleanup * Remove ctx-pool-slots, make ctx-per-slot an int * refactor it --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
llama.cpp's `--ctx-size` is not a per-request cap: it is the entire KV pool, and `--parallel N` statically carves it into N slots holding `ctx/N` tokens each. The unified KV cache that would share one buffer instead is only enabled when the slot count is `auto`, and `start_llm` always passes `--parallel` explicitly — so the division always applied. Forwarding the operator's number raw meant `grid join --max-concurrency 4 --ctx-size 32000` gave each request 8000 tokens, while `remote/serve.py` advertised the undivided 32000 to the grid as `context_window`. The auto-router then ranked against a window four times larger than any single request could use. Scaling `-c` by the slot count makes the flag mean what every other runtime Grid fronts already means by it: vLLM's `--max-model-len` and SGLang's `--context-length` are both per-sequence caps, with the KV pool sized separately by a memory fraction. `context_window` needs no change — it sends the raw per-request number, which is now the truth rather than an N-fold over-promise. What does NOT carry over is the cost model. Those two page the KV cache and treat the number as a ceiling; llama.cpp reserves every slot's share up front. So `--max-concurrency 8 --ctx-size 32000` really does allocate 256k of KV and, per this launcher's existing fail-loud policy, refuses to boot rather than shrinking. `LlamaProcess` now carries the launch parameters so `wait_for_models` can say so: llama.cpp's own log reports only that it could not allocate 128000 tokens, and nothing in it connects that back to the 32000 the operator typed. BEHAVIOR CHANGE: anyone who already compensated by hand — `--ctx-size 128000 --max-concurrency 4` to get 32k per request — now asks for 512k and fails to start. Refs: ggml-org/llama.cpp#11681 (the division), ggml-org/llama.cpp#24124 (`--kv-unified-per-slot`, the fuller alignment — needs a build-floor check against MIN_LLAMA_SERVER_BUILD before it can be adopted). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WgR4BpPKeCLuunxjLKZogh
Overview
Adds two new arguments for llama-server context provisioning
kv-unified-per-slot
First argument,
--kv-unified-per-slotThis specifies how much context each slot is limited to in the
kvumode, so for example if a user wants to have N slots with each having 4096 context, they can specify--kv-unified-per-slot 4096, it will provision 16384 total context:If specified with
-c, you will then limit the total context:If you set
-cto a value lower than--kv-unified-per-slot, it will warn and cap to-c's value:Details
ctx-pool-slots
Also adds another new flag,
--ctx-pool-slots, which specifies how many slots worth of context should be allocated, this is similar to setting-cto an explicit value but doesn't require the user to manually do the math ahead of time:If you specify more pool slots than
np, it will warn and clamp tonp:Since this is purely a helper argument, I wouldn't mind dropping it from the PR, I think it would be nice to have but I've been told the server arguments are already getting a bit bloated...
Requirements