Summary
With the RPC backend active (--rpc host:port,...), llama-server's prompt-cache
save path aborts the entire server the first time it tries to save a slot's
prompt state. The save path (server_slot prompt saving, enabled by default
via --cache-ram) calls rpc_buffer_get_tensor, which hits an abort in the
RPC backend. Launching with --cache-ram 0 fully avoids the crash.
This is easy to misread as a random crash under load: the server starts,
answers a short first request, then dies on the first request large enough to
trigger a cache save. It is deterministic once you know the trigger.
Environment
- build: 11924d4, GNU 13.3.0, Linux aarch64
- 3× NVIDIA GB10 (DGX Spark), CUDA backend on all nodes
- coordinator + 2×
ggml-rpc-server workers (one over a ConnectX-7
point-to-point link, one over gigabit LAN)
- model: Qwen3-235B-A22B-Instruct-2507 Q6_K (split GGUF, 4 shards)
Repro
- Start two rpc-server workers.
- Start the server with the RPC backend and defaults otherwise:
llama-server -m Qwen3-235B-A22B-Instruct-2507-Q6_K-00001-of-00004.gguf \
--rpc 169.254.1.2:50052,192.0.2.35:50052 \
--tensor-split 45,40,15 -ngl 999 --split-mode layer -fit off \
-c 65536 -np 1 --host 0.0.0.0 --port 8000
- Send a chat completion with a prompt of a few thousand tokens, then a second
request (anything that makes the slot's prompt eligible for cache save).
- Server aborts (rpc_buffer_get_tensor path).
Expected
Either the prompt cache works over RPC, or the server declines to enable
prompt-cache save when RPC devices are present (with a log line), rather than
aborting mid-service.
Workaround
--cache-ram 0 — with it, the same 3-node setup has served multi-hour
sessions without a crash.
Notes
- Suggestion: gate the cache-save feature on backend capability, or make
rpc_buffer_get_tensor failures non-fatal for the cache path specifically.
- Happy to re-run with extra logging or a patch — this cluster reproduces it
on demand, and I can attach the full abort log from a fresh repro if useful.
Summary
With the RPC backend active (
--rpc host:port,...), llama-server's prompt-cachesave path aborts the entire server the first time it tries to save a slot's
prompt state. The save path (
server_slotprompt saving, enabled by defaultvia
--cache-ram) callsrpc_buffer_get_tensor, which hits an abort in theRPC backend. Launching with
--cache-ram 0fully avoids the crash.This is easy to misread as a random crash under load: the server starts,
answers a short first request, then dies on the first request large enough to
trigger a cache save. It is deterministic once you know the trigger.
Environment
ggml-rpc-serverworkers (one over a ConnectX-7point-to-point link, one over gigabit LAN)
Repro
request (anything that makes the slot's prompt eligible for cache save).
Expected
Either the prompt cache works over RPC, or the server declines to enable
prompt-cache save when RPC devices are present (with a log line), rather than
aborting mid-service.
Workaround
--cache-ram 0— with it, the same 3-node setup has served multi-hoursessions without a crash.
Notes
rpc_buffer_get_tensorfailures non-fatal for the cache path specifically.on demand, and I can attach the full abort log from a fresh repro if useful.