Skip to content

server: prompt-cache save aborts with RPC backend (rpc_buffer_get_tensor abort); --cache-ram 0 works around it #26529

Description

@ClosedLadder

Summary

With the RPC backend active (--rpc host:port,...), llama-server's prompt-cache
save path aborts the entire server the first time it tries to save a slot's
prompt state. The save path (server_slot prompt saving, enabled by default
via --cache-ram) calls rpc_buffer_get_tensor, which hits an abort in the
RPC backend. Launching with --cache-ram 0 fully avoids the crash.

This is easy to misread as a random crash under load: the server starts,
answers a short first request, then dies on the first request large enough to
trigger a cache save. It is deterministic once you know the trigger.

Environment

  • build: 11924d4, GNU 13.3.0, Linux aarch64
  • 3× NVIDIA GB10 (DGX Spark), CUDA backend on all nodes
  • coordinator + 2× ggml-rpc-server workers (one over a ConnectX-7
    point-to-point link, one over gigabit LAN)
  • model: Qwen3-235B-A22B-Instruct-2507 Q6_K (split GGUF, 4 shards)

Repro

  1. Start two rpc-server workers.
  2. Start the server with the RPC backend and defaults otherwise:
llama-server -m Qwen3-235B-A22B-Instruct-2507-Q6_K-00001-of-00004.gguf \
  --rpc 169.254.1.2:50052,192.0.2.35:50052 \
  --tensor-split 45,40,15 -ngl 999 --split-mode layer -fit off \
  -c 65536 -np 1 --host 0.0.0.0 --port 8000
  1. Send a chat completion with a prompt of a few thousand tokens, then a second
    request (anything that makes the slot's prompt eligible for cache save).
  2. Server aborts (rpc_buffer_get_tensor path).

Expected

Either the prompt cache works over RPC, or the server declines to enable
prompt-cache save when RPC devices are present (with a log line), rather than
aborting mid-service.

Workaround

--cache-ram 0 — with it, the same 3-node setup has served multi-hour
sessions without a crash.

Notes

  • Suggestion: gate the cache-save feature on backend capability, or make
    rpc_buffer_get_tensor failures non-fatal for the cache path specifically.
  • Happy to re-run with extra logging or a patch — this cluster reproduces it
    on demand, and I can attach the full abort log from a fresh repro if useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions