server /props: memory_breakdown (weights, KV, compute per buffer type) — the lane's footprint as the engine measures it - #25
Conversation
…lds, per buffer type (weights, KV cache, compute), as the engine allocated it A scheduler needs a lane's footprint. On Windows (WDDM) there is no per-process GPU reading, and an adopted lane has no spawn baseline for a device delta, so the continuum core logged "the lane's footprint could not be read" every minute on the 5090 (card 27fe9f8b; it holds continuum ggml-org#4396). The engine knows exactly what it allocated: llama_get_memory_breakdown(ctx) (model / context / compute bytes per buffer type). Taken once at load, after the context exists (KV and compute are sized there), and served from the cached meta, so /props stays readable while the server sleeps and never touches the context from the HTTP thread. 1.5B Q4_K_M, -c 4096 -np 2 on the 5090: CUDA0 model 980104704 (= model_weight_buffers' CUDA0), context 117440512, compute 66594944; CPU_Mapped model 191439360; CUDA_Host compute 8400928. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
Fable: APPROVED at f617a58 (the word is on airc). A small and correct change: it reports every buffer of the serving context, per buffer type, from the meta cached at load. Not blocking: |
|
APPROVED at f617a58 (Cormac). 🤖 Generated with Claude Code |
A scheduler needs a lane's footprint. On Windows (WDDM) there is no per-process GPU reading, and an
adopted lane has no spawn baseline for a device delta, so the continuum core logged "the lane's
footprint could not be read" every minute on the 5090 (card 27fe9f8b; it holds continuum ggml-org#4396).
The engine knows exactly what it allocated: llama_get_memory_breakdown(ctx) (model / context /
compute bytes per buffer type). Taken once at load, after the context exists (KV and compute are
sized there), and served from the cached meta, so /props stays readable while the server sleeps
and never touches the context from the HTTP thread.
1.5B Q4_K_M, -c 4096 -np 2 on the 5090: CUDA0 model 980104704 (= model_weight_buffers' CUDA0),
context 117440512, compute 66594944; CPU_Mapped model 191439360; CUDA_Host compute 8400928.
🤖 Generated with Claude Code
https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc