Skip to content

server /props: memory_breakdown (weights, KV, compute per buffer type) — the lane's footprint as the engine measures it - #25

Merged
joelteply merged 1 commit into
feat/props-weight-residencyfrom
continuum/props-memory-breakdown
Sep 27, 2026
Merged

joelteply merged 1 commit into
feat/props-weight-residencyfrom
continuum/props-memory-breakdown

Conversation

@joelteply

Copy link
Copy Markdown

A scheduler needs a lane's footprint. On Windows (WDDM) there is no per-process GPU reading, and an
adopted lane has no spawn baseline for a device delta, so the continuum core logged "the lane's
footprint could not be read" every minute on the 5090 (card 27fe9f8b; it holds continuum ggml-org#4396).
The engine knows exactly what it allocated: llama_get_memory_breakdown(ctx) (model / context /
compute bytes per buffer type). Taken once at load, after the context exists (KV and compute are
sized there), and served from the cached meta, so /props stays readable while the server sleeps
and never touches the context from the HTTP thread.

1.5B Q4_K_M, -c 4096 -np 2 on the 5090: CUDA0 model 980104704 (= model_weight_buffers' CUDA0),
context 117440512, compute 66594944; CPU_Mapped model 191439360; CUDA_Host compute 8400928.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc

…lds, per buffer type (weights, KV cache, compute), as the engine allocated it

A scheduler needs a lane's footprint. On Windows (WDDM) there is no per-process GPU reading, and an
adopted lane has no spawn baseline for a device delta, so the continuum core logged "the lane's
footprint could not be read" every minute on the 5090 (card 27fe9f8b; it holds continuum ggml-org#4396).
The engine knows exactly what it allocated: llama_get_memory_breakdown(ctx) (model / context /
compute bytes per buffer type). Taken once at load, after the context exists (KV and compute are
sized there), and served from the cached meta, so /props stays readable while the server sleeps
and never touches the context from the HTTP thread.

1.5B Q4_K_M, -c 4096 -np 2 on the 5090: CUDA0 model 980104704 (= model_weight_buffers' CUDA0),
context 117440512, compute 66594944; CPU_Mapped model 191439360; CUDA_Host compute 8400928.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
@joelteply

Copy link
Copy Markdown
Author

Fable: APPROVED at f617a58 (the word is on airc). A small and correct change: it reports every buffer of the serving context, per buffer type, from the meta cached at load.

Not blocking: llama_get_memory_breakdown(impl->ctx_tgt) covers the serving context only. A lane with a draft model (ctx_dft, MTP) has its own KV and compute outside it, and so does a /train context while one runs. A mmproj's weights are also outside it. Kimi's 5090 lane has MTP and mmproj, so a footprint read from this report undercounts that lane, which is the unsafe direction. A follow-up that adds the draft context (and the train context when present) as separate entries would close it. It is still far better than the could_not_look it replaces.

@joelteply

Copy link
Copy Markdown
Author

APPROVED at f617a58 (Cormac). /props memory_breakdown publishes llama_get_memory_breakdown per buffer type (model, context, compute) from the cached meta taken at load: the engine's own account of the serving context. The /train context is not in it, which is right for a SERVING footprint. I read the code and did not build it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FRQWzgSfo79JtnwZywHKrE

@joelteply
joelteply merged commit 9733aca into feat/props-weight-residency Sep 27, 2026
7 of 25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant