Skip to content

llama : experimental mlock pinning of hot MoE expert slices (RFC) - #28545

Closed
aryan0078 wants to merge 1 commit into
ggml-org:masterfrom
aryan0078:expert-slice-pinning
Closed

aryan0078 wants to merge 1 commit into
ggml-org:masterfrom
aryan0078:expert-slice-pinning

Conversation

@aryan0078

Copy link
Copy Markdown

Overview

RFC / experimental. Adds env-gated mlock() pinning of the statistically
hottest MoE expert slices inside the existing model mmap.

When system RAM is smaller than the routed-expert pool, page-cache LRU
thrashes: every token streams n_expert_used slices per layer and evicts
pages the next token needs. Pinning the hottest (layer, expert) slices makes
them ineligible for eviction, so only cold experts stream from disk.

Off by default, Linux-only, no behaviour change when the env var is unset.
Slices are addressed as t->data + expert * t->nb[2] on the host buffer;
device-resident expert tensors are skipped.

Honest status

The mechanism is verified (kernel reports the pinned bytes as Locked in
/proc/<pid>/smaps; partial pinning degrades gracefully when
RLIMIT_MEMLOCK is low), but I do not have a benchmark showing an
end-to-end win: the model this was built for (Qwen3.8-Flash-Next) turned out
to route near-uniformly (98.6% of slots active in ~1.3k tokens), where no
pinning policy can beat plain LRU — that finding is documented in a comment
at the pin site. On a skewed router with RAM < expert pool the expected gain
is real but currently theoretical.

Opening as a draft to ask: (a) is this direction acceptable at all, and
(b) if so, should the env vars be promoted to llama_model_params / CLI args?
Happy to close if maintainers prefer to wait for a demonstrated win.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES — the code and this description were prepared with
    an AI assistant (Claude Code). I reviewed the diff and verified pinning
    behaviour on my machine (i7-12700, 30 GiB RAM, RX 6700).

🤖 Generated with Claude Code

When system RAM is smaller than the routed-expert pool, page-cache LRU
thrashes under per-token expert streaming. This pins the statistically
hottest (layer, expert) slices inside the existing mmap so they cannot be
evicted; only cold experts stream from disk. Env-gated and off by default
(LLAMA_EXPERT_PIN_PROFILE, LLAMA_EXPERT_PIN_MB), Linux-only.

Only helps when routing is skewed; on near-uniform routers (measured:
Qwen3.8-Flash-Next, 98.6% of slots active in 1.3k tokens) no policy beats
LRU - measure first with examples/moe-trace or similar.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 7, 2026

Copy link
Copy Markdown

Hi @aryan0078, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 3 open PRs.

  • AI-generated content: While code is allowed to be generated by AI, please write the PR description and commit messages on your own without the help of AI.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants