Repository navigation
Conversation
When system RAM is smaller than the routed-expert pool, page-cache LRU thrashes under per-token expert streaming. This pins the statistically hottest (layer, expert) slices inside the existing mmap so they cannot be evicted; only cold experts stream from disk. Env-gated and off by default (LLAMA_EXPERT_PIN_PROFILE, LLAMA_EXPERT_PIN_MB), Linux-only. Only helps when routing is skewed; on near-uniform routers (measured: Qwen3.8-Flash-Next, 98.6% of slots active in 1.3k tokens) no policy beats LRU - measure first with examples/moe-trace or similar. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Hi @aryan0078, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
RFC / experimental. Adds env-gated mlock() pinning of the statistically
hottest MoE expert slices inside the existing model mmap.
When system RAM is smaller than the routed-expert pool, page-cache LRU
thrashes: every token streams
n_expert_usedslices per layer and evictspages the next token needs. Pinning the hottest (layer, expert) slices makes
them ineligible for eviction, so only cold experts stream from disk.
LLAMA_EXPERT_PIN_PROFILE— csv oflayer,expert[,count], hottest first(producible with the moe-trace example, PR examples : add moe-trace to log MoE expert routing decisions #28544)
LLAMA_EXPERT_PIN_MB— pin budget (default 8192)Off by default, Linux-only, no behaviour change when the env var is unset.
Slices are addressed as
t->data + expert * t->nb[2]on the host buffer;device-resident expert tensors are skipped.
Honest status
The mechanism is verified (kernel reports the pinned bytes as Locked in
/proc/<pid>/smaps; partial pinning degrades gracefully whenRLIMIT_MEMLOCK is low), but I do not have a benchmark showing an
end-to-end win: the model this was built for (Qwen3.8-Flash-Next) turned out
to route near-uniformly (98.6% of slots active in ~1.3k tokens), where no
pinning policy can beat plain LRU — that finding is documented in a comment
at the pin site. On a skewed router with RAM < expert pool the expected gain
is real but currently theoretical.
Opening as a draft to ask: (a) is this direction acceptable at all, and
(b) if so, should the env vars be promoted to
llama_model_params/ CLI args?Happy to close if maintainers prefer to wait for a demonstrated win.
Requirements
an AI assistant (Claude Code). I reviewed the diff and verified pinning
behaviour on my machine (i7-12700, 30 GiB RAM, RX 6700).
🤖 Generated with Claude Code