Skip to content

Feature Request: Just-in-time MoE expert streaming from storage #27562

Description

@TechRenamed

Prerequisites

  • I am running the latest code. Mention the version if possible as well.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new and useful enhancement to share.

Feature Description

Would llama.cpp consider implementing expert-aware streaming for MoE models, where only the experts selected by the router are fetched from storage for the current token/layer?
The idea would be to keep large MoE models file-backed, fetch only selected expert weight slices just in time, optionally cache frequently used experts in RAM, and overlap storage reads with CPU computation.
A project called Big MoE on Edge is already experimenting with this approach, including Android/CPU-only use. This could make MoE models substantially larger than available RAM practical on devices with fast NVMe/UFS storage.
It would be especially interesting as an upstream llama.cpp feature with options for expert caching and asynchronous I/O/prefetching.

Motivation

llama.cpp currently supports memory-mapped model loading, but sparse MoE models present an opportunity to reduce RAM requirements further by making storage access expert-aware.
For each token and MoE layer, only a subset of experts is selected by the router. Instead of relying entirely on normal OS page caching/mmap behavior, llama.cpp could optionally fetch only the weight slices belonging to the selected experts from storage just in time, while caching frequently used experts in RAM.
This could make MoE models larger than available physical RAM more practical on memory-constrained systems, particularly Android devices with fast UFS storage and PCs with fast NVMe SSDs.
An additional optimization could overlap reads for upcoming expert weights with computation on already-loaded weights, reducing the amount of time inference stalls waiting for storage.
A proof-of-concept project, “Big MoE on Edge,” appears to implement this approach using llama.cpp's evaluation callback, with an optional small hook for per-expert synchronization/overlap. It would be useful to have similar functionality supported upstream in llama.cpp.

Possible Implementation

Add an optional expert-streaming backend for MoE tensors:
Keep expert weights file-backed rather than requiring all experts to remain resident in RAM.
After routing, determine the experts selected for the current layer/token.
Fetch only the required expert weight slices.
Maintain a configurable RAM cache/LRU for frequently selected experts.
Optionally prefetch/overlap storage I/O with computation.
Fall back to the existing mmap behavior when expert streaming is disabled.
Ideally this would remain optional so existing llama.cpp behavior and performance are unchanged by default.
Reference implementation / prior work: https://github.com/Helldez/BigMoeOnEdge

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions