Skip to content

Feature Request: SSD streaming of MoE routed expert weights (run MoE models larger than system RAM) #25257

Description

@freedomljc

Prerequisites

  • I am running the latest code. Mention the version if possible as well.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new and useful enhancement to share.

Feature Description

By manually managing a small cache of experts in RAM and explicitly streaming weights from the SSD only when needed, we can run models significantly larger than system RAM with better performance, compared with existing implemenation.

Motivation

Currently, running massive Mixture of Experts (MoE) models (e.g., Mixtral, GLM-5.2) on systems with limited RAM relies on the operating system's virtual memory (mmap). Because the OS doesn't understand which experts the neural network will need, it constantly swaps data in and out randomly, leading to I/O thrashing and slow down generation speed.

Possible Implementation

There could be multiple phases:

  • Only fetch the needed expert weights from the GGUF file directly into the cache.
  • Async read an expert from the SSD, to improve the decode speed.
  • Break the prompt computation into multiple waves. In the first wave, we only calculate the math for the experts currently sitting in the cache. While that math is running, the background threads load the experts needed for the next wave. We sum the results together at the end. It would increase the prefill speed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions