Prerequisites
Feature Description
By manually managing a small cache of experts in RAM and explicitly streaming weights from the SSD only when needed, we can run models significantly larger than system RAM with better performance, compared with existing implemenation.
Motivation
Currently, running massive Mixture of Experts (MoE) models (e.g., Mixtral, GLM-5.2) on systems with limited RAM relies on the operating system's virtual memory (mmap). Because the OS doesn't understand which experts the neural network will need, it constantly swaps data in and out randomly, leading to I/O thrashing and slow down generation speed.
Possible Implementation
There could be multiple phases:
- Only fetch the needed expert weights from the GGUF file directly into the cache.
- Async read an expert from the SSD, to improve the decode speed.
- Break the prompt computation into multiple waves. In the first wave, we only calculate the math for the experts currently sitting in the cache. While that math is running, the background threads load the experts needed for the next wave. We sum the results together at the end. It would increase the prefill speed.
Prerequisites
Feature Description
By manually managing a small cache of experts in RAM and explicitly streaming weights from the SSD only when needed, we can run models significantly larger than system RAM with better performance, compared with existing implemenation.
Motivation
Currently, running massive Mixture of Experts (MoE) models (e.g., Mixtral, GLM-5.2) on systems with limited RAM relies on the operating system's virtual memory (
mmap). Because the OS doesn't understand which experts the neural network will need, it constantly swaps data in and out randomly, leading to I/O thrashing and slow down generation speed.Possible Implementation
There could be multiple phases: