Skip to content

LLama-wackMall-fusion expert cache. - #30059

Open
miltos22 wants to merge 2 commits into
ggml-org:masterfrom
miltos22:moe-cache
Open

miltos22 wants to merge 2 commits into
ggml-org:masterfrom
miltos22:moe-cache

Conversation

@miltos22

@miltos22 miltos22 commented Oct 6, 2026 •

Copy link
Copy Markdown

Overview

This is expert caching combining ideas from several different prs to combine into my interpretation of the best of all worlds in stability, performance, user friendliness, and maintainability
The diff is at the size i think where its still acceptable unlike my early attempts, and it has resolved all issues anyone has had

Full (extensive*) credits table

Attribution table:

What From
Running host-resident MoE experts on the GPU from a device cache: LRU with miss-only uploads, small-batch gate (32 tokens), separate banks per expert layout qvac-fabric, via #29887 (am17an)
Scheduler integration (routed ids read back and remapped, one upload shared by gate/up/down), cache disabled for the draft context, --fit accounting for the cache #29887 (am17an)
--moe-cache-layers N mirroring --n-cpu-moe, host experts placed through tensor overrides #15077, #11397 (slaren)
Staging uploads that copy only the used experts, merged into runs #15346 (slaren)
Fused gate_up expert tensors #20416 (ddh0)
Releasing only whole pages, MADV_DONTNEED rounded inward to page boundaries #26003 (pwilkin)
Eviction that weighs use frequency instead of plain LRU pashko-k (in #29887), #21609 (e1n00r)
Sizing the cache from free VRAM next to --n-cpu-moe #21614 (e1n00r)
Synchronizing before evicted experts are retired pwasiewi (in #27861)
Distinct ids per token in MUL_MAT_ID, no shared dummy slot Inovello (in #27861)
Hit/miss counters #24524 (leloch)
Big batches kept off the cache, lessons from staging-tensor residency #23170 (avifenesh), #24524 (leloch)
Move mode aimed at mmap, measured against plain mmap with the page cache pwilkin (in #28414)
The host buffer types move mode has to handle (mapped, pinned CUDA_Host, plain) #28223 (Inovello), #26659 (bluechiperic)
Releasing the host pages of cached experts rather than pinning hot ones #26414 (Ghimli), #28545 (aryan0078)
Routing traces showing experts are reused enough to cache #28544 (aryan0078)
The big-batch op offload path the staging planes plug into #13386 (hjc4869), #18535 (DocShotgun)

The first attempt:
#26563

*Some of those are already merged. I still gave credits if I felt my ideas took from them

Original code to wackMall:

  • Routing heatmap: per-expert usage tracked with decay, plus the heat log; admission and eviction governed by hysteresis, dwell, a sync period and a no-evict mode.
  • Heat-weighted eviction order for the device cache.
  • Move mode's release gate: host pages are released only once an expert has stayed cached long enough and sits clear of the coldest quarter, with rate-limited background read-back of released experts that turn cold.
  • Copy mode by default; move mode runs only on mmap and falls back to copy.
  • The --moe-cache-layers N|auto option, auto resolved through --fit.

Performance

Generation speed (TG, tokens/s), single prompt, temperature 0, RTX 3070 Laptop (8 GB VRAM) + 32 GB RAM.
Stock = cache off; Cache on = --moe-cache-layers auto.

Model Stock Cache on Speedup
LFM2.5-8B-A1B Q4_K_M 238.2 242.8 1.02×
Qwen3.6-35B-A3B IQ2_M 35.0 82.2 2.35×
Gemma-4-26B-A4B Q5_K_S 15.0 46.0 3.06×
Qwen3.6-35B-A3B Q6 13.0 24.7 1.90×
Qwen3.5-122B-A10B IQ2_M 5.4 13.7 2.56×

Usage

--moe-cache-layers auto OR N
N is how many MoE layers get a GPU cache for their experts (--moe-cache-layers 8 = the experts of 8 layers). auto sizes it from the free VRAM. Default is 0 = off, everything as stock. Only affects models whose experts stay in host memory.

Requirements

Add --moe-cache-layers N|auto, a GPU cache for the experts of host-resident
MoE weights. Small batches run from the store slots, bigger ones use the
copy callback; --fit resolves auto by bisection and llama-bench gets the
option.

Assisted-by: deepseek v4.1 flash
@miltos22

miltos22 commented Oct 6, 2026

Copy link
Copy Markdown
Author

Features ready for next pr, already tested:

--moe-cache-mode move:
Moving experts based on velocity in and out of ram to recover ram when they exist on the GPU for prolonged periods of time as well as pre fetching them when they are about to move out of gpu into ram. About 200 more lines.
llama-server -m model.gguf --moe-cache-layers N --moe-cache-mode move

Speculative prediction methods which i'm still testing to figure out performance wise but all work fine

@CISC

CISC commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Merging PRs like this is not OK.

Nvm, I see that was not what you were doing (OP is confusing, you should probably reword it).

@miltos22

miltos22 commented Oct 6, 2026 •

Copy link
Copy Markdown
Author

Merging PRs like this is not OK.

̶I̶ ̶d̶o̶n̶t̶ ̶w̶a̶n̶t̶ ̶t̶h̶i̶s̶ ̶t̶o̶ ̶c̶o̶m̶e̶ ̶a̶c̶r̶o̶s̶s̶ ̶a̶s̶ ̶i̶n̶t̶e̶n̶t̶i̶o̶n̶a̶l̶l̶y̶ ̶d̶o̶i̶n̶g̶ ̶s̶o̶m̶e̶t̶h̶i̶n̶g̶ ̶w̶r̶o̶n̶g̶ ̶b̶u̶t̶ ̶I̶ ̶h̶a̶v̶e̶ ̶b̶e̶e̶n̶ ̶u̶n̶d̶e̶r̶ ̶t̶h̶e̶ ̶i̶m̶p̶r̶e̶s̶s̶i̶o̶n̶ ̶t̶h̶a̶t̶ ̶o̶p̶e̶n̶s̶o̶u̶r̶c̶e̶ ̶i̶s̶ ̶a̶l̶l̶ ̶a̶b̶o̶u̶t̶ ̶g̶e̶t̶t̶i̶n̶g̶ ̶t̶h̶e̶ ̶b̶e̶s̶t̶ ̶r̶e̶s̶u̶l̶t̶ ̶o̶u̶t̶ ̶o̶f̶ ̶e̶v̶e̶r̶y̶o̶n̶e̶'̶s̶ ̶e̶f̶f̶o̶r̶t̶

Nevermind I see your edit (ok i have)

@JohannesGaessler

Copy link
Copy Markdown
Contributor

As you already wrote yourself in the other PR #29887 : @am17an has already started working on this and I would go with his PR unless he himself is of a different opinion.

@github-actions github-actions Bot added documentation Improvements or additions to documentation testing Everything test related server ggml changes relating to the ggml tensor library for machine learning labels Oct 6, 2026
@miltos22

miltos22 commented Oct 6, 2026 •

Copy link
Copy Markdown
Author

Just to provide a bit more context on the architecture here: my main goal was to solve the bottlenecking issues associated with standard LRU (by implementing heat-weighted eviction and hysteresis), but I wanted to make sure it didn't bloat the codebase. I managed to keep the footprint not that much worse in comparison

Although I would really appreciate someone including am17an to have a look at it now as I've been trying to do this for months
Edit: Hysteresis is planned for the next pr which includes move mode

As you already wrote yourself in the other PR #29887 : @am17an has already started working on this and I would go with his PR unless he himself is of a different opinion.

Some backends (Vulkan) leave view tensors uninitialized when a no_alloc
context is allocated with ggml_backend_alloc_ctx_tensors_from_buft(), so
the cached weight view of the store had a null buffer and the scheduler
aborted on a MUL_MAT_ID with host experts. Initialize any uninitialized
bank view right after the allocation; a no-op elsewhere.

Assisted-by: deepseek v4.1 flash
@miltos22 miltos22 changed the title LLama-wackMall-fusion. LLama-wackMall-fusion expert cache. Oct 7, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning server testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants