Skip to content

llama: support MoE cache over multiple GPUs - #30112

Merged
am17an merged 5 commits into
masterfrom
aman/moe-cache-multigpu
Oct 8, 2026
Merged

am17an merged 5 commits into
masterfrom
aman/moe-cache-multigpu

Conversation

@am17an

@am17an am17an commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Cont #29887, adding multi-GPU support. Benchmarks on 2x 4090s

Performance

2x RTX 4090 (layer split), Qwen3.8-Flash-Next Q4_0 (93.7 GiB, 65 GiB of experts), EPYC 7742 (16 threads, one NUMA node)
SPEED-Bench qualitative, 1 sample per category (11 prompts), --osl 512, temperature 0, concurrency 1

config decode t/s vs master prompt t/s vs master cache hit rate (decode)
master --fit 34.78 1.00x 440.7 1.00x -
--fit --moe-cache-mib 6544 60.38 1.74x 402.0 0.91x 85.9%
-cmoe --moe-cache-mib 15000 64.39 1.85x 321.5 0.73x 95.8%

--moe-cache-mib is per GPU. With --fit, CUDA0 holds 15 full layers, so only CUDA1 has host experts and a cache. With -cmoe, CUDA0 caches 25 layers and CUDA1 caches 23.

Additional information

Requirements

@am17an
am17an requested review from CISC and ggerganov as code owners October 7, 2026 18:50
@github-actions github-actions Bot added the testing Everything test related label Oct 7, 2026
@pwilkin

pwilkin commented Oct 7, 2026

Copy link
Copy Markdown
Member

CI is failing on build:

/__w/llama.cpp/llama.cpp/src/llama-moe-cache.cpp: In constructor 'llama_moe_cache::impl::impl(const llama_model&, const std::vector<ggml_backend*>&, const std::vector<ggml_backend_buffer_type*>&, size_t)':
/__w/llama.cpp/llama.cpp/src/llama-moe-cache.cpp:225:34: error: missing initializer for member 'llama_moe_cache::impl::device::ctx' [-Werror=missing-field-initializers]
  225 |                 devices.push_back({ backends[i], bufts[i] });
      |                 ~~~~~~~~~~~~~~~~~^~~~~~~~~~~~~~~~~~~~~~~~~~~
/__w/llama.cpp/llama.cpp/src/llama-moe-cache.cpp:225:34: error: missing initializer for member 'llama_moe_cache::impl::device::buf' [-Werror=missing-field-initializers]

@vbooka1

vbooka1 commented Oct 7, 2026

Copy link
Copy Markdown

93.7 GiB, 65 GiB of experts

pardon my ignorance, how do I calculate the weight of experts only?

@am17an am17an added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Oct 8, 2026
@am17an

am17an commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

@vbooka1 if you run with --moe-cache-mib there should be log line with the amount of experts added.

LLAMA_LOG_INFO("%s: %10s MoE cache size = %8.2f MiB for %.2f MiB of host experts\n", __func__,

@ggerganov
ggerganov requested review from a team and ngxson as code owners October 8, 2026 06:52
@Green-Sky

Copy link
Copy Markdown
Collaborator

per device

Would it not be better to have a comma seperated list to support heterogene setups?

@github-actions github-actions Bot added documentation Improvements or additions to documentation server labels Oct 8, 2026
@ggerganov

ggerganov commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Or maybe the best would be to specify the cache size "per layer"? I.e. a single number that has the meaning of the "moe cache size for a layer". This way, each device will allocate moe cache based on how many layers it works on.

@am17an

am17an commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Per layer leaves out an optimisation opportunity to share slots between layers which have the same expert types, some layers can use more experts than others.

@ggerganov

Copy link
Copy Markdown
Member

some layers can use more experts than others.

Is this something we know for sure? I would expect expert usage to be rather uniform across the different layers.

@am17an

am17an commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Yes it increases the hit rate by 5% (it's already implemented in the implementation)

@ggerganov

Copy link
Copy Markdown
Member

So do we keep the CLI arg like this, or do we do it a list of sizes for each device? The latter would be more complicated in terms of UX.

@am17an

am17an commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

We should support it per device, one single value can mean it's that value per device. Comma separated values can be per device values. What do you think?

@ggerganov

Copy link
Copy Markdown
Member

There is also the option to have a single global cache size (i.e. not per-device) and use the existing --tensor-split to infer how to split the cache among the devices. Would be inline with the rest "per-device" splitting logic.

@am17an

am17an commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Yeah that's also good, more intuitive I guess

@am17an
am17an merged commit c811cb8 into master Oct 8, 2026
17 of 18 checks passed
feal87 added a commit to feal87/myllama.cpp that referenced this pull request Oct 8, 2026
Keep the fork MoE expert cache (src/llama-moecache.*). Drop the upstream one (src/llama-moe-cache.*, PR ggml-org#29887 and ggml-org#30112) and its moe_cache_size plumbing, see scripts/fork-post-merge.sh.

Keep the fork server checkpoint store and combine it with the upstream slot save/restore checkpoint appendix (PR ggml-org#26004): load_tgt / load_dft now return bool, an in-memory spec checkpoint aborts when a restore fails and a slot-file checkpoint falls back to full prompt re-processing.

Assisted-by: pi (deepseek-flash)
@Yozam-87

Yozam-87 commented Oct 8, 2026

Copy link
Copy Markdown

Please consider supporting per-device / comma-separated values (like --moe-cache-mib 512,4096)!
On asymmetric multi-GPU setups (e.g. a modern fast GPU for Flash Attention with -ts 1,0 and an older secondary GPU with lots of free VRAM), coupling the cache to --tensor-split makes the MoE cache unusable on the secondary card. With -ts 1,0, the secondary card gets 0 MB cache, even though it has gigabytes of free VRAM that could easily cache CPU experts without taking on attention layers.

@vbooka1

vbooka1 commented Oct 8, 2026 •

Copy link
Copy Markdown

Does this patch apply to Qwen-Next family only? GLM5.3 Flash both PP and TG became slower, PP about 10% slower and TG became 200% slower, 33 -> 16 t/s

0.33.942.823 I load_tensors:        CUDA0 model buffer size = 91667.78 MiB
0.33.942.824 I load_tensors:        CUDA1 model buffer size = 81986.62 MiB
0.33.942.825 I load_tensors:    CUDA_Host model buffer size = 104650.31 MiB
...
0.55.262.083 I impl:      CUDA1 MoE cache size =  9990.75 MiB for 99648.00 MiB of host experts
...
5.29.027.169 I llama_moe_cache: ubatch <= 8: hits = 50785, misses = 65842, hit rate = 43.54%, uploaded = 1423833.25 MiB
5.29.027.175 I llama_moe_cache: ubatch  > 8: hits = 120, misses = 1619, hit rate = 6.90%, uploaded = 35010.88 MiB

EmeraldBitTwizzler pushed a commit to EmeraldBitTwizzler/llama.cpp that referenced this pull request Oct 9, 2026
…29887 and ggml-org#30112 merged

ggml-org#29887 (the MoE expert cache) merged on 10-07 and ggml-org#30112 (the cache over
multiple GPUs) on 10-08. On b11496 and later the old Inkling head refused in six
files and unslothai#251, which carried its own copy of ggml-org#29887, refused in eight.

Both branches now merge upstream master c811cb8. unslothai#251 keeps upstream's cache
lines intact, so the merged tree holds both pins in full.
@sasa7812

sasa7812 commented Oct 9, 2026

Copy link
Copy Markdown

@Yozam-87 the way it is made now, from what I see, if you push the cache to the second GPU, the layers cached on that GPU will also get attention. If this is not what you want, you probably should be good with ts 1,0.

@BriNoB

BriNoB commented Oct 9, 2026

Copy link
Copy Markdown

@Yozam-87 the way it is made now, from what I see, if you push the cache to the second GPU, the layers cached on that GPU will also get attention. If this is not what you want, you probably should be good with ts 1,0.

On my setup model doesn't even load with ts 1,0, if I would be able to have second value - I would be able to use cache. I have 16gb and 8gb GPUs - qwen3.8-flash-next

fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Oct 11, 2026
Nach den Verify-/Budget-Gates: MTP+Gather+Admission sauber (Acceptance
0.54-0.59), Budget-Sweep zeigt 6 GB Sweet Spot (~694 t/s pp), und der
separate Upstream-Build (ggml-org#29887/ggml-org#30112, graph-level MoE-Cache) erreicht
auf identischer Hardware nur ~80-100 t/s pp (Slot-Indirektion nur <=32
Tokens, Pipeline-Parallel deaktiviert bei aktivem Cache) vs. Fork-Gather
~694 t/s. Jeder Gather-Reject/OOM faellt atomar auf den Baseline-Upload
zurueck; worst-case Gather (8 GB Budget, Admission-Churn) bleibt ~5x
ueber Baseline.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. server testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants