Repository navigation
llama: support MoE cache over multiple GPUs - #30112
Conversation
|
CI is failing on build: |
pardon my ignorance, how do I calculate the weight of experts only? |
|
@vbooka1 if you run with llama.cpp/src/llama-moe-cache.cpp Line 366 in 2de17c2 |
Would it not be better to have a comma seperated list to support heterogene setups? |
|
Or maybe the best would be to specify the cache size "per layer"? I.e. a single number that has the meaning of the "moe cache size for a layer". This way, each device will allocate moe cache based on how many layers it works on. |
|
Per layer leaves out an optimisation opportunity to share slots between layers which have the same expert types, some layers can use more experts than others. |
Is this something we know for sure? I would expect expert usage to be rather uniform across the different layers. |
|
Yes it increases the hit rate by 5% (it's already implemented in the implementation) |
|
So do we keep the CLI arg like this, or do we do it a list of sizes for each device? The latter would be more complicated in terms of UX. |
|
We should support it per device, one single value can mean it's that value per device. Comma separated values can be per device values. What do you think? |
|
There is also the option to have a single global cache size (i.e. not per-device) and use the existing |
|
Yeah that's also good, more intuitive I guess |
Keep the fork MoE expert cache (src/llama-moecache.*). Drop the upstream one (src/llama-moe-cache.*, PR ggml-org#29887 and ggml-org#30112) and its moe_cache_size plumbing, see scripts/fork-post-merge.sh. Keep the fork server checkpoint store and combine it with the upstream slot save/restore checkpoint appendix (PR ggml-org#26004): load_tgt / load_dft now return bool, an in-memory spec checkpoint aborts when a restore fails and a slot-file checkpoint falls back to full prompt re-processing. Assisted-by: pi (deepseek-flash)
|
Please consider supporting per-device / comma-separated values (like --moe-cache-mib 512,4096)! |
|
Does this patch apply to Qwen-Next family only? GLM5.3 Flash both PP and TG became slower, PP about 10% slower and TG became 200% slower, 33 -> 16 t/s |
…29887 and ggml-org#30112 merged ggml-org#29887 (the MoE expert cache) merged on 10-07 and ggml-org#30112 (the cache over multiple GPUs) on 10-08. On b11496 and later the old Inkling head refused in six files and unslothai#251, which carried its own copy of ggml-org#29887, refused in eight. Both branches now merge upstream master c811cb8. unslothai#251 keeps upstream's cache lines intact, so the merged tree holds both pins in full.
|
@Yozam-87 the way it is made now, from what I see, if you push the cache to the second GPU, the layers cached on that GPU will also get attention. If this is not what you want, you probably should be good with ts 1,0. |
On my setup model doesn't even load with ts 1,0, if I would be able to have second value - I would be able to use cache. I have 16gb and 8gb GPUs - qwen3.8-flash-next |
Nach den Verify-/Budget-Gates: MTP+Gather+Admission sauber (Acceptance 0.54-0.59), Budget-Sweep zeigt 6 GB Sweet Spot (~694 t/s pp), und der separate Upstream-Build (ggml-org#29887/ggml-org#30112, graph-level MoE-Cache) erreicht auf identischer Hardware nur ~80-100 t/s pp (Slot-Indirektion nur <=32 Tokens, Pipeline-Parallel deaktiviert bei aktivem Cache) vs. Fork-Gather ~694 t/s. Jeder Gather-Reject/OOM faellt atomar auf den Baseline-Upload zurueck; worst-case Gather (8 GB Budget, Admission-Churn) bleibt ~5x ueber Baseline.
Overview
Cont #29887, adding multi-GPU support. Benchmarks on 2x 4090s
Performance
2x RTX 4090 (layer split), Qwen3.8-Flash-Next Q4_0 (93.7 GiB, 65 GiB of experts), EPYC 7742 (16 threads, one NUMA node)
SPEED-Bench
qualitative, 1 sample per category (11 prompts),--osl 512, temperature 0, concurrency 1--fit--fit --moe-cache-mib 6544-cmoe --moe-cache-mib 15000--moe-cache-mibis per GPU. With--fit, CUDA0 holds 15 full layers, so only CUDA1 has host experts and a cache. With-cmoe, CUDA0 caches 25 layers and CUDA1 caches 23.Additional information
Requirements