Repository navigation
Conversation
Add --moe-cache-layers N|auto, a GPU cache for the experts of host-resident MoE weights. Small batches run from the store slots, bigger ones use the copy callback; --fit resolves auto by bisection and llama-bench gets the option. Assisted-by: deepseek v4.1 flash
|
Features ready for next pr, already tested: --moe-cache-mode move: Speculative prediction methods which i'm still testing to figure out performance wise but all work fine |
|
Nvm, I see that was not what you were doing (OP is confusing, you should probably reword it). |
̶I̶ ̶d̶o̶n̶t̶ ̶w̶a̶n̶t̶ ̶t̶h̶i̶s̶ ̶t̶o̶ ̶c̶o̶m̶e̶ ̶a̶c̶r̶o̶s̶s̶ ̶a̶s̶ ̶i̶n̶t̶e̶n̶t̶i̶o̶n̶a̶l̶l̶y̶ ̶d̶o̶i̶n̶g̶ ̶s̶o̶m̶e̶t̶h̶i̶n̶g̶ ̶w̶r̶o̶n̶g̶ ̶b̶u̶t̶ ̶I̶ ̶h̶a̶v̶e̶ ̶b̶e̶e̶n̶ ̶u̶n̶d̶e̶r̶ ̶t̶h̶e̶ ̶i̶m̶p̶r̶e̶s̶s̶i̶o̶n̶ ̶t̶h̶a̶t̶ ̶o̶p̶e̶n̶s̶o̶u̶r̶c̶e̶ ̶i̶s̶ ̶a̶l̶l̶ ̶a̶b̶o̶u̶t̶ ̶g̶e̶t̶t̶i̶n̶g̶ ̶t̶h̶e̶ ̶b̶e̶s̶t̶ ̶r̶e̶s̶u̶l̶t̶ ̶o̶u̶t̶ ̶o̶f̶ ̶e̶v̶e̶r̶y̶o̶n̶e̶'̶s̶ ̶e̶f̶f̶o̶r̶t̶ Nevermind I see your edit (ok i have) |
|
Just to provide a bit more context on the architecture here: my main goal was to solve the bottlenecking issues associated with standard LRU (by implementing heat-weighted eviction and hysteresis), but I wanted to make sure it didn't bloat the codebase. I managed to keep the footprint not that much worse in comparison Although I would really appreciate someone including am17an to have a look at it now as I've been trying to do this for months
|
Some backends (Vulkan) leave view tensors uninitialized when a no_alloc context is allocated with ggml_backend_alloc_ctx_tensors_from_buft(), so the cached weight view of the store had a null buffer and the scheduler aborted on a MUL_MAT_ID with host experts. Initialize any uninitialized bank view right after the allocation; a no-op elsewhere. Assisted-by: deepseek v4.1 flash
Overview
This is expert caching combining ideas from several different prs to combine into my interpretation of the best of all worlds in stability, performance, user friendliness, and maintainability
The diff is at the size i think where its still acceptable unlike my early attempts, and it has resolved all issues anyone has had
Full (extensive*) credits table
Attribution table:
--fitaccounting for the cache--moe-cache-layers Nmirroring--n-cpu-moe, host experts placed through tensor overridesgate_upexpert tensorsMADV_DONTNEEDrounded inward to page boundaries--n-cpu-moeMUL_MAT_ID, no shared dummy slotCUDA_Host, plain)The first attempt:
#26563
*Some of those are already merged. I still gave credits if I felt my ideas took from them
Original code to wackMall:
Performance
Generation speed (TG, tokens/s), single prompt, temperature 0, RTX 3070 Laptop (8 GB VRAM) + 32 GB RAM.
Stock = cache off; Cache on =
--moe-cache-layers auto.Usage
--moe-cache-layers auto OR N
N is how many MoE layers get a GPU cache for their experts (--moe-cache-layers 8 = the experts of 8 layers). auto sizes it from the free VRAM. Default is 0 = off, everything as stock. Only affects models whose experts stay in host memory.
Requirements