Conversation
If you know that, why not split them into two PRs? |
|
Hello, @mxxm-t At least parallel loading works. Once a model is in RAM, it loads almost instantly. |
|
I have 1080 ti multi. Very interested in this PR. Results: Update Seems CUDA supported, but not switched ON for all GPUs? What to do? |
Try putting GGML_CUDA_HOST_STAGE=1 in front of your command. |
No it is not working on Vulkan. |
Enabled, but not sure that it actually work. At least I don't see more tokens. p.s. with full -fit-ctx it crashed, but when I change to fixed context with -c it work. Maybe issue not related to your code. |
|
I actually see significant drop in PP compared to vanilla on my 4x MI50 system with qwen3.5-120B-A10B. This with P2P enabled, cards connected via PCIe x4 (though some working only with x2 for some reason). There is also rx 7900xt in the system but not used for the test. PR Vanilla |
Try it without -ts or use it so its as equal as possible between GPU's. Also if you have different PCIe width used on cards try setting fastest connected GPU as main GPU. If you say you are using p2p then it should default to off and is not actually working. You can use -v flag then it tells you if it is actually on with similar logs to this: Also if you have any other models you can try those because I didn't get myself Qwen3.5 models working on that version of llama.cpp what my branch is based on. |
|
I see following in the logs: I tried all 5 gpus with no -ts - qwen3.5 is still behind vanilla, but on qwen3 it's comparable (but no real improvement). Maybe tomorrow I'll do some additional tests with P2P disabled. PR Vanilla |
|
Can you please submit PR with only quick model loading. I hope it can be already merged and you will be able to test and work on parallelism? |
|
Closing this PR and continue submitting smaller more focused PR's so its easier to review. Will leave branch in place for anyone that is interested, |
…hed ⏭️ + AtomicBot-ai#76 CPU Fusion ⏭️ 3 parallele Tiefen-Evals für Tier-3 Items: AtomicBot-ai#74 Vulkan Descriptor Indexing (Bindless) ❌ VERWORFEN: - Redundant mit AtomicBot-ai#85 Push Descriptors (✅ implementiert 2026-07-14) - Push Descriptors eliminieren dieselben CPU-Aufrufe - AtomicBot-ai#85-Benchmark auf Mars RADV: ±0.1-0.3% (Rauschen) - Bindless würde über Push-Descriptors hinaus <0.5% bringen - Workload-Mismatch: Bindless für draw-heavy Rendering, nicht Compute - Mars/Venus bandwidth-bound, nicht descriptor-bound - Aufwand revidiert: 2-4 → 3-5 Wochen (Shader-Rewrite aller .comp-Files) AtomicBot-ai#75 Non-blocking Pipeline Scheduling ⏭️ SPÄTER: - PR ggml-org#19922 closed (2026-03-03, unmerged, 4+ Mo stale) - Fork hat bereits Upstream-Pipeline-Parallelismus - Konflikt mit AtomicBot-ai#79 TP (✅+23-32% tg, split-mode-exklusiv) - NVIDIA ungetestet, PP-Regression auf 4x MI50 gemeldet - 2-GPU-Setup → geringer Bubble-Hebel - Aufwand revidiert: 3-4 → 4-6 Wochen AtomicBot-ai#76 CPU Backend Operator Fusion ⏭️ SPÄTER: - RMS_NORM+MUL Fusion bereits im Fork (PR ggml-org#22423 upstream-merged) - MoE Gated FFN riskant: PR ggml-org#20596 zeigt Regressionen auf Consumer-CPUs (M2: 0.98-1.00x, qwen3moe 30B: 0.75-0.96x bei t=2-4) - Nur auf 96-Core-EPYC konsistente Gains (1.04-1.08x) - Styx/Uranus haben Consumer-CPUs → wahrscheinlich Regression - Re-Eval wenn PR ggml-org#20596 gemerged mit Regression-Freiheit
This PR includes two changes/features (I know two changes in 1 PR is bad):
Backstory
I have a system with 10x Radeon MI50 (32 GB) GPUs. When I started digging into the llama.cpp codebase to improve MI50 performance, I noticed that a large portion of the test time was spent on model loading, not inference.
So I focused on improving multi‑GPU model loading to better saturate my NVMe drive(s) and get closer to full disk throughput.
Back to optimizing the speed… Old news is that you should try to fit the model on the least number of GPUs to get maximum performance, especially in layer-split mode.
Looking at GPU utilization during inference we can see that pretty much most of the time one GPU is doing work while other GPUs are sitting idle.
Idle time = waste of performance, and the more GPUs you use the more waste of potential performance you can end up with.
When I started to look into the codebase and added debug logging and timings for how much time each part takes, I found out that most of the time on prompt processing was spent on the CPU hitting blocking sync points between split submissions, which created pipeline bubbles (some GPUs idle) and prevented keeping multiple microbatches in flight.
This PR improves that by moving more of the coordination to GPU-side events/streams, so the CPU can keep preparing/submitting the next work instead of blocking, and the GPUs can self-synchronize with less idle time.
Testing / Feedback wanted
I know these changes are big and they change a lot of core logic, so it needs a lot of testing, but I think it's definitely worth checking out. I think this also opens up opportunities for further performance improvements.
We also need to figure out if the current implementation is fine as-is, or if this should be an optional toggle and by default keep the current/vanilla route for cross‑GPU execution.
So far this has mostly been tested on my system and a few other multi‑MI50 systems.
What I'm looking for now is more people with different builds / hardware combinations to test it, to see how it behaves and if there are any drawbacks or bugs.
Since I don't have multiple NVIDIA GPUs myself, I don't really know:
on NVIDIA GPUs. In theory there shouldn't be a difference, but I'd like real-world confirmation.
AI generated summary
1) Non‑blocking pipeline scheduling (event‑driven dependencies)
2) Queue throttle for stability under full async pipelining
3) Faster / safer cross‑GPU copies (direct P2P when available, host‑staged when needed)
4) Parallel multi‑GPU model loading (
-tl/--threads-load)5) Load-time async uploads toggle (CUDA + HIP) + HIP default safety
User-facing controls (flags / env vars)
CLI flags
-tl, --threads-load NN=1forces sequential loading4-hs, --host-stage {auto,0,1}GGML_CUDA_HOST_STAGE)auto(default): direct P2P if available, otherwise host‑staged0: force direct P2P1: force host‑stagedEnvironment variables
Cross‑GPU transfer policy / tuning
GGML_CUDA_HOST_STAGEunset: AUTO (direct if P2P available else staged)0: force directGGML_CUDA_HOST_STAGE_TILE_MB16GGML_CUDA_ACTXFER_GATEunset/non‑zero: enabled0: disabledGGML_CUDA_FORCE_XDEV_SYNC1: enabledLoad-time async uploads
LLAMA_ASYNC_UPLOADS0: disable async uploads during loadingunset:Notes
Benchmarks / Tests
GPT‑OSS 120B (MXFP4 MoE)
Baseline: scaling vs number of GPUs
This PR: scaling vs number of GPUs
This PR: different pp sizes
GPT‑OSS 20B (MXFP4 MoE)
This PR: different pp sizes (and tg sanity check)
Real-world test:
https://www.youtube.com/watch?v=xXVaWT8UoV0