Skip to content

llama/ggml: multi-GPU pipeline parallelism (xdev host staging) + faster model loading - #19922

Closed
mxxm-t wants to merge 1 commit into
ggml-org:masterfrom
mxxm-t:pipeline-parallelism
Closed

mxxm-t wants to merge 1 commit into
ggml-org:masterfrom
mxxm-t:pipeline-parallelism

Conversation

@mxxm-t

@mxxm-t mxxm-t commented Feb 26, 2026

Copy link
Copy Markdown

This PR includes two changes/features (I know two changes in 1 PR is bad):

  1. Parallel multi‑GPU model loading, improving load time by using multiple loading threads (limited by '-tl').
  2. An improvement/change to cross‑GPU transfers used by multi‑GPU execution.

Backstory

I have a system with 10x Radeon MI50 (32 GB) GPUs. When I started digging into the llama.cpp codebase to improve MI50 performance, I noticed that a large portion of the test time was spent on model loading, not inference.
So I focused on improving multi‑GPU model loading to better saturate my NVMe drive(s) and get closer to full disk throughput.

Back to optimizing the speed… Old news is that you should try to fit the model on the least number of GPUs to get maximum performance, especially in layer-split mode.

Looking at GPU utilization during inference we can see that pretty much most of the time one GPU is doing work while other GPUs are sitting idle.
Idle time = waste of performance, and the more GPUs you use the more waste of potential performance you can end up with.

When I started to look into the codebase and added debug logging and timings for how much time each part takes, I found out that most of the time on prompt processing was spent on the CPU hitting blocking sync points between split submissions, which created pipeline bubbles (some GPUs idle) and prevented keeping multiple microbatches in flight.

This PR improves that by moving more of the coordination to GPU-side events/streams, so the CPU can keep preparing/submitting the next work instead of blocking, and the GPUs can self-synchronize with less idle time.

Testing / Feedback wanted

I know these changes are big and they change a lot of core logic, so it needs a lot of testing, but I think it's definitely worth checking out. I think this also opens up opportunities for further performance improvements.

We also need to figure out if the current implementation is fine as-is, or if this should be an optional toggle and by default keep the current/vanilla route for cross‑GPU execution.

So far this has mostly been tested on my system and a few other multi‑MI50 systems.

What I'm looking for now is more people with different builds / hardware combinations to test it, to see how it behaves and if there are any drawbacks or bugs.

Since I don't have multiple NVIDIA GPUs myself, I don't really know:

  1. if it works
  2. how well it works

on NVIDIA GPUs. In theory there shouldn't be a difference, but I'd like real-world confirmation.

AI generated summary

1) Non‑blocking pipeline scheduling (event‑driven dependencies)

  • Improves:
    • Overlap between CPU submission and GPU execution
    • Less time waiting between split submissions
  • How:
    • Uses streams/events for cross‑split dependencies so GPUs self‑synchronize while the CPU prepares the next microbatch

2) Queue throttle for stability under full async pipelining

  • Improves:
    • Steadier latency and avoids the CPU getting too far ahead of GPU work
  • How:
    • Adds a lightweight throttle to keep the in‑flight queue bounded

3) Faster / safer cross‑GPU copies (direct P2P when available, host‑staged when needed)

  • Improves:
    • Robustness/perf on systems where P2P is unavailable/slow or device pairs differ
  • How:
    • AUTO selects direct P2P vs host‑staged tiled copies per device‑pair
    • Optional gating to avoid producer/consumer races on the host‑staged path

4) Parallel multi‑GPU model loading (-tl/--threads-load)

  • Improves:
    • Model load time when multiple GPU contexts are used
  • How:
    • Loads tensors for multiple GPU contexts in parallel with a configurable thread limit

5) Load-time async uploads toggle (CUDA + HIP) + HIP default safety

  • Improves:
    • Controlled enable/disable of async upload behavior during loading
  • Note (ROCm/HIP):
    • Async uploads are disabled by default because enabling them can severely reduce multi‑GPU inference performance on some ROCm setups (override via env var)

User-facing controls (flags / env vars)

CLI flags

  1. -tl, --threads-load N
  • Controls parallel model loading worker limit
  • N=1 forces sequential loading
  • Default: 4
  1. -hs, --host-stage {auto,0,1}
  • Controls cross‑GPU transfer mode (sets GGML_CUDA_HOST_STAGE)
  • auto (default): direct P2P if available, otherwise host‑staged
  • 0: force direct P2P
  • 1: force host‑staged

Environment variables

Cross‑GPU transfer policy / tuning

  1. GGML_CUDA_HOST_STAGE
  • unset: AUTO (direct if P2P available else staged)
  • 0: force direct
  • non‑zero: force staged
  1. GGML_CUDA_HOST_STAGE_TILE_MB
  • Host‑staged tile size in MiB
  • Default: 16
  1. GGML_CUDA_ACTXFER_GATE
  • Producer/consumer gating for host‑staged transfers
  • unset/non‑zero: enabled
  • 0: disabled
  1. GGML_CUDA_FORCE_XDEV_SYNC
  • Debug: force CPU-side cross-device sync
  • 1: enabled
  • Default: auto-probe

Load-time async uploads

  1. LLAMA_ASYNC_UPLOADS
  • 0: disable async uploads during loading
  • non‑zero: enable async uploads during loading (still requires backend support)
  • unset:
    • HIP: defaults OFF
    • CUDA: unchanged default behavior (enabled if supported)

Notes

  • The host‑staged path is designed to be data-safe under async execution using explicit ordering/gating and per device‑pair coordination.
  • HIP vs CUDA: behavior is aligned wherever possible; HIP keeps the conservative default for load‑time async uploads due to observed ROCm behavior.
Benchmarks / Tests

GPT‑OSS 120B (MXFP4 MoE)

Baseline: scaling vs number of GPUs

| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1  |   1 |         pp16384 |        890.38 ± 0.56 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1  |   1 |          tg1024 |         82.44 ± 0.38 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3 |   1 |         pp16384 |        639.35 ± 0.95 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3 |   1 |          tg1024 |         76.84 ± 0.24 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5 |   1 |         pp16384 |        583.18 ± 0.75 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5 |   1 |          tg1024 |         73.43 ± 0.04 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5/ROCm6/ROCm7 |   1 |         pp16384 |        549.62 ± 0.53 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5/ROCm6/ROCm7 |   1 |          tg1024 |         70.50 ± 0.21 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5/ROCm6/ROCm7/ROCm8/ROCm9/ROCm9 |   1 |         pp16384 |        577.47 ± 0.71 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5/ROCm6/ROCm7/ROCm8/ROCm9/ROCm9 |   1 |          tg1024 |         67.30 ± 0.07 |

This PR: scaling vs number of GPUs

| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1  |   1 |         pp16384 |        907.38 ± 0.65 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1  |   1 |          tg1024 |         84.35 ± 0.05 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3 |   1 |         pp16384 |       1409.88 ± 1.82 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3 |   1 |          tg1024 |         83.44 ± 0.01 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5 |   1 |         pp16384 |       1799.02 ± 3.41 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5 |   1 |          tg1024 |         81.98 ± 0.03 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5/ROCm6/ROCm7 |   1 |         pp16384 |       2223.13 ± 2.23 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5/ROCm6/ROCm7 |   1 |          tg1024 |         80.68 ± 0.05 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5/ROCm6/ROCm7/ROCm8/ROCm9/ROCm9 |   1 |         pp16384 |       2465.36 ± 4.83 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 | ROCm0/ROCm1/ROCm2/ROCm3/ROCm4/ROCm5/ROCm6/ROCm7/ROCm8/ROCm9/ROCm9 |   1 |          tg1024 |         78.83 ± 0.04 |

This PR: different pp sizes

| model                          |       size |     params | backend    | ngl | n_ubatch | fa | dio |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | -: | --: | --------------: | -------------------: |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 |   1 |          pp1024 |        614.58 ± 3.16 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 |   1 |          pp2048 |       1058.78 ± 4.42 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 |   1 |          pp4096 |       1630.90 ± 2.39 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 |   1 |          pp8192 |       2166.38 ± 3.17 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 |   1 |         pp16384 |       2479.42 ± 2.20 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 |   1 |         pp32768 |       2435.75 ± 1.57 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 |   1 |         pp65538 |       2047.69 ± 0.96 |
| gpt-oss 120B MXFP4 MoE         |  59.02 GiB |   116.83 B | ROCm       |  99 |     1024 |  1 |   1 |        pp131072 |       1484.23 ± 0.90 |

GPT‑OSS 20B (MXFP4 MoE)

This PR: different pp sizes (and tg sanity check)

model                          |       size |     params | backend    | ngl | n_ubatch | fa | dio |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | -: | --: | --------------: | -------------------: |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | ROCm       |  99 |     1024 |  1 |   1 |          pp1024 |       1235.14 ± 7.73 |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | ROCm       |  99 |     1024 |  1 |   1 |          pp2048 |       2097.51 ± 8.46 |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | ROCm       |  99 |     1024 |  1 |   1 |          pp4096 |       3174.59 ± 2.63 |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | ROCm       |  99 |     1024 |  1 |   1 |          pp8192 |       4163.83 ± 6.84 |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | ROCm       |  99 |     1024 |  1 |   1 |         pp16384 |       4606.63 ± 2.96 |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | ROCm       |  99 |     1024 |  1 |   1 |         pp32768 |       4345.74 ± 1.52 |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | ROCm       |  99 |     1024 |  1 |   1 |         pp65538 |       3511.52 ± 3.79 |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | ROCm       |  99 |     1024 |  1 |   1 |        pp131072 |       2442.09 ± 0.27 |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | ROCm       |  99 |     1024 |  1 |   1 |          tg2048 |        112.05 ± 0.15 |

Real-world test:

https://www.youtube.com/watch?v=xXVaWT8UoV0

@0cc4m

0cc4m commented Feb 26, 2026

Copy link
Copy Markdown
Contributor

This PR includes two changes/features (I know two changes in 1 PR is bad)

If you know that, why not split them into two PRs?

@github-actions github-actions Bot added Nvidia GPU Issues specific to Nvidia GPUs ggml changes relating to the ggml tensor library for machine learning labels Feb 26, 2026
@savvadesogle

Copy link
Copy Markdown

Hello, @mxxm-t
Am I correct in understanding that this doesn't work with the Vulkan backend?

At least parallel loading works. Once a model is in RAM, it loads almost instantly.

@bchtrue

bchtrue commented Feb 27, 2026 •

Copy link
Copy Markdown

I have 1080 ti multi. Very interested in this PR.
I will try to test.
Thanks

Results:
It load really fast!! Thanks :-)
About performance improvement, don't see any currently. maybe need to tweek settings? Is it only for ROCm, not for CUDA? @mxxm-t Please explain. Is it hard to add for CUDA (nvidia) cards?

Update

pair=0<->1 host_stage=OFF (mode=AUTO, p2p=1, GGML_CUDA_HOST_STAGE=<unset>)
ggml_cuda_xdev: pair=1<->2 host_stage=OFF (mode=AUTO, p2p=1, GGML_CUDA_HOST_STAGE=<unset>)
ggml_cuda_xdev: pair=0<->2 host_stage=OFF (mode=AUTO, p2p=1, GGML_CUDA_HOST_STAGE=<unset>)
ggml_cuda_xdev: pair=2<->3 host_stage=OFF (mode=AUTO, p2p=1, GGML_CUDA_HOST_STAGE=<unset>)
ggml_cuda_xdev: pair=0<->3 host_stage=OFF (mode=AUTO, p2p=1, GGML_CUDA_HOST_STAGE=<unset>)
ggml_cuda_xdev: pair=3<->4 host_stage=ON (mode=AUTO, p2p=0, GGML_CUDA_HOST_STAGE=<unset>)
ggml_cuda_xdev_wait: cross-device cudaStreamWaitEvent supported -> using GPU-side waits (non-blocking)
ggml_cuda_xdev: pair=0<->4 host_stage=ON (mode=AUTO, p2p=0, GGML_CUDA_HOST_STAGE=<unset>)
ggml_cuda_xdev: pair=4<->5 host_stage=OFF (mode=AUTO, p2p=1, GGML_CUDA_HOST_STAGE=<unset>)

Seems CUDA supported, but not switched ON for all GPUs? What to do?

@mxxm-t

mxxm-t commented Feb 27, 2026

Copy link
Copy Markdown
Author

Seems CUDA supported, but not switched ON for all GPUs? What to do?

Try putting GGML_CUDA_HOST_STAGE=1 in front of your command.

@mxxm-t

mxxm-t commented Feb 27, 2026

Copy link
Copy Markdown
Author

Hello, @mxxm-t Am I correct in understanding that this doesn't work with the Vulkan backend?

At least parallel loading works. Once a model is in RAM, it loads almost instantly.

No it is not working on Vulkan.

@bchtrue

bchtrue commented Feb 27, 2026 •

Copy link
Copy Markdown

@mxxm-t

common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
�[0mggml_cuda_xdev: pair=0<->1 host_stage=ON (mode=FORCE_STAGED, p2p=1, GGML_CUDA_HOST_STAGE=1)
ggml_cuda_xdev_wait: cross-device cudaStreamWaitEvent supported -> using GPU-side waits (non-blocking)
ggml_cuda_xdev: pair=1<->2 host_stage=ON (mode=FORCE_STAGED, p2p=1, GGML_CUDA_HOST_STAGE=1)
ggml_cuda_xdev: pair=0<->2 host_stage=ON (mode=FORCE_STAGED, p2p=1, GGML_CUDA_HOST_STAGE=1)
ggml_cuda_xdev: pair=2<->3 host_stage=ON (mode=FORCE_STAGED, p2p=1, GGML_CUDA_HOST_STAGE=1)
ggml_cuda_xdev: pair=0<->3 host_stage=ON (mode=FORCE_STAGED, p2p=1, GGML_CUDA_HOST_STAGE=1)
ggml_cuda_xdev: pair=3<->4 host_stage=ON (mode=FORCE_STAGED, p2p=0, GGML_CUDA_HOST_STAGE=1)
ggml_cuda_xdev: pair=0<->4 host_stage=ON (mode=FORCE_STAGED, p2p=0, GGML_CUDA_HOST_STAGE=1)
ggml_cuda_xdev: pair=4<->5 host_stage=ON (mode=FORCE_STAGED, p2p=1, GGML_CUDA_HOST_STAGE=1)
ggml_cuda_xdev: pair=0<->5 host_stage=ON (mode=FORCE_STAGED, p2p=0, GGML_CUDA_HOST_STAGE=1)
srv    load_model: initializing slots, n_slots = 1
common_speculative_is_compat: the target context does not support partial sequence removal
�[0msrv    load_model: speculative decoding not supported by this context

Enabled, but not sure that it actually work. At least I don't see more tokens.
Maybe this is because of current model? "speculative decoding not supported by this context".

p.s. with full -fit-ctx it crashed, but when I change to fixed context with -c it work. Maybe issue not related to your code.

@Andryusz

Copy link
Copy Markdown

I actually see significant drop in PP compared to vanilla on my 4x MI50 system with qwen3.5-120B-A10B. This with P2P enabled, cards connected via PCIe x4 (though some working only with x2 for some reason). There is also rx 7900xt in the system but not used for the test.

PR

./llama-bench -m Qwen3.5-122B-A10B-Q4_1-00001-of-00003.gguf -ts 0/32/32/32/32 --mmap 0 -fa 1 -ngl 99 -r 1 -p 16384
ggml_cuda_init: found 5 ROCm devices:
  Device 0: Radeon RX 7900 XT, gfx1100 (0x1100), VMM: no, Wave Size: 32
  Device 1: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 2: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 3: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 4: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
| model                          |       size |     params | backend    | ngl | fa | ts           |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -: | ------------ | --------------: | -------------------: |
| qwen35moe 80B.A3B Q4_1         |  71.40 GiB |   122.11 B | ROCm       |  99 |  1 | 0.00/32.00/32.00/32.00/32.00 |         pp16384 |        434.03 ± 0.00 |
| qwen35moe 80B.A3B Q4_1         |  71.40 GiB |   122.11 B | ROCm       |  99 |  1 | 0.00/32.00/32.00/32.00/32.00 |           tg128 |         32.49 ± 0.00 |

Vanilla

./llama-bench -m Qwen3.5-122B-A10B-Q4_1-00001-of-00003.gguf -ts 0/32/32/32/32 --mmap 0 -fa 1 -ngl 99 -r 1 -p 16384
ggml_cuda_init: found 5 ROCm devices:
  Device 0: Radeon RX 7900 XT, gfx1100 (0x1100), VMM: no, Wave Size: 32
  Device 1: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 2: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 3: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 4: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
| model                          |       size |     params | backend    | ngl | fa | ts           |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -: | ------------ | --------------: | -------------------: |
| qwen35moe 80B.A3B Q4_1         |  71.40 GiB |   122.11 B | ROCm       |  99 |  1 | 0.00/32.00/32.00/32.00/32.00 |         pp16384 |        733.64 ± 0.00 |
| qwen35moe 80B.A3B Q4_1         |  71.40 GiB |   122.11 B | ROCm       |  99 |  1 | 0.00/32.00/32.00/32.00/32.00 |           tg128 |         33.03 ± 0.00 |

@mxxm-t

mxxm-t commented Feb 27, 2026

Copy link
Copy Markdown
Author

I actually see significant drop in PP compared to vanilla on my 4x MI50 system with qwen3.5-120B-A10B. This with P2P enabled, cards connected via PCIe x4 (though some working only with x2 for some reason). There is also rx 7900xt in the system but not used for the test.

PR

./llama-bench -m Qwen3.5-122B-A10B-Q4_1-00001-of-00003.gguf -ts 0/32/32/32/32 --mmap 0 -fa 1 -ngl 99 -r 1 -p 16384
ggml_cuda_init: found 5 ROCm devices:
  Device 0: Radeon RX 7900 XT, gfx1100 (0x1100), VMM: no, Wave Size: 32
  Device 1: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 2: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 3: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 4: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
| model                          |       size |     params | backend    | ngl | fa | ts           |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -: | ------------ | --------------: | -------------------: |
| qwen35moe 80B.A3B Q4_1         |  71.40 GiB |   122.11 B | ROCm       |  99 |  1 | 0.00/32.00/32.00/32.00/32.00 |         pp16384 |        434.03 ± 0.00 |
| qwen35moe 80B.A3B Q4_1         |  71.40 GiB |   122.11 B | ROCm       |  99 |  1 | 0.00/32.00/32.00/32.00/32.00 |           tg128 |         32.49 ± 0.00 |

Vanilla

./llama-bench -m Qwen3.5-122B-A10B-Q4_1-00001-of-00003.gguf -ts 0/32/32/32/32 --mmap 0 -fa 1 -ngl 99 -r 1 -p 16384
ggml_cuda_init: found 5 ROCm devices:
  Device 0: Radeon RX 7900 XT, gfx1100 (0x1100), VMM: no, Wave Size: 32
  Device 1: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 2: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 3: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 4: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
| model                          |       size |     params | backend    | ngl | fa | ts           |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -: | ------------ | --------------: | -------------------: |
| qwen35moe 80B.A3B Q4_1         |  71.40 GiB |   122.11 B | ROCm       |  99 |  1 | 0.00/32.00/32.00/32.00/32.00 |         pp16384 |        733.64 ± 0.00 |
| qwen35moe 80B.A3B Q4_1         |  71.40 GiB |   122.11 B | ROCm       |  99 |  1 | 0.00/32.00/32.00/32.00/32.00 |           tg128 |         33.03 ± 0.00 |

Try it without -ts or use it so its as equal as possible between GPU's. Also if you have different PCIe width used on cards try setting fastest connected GPU as main GPU. If you say you are using p2p then it should default to off and is not actually working. You can use -v flag then it tells you if it is actually on with similar logs to this:
ggml_cuda_xdev: pair=1<->2 host_stage=ON (mode=FORCE_STAGED, p2p=1, GGML_CUDA_HOST_STAGE=1)

Also if you have any other models you can try those because I didn't get myself Qwen3.5 models working on that version of llama.cpp what my branch is based on.

@Andryusz

Copy link
Copy Markdown

I see following in the logs:

ggml_cuda_xdev: pair=0<->1 host_stage=OFF (mode=AUTO, p2p=1, GGML_CUDA_HOST_STAGE=<unset>)

I tried all 5 gpus with no -ts - qwen3.5 is still behind vanilla, but on qwen3 it's comparable (but no real improvement). Maybe tomorrow I'll do some additional tests with P2P disabled.

PR

llama-bench -m Qwen3-VL-235B-A22B-Instruct-Q4_0-00001-of-00003.gguf --mmap 0 -fa 1 -ngl 99 -r 1 -p 16384 -ub 512,1024 -mg 0
ggml_cuda_init: found 5 ROCm devices:
  Device 0: Radeon RX 7900 XT, gfx1100 (0x1100), VMM: no, Wave Size: 32
  Device 1: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 2: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 3: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 4: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
| model                          |       size |     params | backend    | ngl | n_ubatch | fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | -: | --------------: | -------------------: |
| qwen3vlmoe 235B.A22B Q4_0      | 123.98 GiB |   235.09 B | ROCm       |  99 |      512 |  1 |         pp16384 |        405.72 ± 0.00 |
| qwen3vlmoe 235B.A22B Q4_0      | 123.98 GiB |   235.09 B | ROCm       |  99 |      512 |  1 |           tg128 |         26.69 ± 0.00 |
| qwen3vlmoe 235B.A22B Q4_0      | 123.98 GiB |   235.09 B | ROCm       |  99 |     1024 |  1 |         pp16384 |        352.24 ± 0.00 |
| qwen3vlmoe 235B.A22B Q4_0      | 123.98 GiB |   235.09 B | ROCm       |  99 |     1024 |  1 |           tg128 |         26.69 ± 0.00 |

Vanilla

llama-bench -m Qwen3-VL-235B-A22B-Instruct-Q4_0-00001-of-00003.gguf --mmap 0 -fa 1 -ngl 99 -r 1 -p 16384 -ub 512,1024 -mg 0
ggml_cuda_init: found 5 ROCm devices:
  Device 0: Radeon RX 7900 XT, gfx1100 (0x1100), VMM: no, Wave Size: 32
  Device 1: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 2: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 3: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
  Device 4: AMD Radeon Graphics, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64
| model                          |       size |     params | backend    | ngl | n_ubatch | fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | -: | --------------: | -------------------: |
| qwen3vlmoe 235B.A22B Q4_0      | 123.98 GiB |   235.09 B | ROCm       |  99 |      512 |  1 |         pp16384 |        376.78 ± 0.00 |
| qwen3vlmoe 235B.A22B Q4_0      | 123.98 GiB |   235.09 B | ROCm       |  99 |      512 |  1 |           tg128 |         26.74 ± 0.00 |
| qwen3vlmoe 235B.A22B Q4_0      | 123.98 GiB |   235.09 B | ROCm       |  99 |     1024 |  1 |         pp16384 |        416.65 ± 0.00 |
| qwen3vlmoe 235B.A22B Q4_0      | 123.98 GiB |   235.09 B | ROCm       |  99 |     1024 |  1 |           tg128 |         26.73 ± 0.00 |

@bchtrue

bchtrue commented Feb 28, 2026

Copy link
Copy Markdown

Can you please submit PR with only quick model loading. I hope it can be already merged and you will be able to test and work on parallelism?

@mxxm-t

mxxm-t commented Mar 3, 2026

Copy link
Copy Markdown
Author

Closing this PR and continue submitting smaller more focused PR's so its easier to review. Will leave branch in place for anyone that is interested,

@mxxm-t mxxm-t closed this Mar 3, 2026
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 20, 2026
…hed ⏭️ + AtomicBot-ai#76 CPU Fusion ⏭️

3 parallele Tiefen-Evals für Tier-3 Items:

AtomicBot-ai#74 Vulkan Descriptor Indexing (Bindless) ❌ VERWORFEN:
- Redundant mit AtomicBot-ai#85 Push Descriptors (✅ implementiert 2026-07-14)
- Push Descriptors eliminieren dieselben CPU-Aufrufe
- AtomicBot-ai#85-Benchmark auf Mars RADV: ±0.1-0.3% (Rauschen)
- Bindless würde über Push-Descriptors hinaus <0.5% bringen
- Workload-Mismatch: Bindless für draw-heavy Rendering, nicht Compute
- Mars/Venus bandwidth-bound, nicht descriptor-bound
- Aufwand revidiert: 2-4 → 3-5 Wochen (Shader-Rewrite aller .comp-Files)

AtomicBot-ai#75 Non-blocking Pipeline Scheduling ⏭️ SPÄTER:
- PR ggml-org#19922 closed (2026-03-03, unmerged, 4+ Mo stale)
- Fork hat bereits Upstream-Pipeline-Parallelismus
- Konflikt mit AtomicBot-ai#79 TP (✅+23-32% tg, split-mode-exklusiv)
- NVIDIA ungetestet, PP-Regression auf 4x MI50 gemeldet
- 2-GPU-Setup → geringer Bubble-Hebel
- Aufwand revidiert: 3-4 → 4-6 Wochen

AtomicBot-ai#76 CPU Backend Operator Fusion ⏭️ SPÄTER:
- RMS_NORM+MUL Fusion bereits im Fork (PR ggml-org#22423 upstream-merged)
- MoE Gated FFN riskant: PR ggml-org#20596 zeigt Regressionen auf Consumer-CPUs
  (M2: 0.98-1.00x, qwen3moe 30B: 0.75-0.96x bei t=2-4)
- Nur auf 96-Core-EPYC konsistente Gains (1.04-1.08x)
- Styx/Uranus haben Consumer-CPUs → wahrscheinlich Regression
- Re-Eval wenn PR ggml-org#20596 gemerged mit Regression-Freiheit
@mxxm-t
mxxm-t deleted the pipeline-parallelism branch September 4, 2026 07:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Nvidia GPU Issues specific to Nvidia GPUs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants