Improve CUDA graph capture - #19754
Conversation
Currently, CUDA graphs are eagerly enabled on the first call to ggml_backend_cuda_graph_compute. If the graph properties keep changing (4+ consecutive updates), the graph is permanently disabled. This is suboptimal because: - The first call always incurs CUDA graph capture overhead even if the graph is unstable - Once permanently disabled, CUDA graphs never re-enable even after the graph stabilizes (e.g., switching from prompt processing to decode) The new approach delays CUDA graph activation until warmup completes: the same cgraph must be called at least twice with matching properties before CUDA graph capture begins. This avoids wasted capture overhead on volatile graphs and allows graphs to become eligible once they stabilize. This also fixes issues such as ggml-org#19708
|
How often does a CUDA graph need to be evaluated to offset the overhead? If that number is low it may make more sense to just enable them more generally since we can now cache more than a single one. |
|
Actually, I think the problem was that due to the increasing context size CUDA graphs were not reusable for the prefill phase so it didn't make sense to use them. |
Exactly. I think PR #19645 enabled it for pre-fill phase, but it was causing CUDA graph to get disabled altogether after 4 consecutive updates. |
Ah yes, thanks for flagging that - I missed it. |
|
Regarding the CUDA Graph logic, we currently have the following state in the code:
The counter + disablement logic fires on both 2a) and 2b), whereas 2a) has performance parity to using cudaLaunch API over cudaGraph API, see the following numbers (tg 200 on a B6000):
For 2a/2b, I made From this data, I would recommend the following:
|
|
#19757 This is a WIP based on my insights above, which would enable cudaGraphs for PP on dense models that do not use |
One of the reason to take this approach in the PR was to fix issues like #19708. Basically, there could be cases where one |
|
@gaugarg-nv, @ORippler, I am just curious whether it is feasible to "freeze" the cuda graph. In some applications, e.g. diffusion, the same cgraph runs repeatedly without topography changes. In this case, even |
Unfortunately ggml does not allow for this yet. We wanted to tackle this as part of the graph-plan API, which was postponed indefinitely due to the hiatus of slaren. llama-context, the LLM orchestration loop of llama.cpp, already avoids rebuilding the graph if topology is consistent since #14482, and would simply have to forward this information to the backends somehow. Quoting from #14482:
|
@ORippler, thanks for the #14482 link. Didn't know such capability already exists. |
Following up on this with a TLDR: Given the limited perf-potential (talking about up to ~2% on Windows based on For the interested, continue reading along:Let's start with the collected perf numbers:Surprisingly, we don't see perf gains on Windows, despite there being a 2% gap between Windows and Linux in cudaLaunch-API-mode. Why? Let's take an nsight systems report to figure out
Ahh, we hit 2b) way more often than expected (my expectation was to hit only 2a). in |
| graph_key = ggml_cuda_graph_get_key(cgraph); | ||
|
|
||
| use_cuda_graph = ggml_cuda_graph_set_enabled(cuda_ctx, graph_key); | ||
| ggml_cuda_graph_set_enabled(cuda_ctx, graph_key); |
There was a problem hiding this comment.
strictly speaking this function solely checks whether the GPU supports cudaGraphs, but we can change the name in a separate PR
JohannesGaessler
left a comment
There was a problem hiding this comment.
Your code comments contain EM dashes. Unless there is a good reason not to, please stick to ASCII characters.
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Fixed now. |
am17an
left a comment
There was a problem hiding this comment.
nice and elegant solution!
Co-authored-by: Aman Gupta <amangupta052@gmail.com>
* Improve CUDA graph capture Currently, CUDA graphs are eagerly enabled on the first call to ggml_backend_cuda_graph_compute. If the graph properties keep changing (4+ consecutive updates), the graph is permanently disabled. This is suboptimal because: - The first call always incurs CUDA graph capture overhead even if the graph is unstable - Once permanently disabled, CUDA graphs never re-enable even after the graph stabilizes (e.g., switching from prompt processing to decode) The new approach delays CUDA graph activation until warmup completes: the same cgraph must be called at least twice with matching properties before CUDA graph capture begins. This avoids wasted capture overhead on volatile graphs and allows graphs to become eligible once they stabilize. This also fixes issues such as ggml-org#19708 * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Remove EM dashes * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Aman Gupta <amangupta052@gmail.com> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de> Co-authored-by: Aman Gupta <amangupta052@gmail.com>
|
@gaugarg-nv I'm seeing a performance regression for pp512 from this PR:
|
Thanks for reporting this. What's the command-line option used? Do you have a warm-up iteration? What is the number of runtime iterations? I guess that when you are testing pp512 with a micro-batch size of 512, the CUDA graph won't change across iterations, and the CUDA graph captured during warm-up will get reused. Can you try removing the warm-up loop or reducing runtime iteration and see if it has any impact? |
|
I tested the performance like this: |
|
I think I know what's going on. If I use 10 rather than 1 repetition for the benchmark there is basically no performance difference:
On the warmup run for pp512 CUDA graphs are not used because it's the first run. On the first benchmark run the same ggml graph is run for a second time so a CUDA graph is captured which introduces overhead. With 10 benchmark runs the overhead amortizes so there is no difference. So I would see this not as an issue with the code but rather with how we are benchmarking it. |
|
Right, I think one way to fix this is to increase the number of warmup iterations to 2 instead of the current 1. With that change, both implementations should show the same perf. |
* Improve CUDA graph capture Currently, CUDA graphs are eagerly enabled on the first call to ggml_backend_cuda_graph_compute. If the graph properties keep changing (4+ consecutive updates), the graph is permanently disabled. This is suboptimal because: - The first call always incurs CUDA graph capture overhead even if the graph is unstable - Once permanently disabled, CUDA graphs never re-enable even after the graph stabilizes (e.g., switching from prompt processing to decode) The new approach delays CUDA graph activation until warmup completes: the same cgraph must be called at least twice with matching properties before CUDA graph capture begins. This avoids wasted capture overhead on volatile graphs and allows graphs to become eligible once they stabilize. This also fixes issues such as ggml-org#19708 * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Remove EM dashes * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Aman Gupta <amangupta052@gmail.com> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de> Co-authored-by: Aman Gupta <amangupta052@gmail.com>
* Improve CUDA graph capture Currently, CUDA graphs are eagerly enabled on the first call to ggml_backend_cuda_graph_compute. If the graph properties keep changing (4+ consecutive updates), the graph is permanently disabled. This is suboptimal because: - The first call always incurs CUDA graph capture overhead even if the graph is unstable - Once permanently disabled, CUDA graphs never re-enable even after the graph stabilizes (e.g., switching from prompt processing to decode) The new approach delays CUDA graph activation until warmup completes: the same cgraph must be called at least twice with matching properties before CUDA graph capture begins. This avoids wasted capture overhead on volatile graphs and allows graphs to become eligible once they stabilize. This also fixes issues such as ggml-org#19708 * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Remove EM dashes * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Aman Gupta <amangupta052@gmail.com> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de> Co-authored-by: Aman Gupta <amangupta052@gmail.com>
* Improve CUDA graph capture Currently, CUDA graphs are eagerly enabled on the first call to ggml_backend_cuda_graph_compute. If the graph properties keep changing (4+ consecutive updates), the graph is permanently disabled. This is suboptimal because: - The first call always incurs CUDA graph capture overhead even if the graph is unstable - Once permanently disabled, CUDA graphs never re-enable even after the graph stabilizes (e.g., switching from prompt processing to decode) The new approach delays CUDA graph activation until warmup completes: the same cgraph must be called at least twice with matching properties before CUDA graph capture begins. This avoids wasted capture overhead on volatile graphs and allows graphs to become eligible once they stabilize. This also fixes issues such as ggml-org#19708 * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Remove EM dashes * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Aman Gupta <amangupta052@gmail.com> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de> Co-authored-by: Aman Gupta <amangupta052@gmail.com>
* Improve CUDA graph capture Currently, CUDA graphs are eagerly enabled on the first call to ggml_backend_cuda_graph_compute. If the graph properties keep changing (4+ consecutive updates), the graph is permanently disabled. This is suboptimal because: - The first call always incurs CUDA graph capture overhead even if the graph is unstable - Once permanently disabled, CUDA graphs never re-enable even after the graph stabilizes (e.g., switching from prompt processing to decode) The new approach delays CUDA graph activation until warmup completes: the same cgraph must be called at least twice with matching properties before CUDA graph capture begins. This avoids wasted capture overhead on volatile graphs and allows graphs to become eligible once they stabilize. This also fixes issues such as ggml-org/llama.cpp#19708 * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Remove EM dashes * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Aman Gupta <amangupta052@gmail.com> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de> Co-authored-by: Aman Gupta <amangupta052@gmail.com>
…AtomicBot-ai#5) 4 parallele Subagents (Vulkan/AMD, CUDA/MoE, arXiv Papers, Multi-GPU/Batching). 51 Items gesammelt, dedupliziert, 5 Quick-Wins verifiziert. Verifikation: 3/5 Quick-Wins bereits im Fork (ggml-org#15524=AtomicBot-ai#6, ggml-org#16391=--cram, ggml-org#19754= warmup). PR ggml-org#25479 (Pascal MMVQ) bereits als AtomicBot-ai#3 ✅. PR ggml-org#22887 (4K per Iter) bereits im Fork. 9 neue ROADMAP-Items hinzugefügt: - AtomicBot-ai#89 Transfer Queue AMD RDNA3 (PR ggml-org#19976) - AtomicBot-ai#90 MUL_MAT_ID Non-Square Tile (Discussion ggml-org#22598) - AtomicBot-ai#91 Internal AllReduce Kernel (PR ggml-org#22299) - AtomicBot-ai#92 RateQuant KV Cache (arXiv:2605.06675) - AtomicBot-ai#93 InnerQ KV Cache (arXiv:2602.23200) - TheTom#94 FineMoE Expert Offloading (arXiv:2502.05370) - TheTom#95 Shared Expert Aux Stream (TensorRT-LLM) - TheTom#96 Speculative Checkpointing (PR ggml-org#19493) - TheTom#97 Backend-agnostic TP Meta Device (PR ggml-org#19378) Vollständiger Report: docs/fork/2026-07-26_OPTIMIZATION_RESEARCH.md
* Improve CUDA graph capture Currently, CUDA graphs are eagerly enabled on the first call to ggml_backend_cuda_graph_compute. If the graph properties keep changing (4+ consecutive updates), the graph is permanently disabled. This is suboptimal because: - The first call always incurs CUDA graph capture overhead even if the graph is unstable - Once permanently disabled, CUDA graphs never re-enable even after the graph stabilizes (e.g., switching from prompt processing to decode) The new approach delays CUDA graph activation until warmup completes: the same cgraph must be called at least twice with matching properties before CUDA graph capture begins. This avoids wasted capture overhead on volatile graphs and allows graphs to become eligible once they stabilize. This also fixes issues such as ggml-org#19708 * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Remove EM dashes * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Aman Gupta <amangupta052@gmail.com> --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de> Co-authored-by: Aman Gupta <amangupta052@gmail.com>
Two CUDA graphs that share a cache key but differ in the first node's extents (e.g. micro-batch ne[2] 1 <-> 4) collapsed onto the same entry. Alternating between them under that key kept invalidating the shared warmup/update state: - before ggml-org#19754 the entry was permanently disabled once the update budget was exhausted, so both variants lost capture entirely - after ggml-org#19754 the entry is re-warmed on every switch instead, starving the warmup budget and never settling into a captured state while the alternation continues Fold the first node's extents (ne[]) into the cache key so each such variant gets its own entry. The eviction scan is untouched. The key still does not cover dtype, layout, or op params: two graphs that share ne[] and differ only in those attributes still collide. That is intentional -- the invariant below is what makes it safe. Safety invariant (unchanged): ggml_cuda_graph_update_required() still takes the full per-node props snapshot as the final validator, so a too-coarse key can only cost re-captures on every variant switch -- the validator itself is unchanged by this patch. Verification: - MMID reproducer (master @ 9cf3bf2): unpatched warmup 63/62 with alternating-phase collapse -> patched complete=2 / reset=0, no collapse - real workload (Qwen2.5-1.5B + 0.5B draft-simple, master @ 4d91760): 173.5 -> 205.3 t/s - no regression: 1024-token single-stream paired run 179.55 vs 179.21 t/s (vanilla vs patched); steady-shape warmup counters identical on both builds (complete=3 / reset<=2) Refs: ggml-org#19754 Assisted-by: glm-5.3-flash
Two CUDA graphs that share a cache key but differ in the first node's extents (e.g. micro-batch ne[2] 1 <-> 4) collapsed onto the same entry. Alternating between them under that key kept invalidating the shared warmup/update state: - before ggml-org#19754 the entry was permanently disabled once the update budget was exhausted, so both variants lost capture entirely - after ggml-org#19754 the entry is re-warmed on every switch instead, starving the warmup budget and never settling into a captured state while the alternation continues Fold the first node's extents (ne[]) into the cache key so each such variant gets its own entry. The eviction scan is untouched. The key still does not cover dtype, layout, or op params: two graphs that share ne[] and differ only in those attributes still collide. That is intentional -- the invariant below is what makes it safe. Safety invariant (unchanged): ggml_cuda_graph_update_required() still takes the full per-node props snapshot as the final validator, so a too-coarse key can only cost re-captures on every variant switch -- the validator itself is unchanged by this patch. Verification: - MMID reproducer (master @ 9cf3bf2): unpatched warmup 63/62 with alternating-phase collapse -> patched complete=2 / reset=0, no collapse - real workload (Qwen2.5-1.5B + 0.5B draft-simple, master @ 4d91760): 173.5 -> 205.3 t/s - no regression: 1024-token single-stream paired run 179.55 vs 179.21 t/s (vanilla vs patched); steady-shape warmup counters identical on both builds (complete=3 / reset<=2) Refs: ggml-org#19754 Assisted-by: glm-5.3-flash
Two CUDA graphs that share a cache key but differ in the first node's extents (e.g. micro-batch ne[2] 1 <-> 4) collapsed onto the same entry. Alternating between them under that key kept invalidating the shared warmup/update state: - before ggml-org#19754 the entry was permanently disabled once the update budget was exhausted, so both variants lost capture entirely - after ggml-org#19754 the entry is re-warmed on every switch instead, so recurring variant switches repeatedly disturb the shared warmup/update state and starve the warmup budget Fold the first node's extents (ne[]) into the cache key so each such variant gets its own entry. The eviction scan is untouched. The key still does not cover dtype, layout, or op params: two graphs that share ne[] and differ only in those attributes still collide. That is intentional -- the change is a partition refinement (P0(A) != P0(B) => P1(A) != P1(B)): domains that were already separate are never merged, and only variants that previously shared one domain are split apart. Entry selection is all this patch changes: the key feeds the map lookup that picks a cache entry, and the unmodified ggml_cuda_graph_update_required() logic (UID fast path and property comparison) runs inside that entry afterwards. This patch does not alter what is validated or when. Verification: - MMID reproducer (master @ 9cf3bf2): unpatched warmup 63/62 with alternating-phase collapse -> patched complete=2 / reset=0, no collapse - real workload (Qwen2.5-1.5B + 0.5B draft-simple, master @ 4d91760): 173.5 -> 205.3 t/s - no regression: 1024-token steady-state paired control 154.25 vs 155.90 t/s with the same steady-state replay count on both builds (graphs reused=1019, logged in ggml-org#28652) Refs: ggml-org#19754 Assisted-by: glm-5.3-flash Assisted-by: gpt-5.6

Currently, CUDA graphs are eagerly enabled on the first call to ggml_backend_cuda_graph_compute. If the graph properties keep changing (4+ consecutive updates), the graph is permanently disabled. This is suboptimal because:
The new approach delays CUDA graph activation until warmup completes: the same cgraph must be called at least twice with matching properties before CUDA graph capture begins. This avoids wasted capture overhead on volatile graphs and allows graphs to become eligible once they stabilize. This also fixes issues such as #19708
Perf improvement for Llama-8b-Q4_K_M on RTX 6000 Ada 300W:
Nsight profile Master:

Nsight profile PR:
