Skip to content

CUDA: make the alloc_deps check batch independent - #29986

Merged
am17an merged 1 commit into
masterfrom
aman/cuda-fix-graph-realloc
Oct 5, 2026
Merged

am17an merged 1 commit into
masterfrom
aman/cuda-fix-graph-realloc

Conversation

@am17an

@am17an am17an commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #29980

Overview

Additional information

Requirements

@am17an
am17an requested a review from a team as a code owner October 5, 2026 10:03
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Oct 5, 2026

@ServeurpersoCom ServeurpersoCom left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tested on a single RTX PRO 6000 with Qwen3.6-35B-A3B Q8_0 (shared expert fusion active) and the prompt from #29980: master aborts under GGML_SCHED_DEBUG_REALLOC=1 and this PR fixes it, prefill and TG are unchanged and greedy output is identical. The 2x prefill drop does not show on one GPU, so it is likely the multi-GPU sync cost on each re-reserve.

Minor: the match now drops the batch independent checks of should_fuse_mul_mat_vec_q too (Pascal cc, F32 types, padding), so graph_optimize can reorder for a fusion that never runs, splitting that function into a static part and a batch part would keep them.

@am17an
am17an merged commit e117148 into master Oct 5, 2026
13 of 17 checks passed
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eval bug: prompt processing ~2x slower on Qwen3.6-35B-A3B since #29184 (fuse shared experts into MMVQ)

3 participants