Skip to content

ggml-cuda: use per-thread stream for buffer-init padding memset - #28782

Merged
JohannesGaessler merged 2 commits into
ggml-org:masterfrom
harkgill-amd:harkgill/hip-buffer-init-perthread-stream
Oct 6, 2026
Merged

JohannesGaessler merged 2 commits into
ggml-org:masterfrom
harkgill-amd:harkgill/hip-buffer-init-perthread-stream

Conversation

@harkgill-amd

Copy link
Copy Markdown
Contributor

Overview

ggml_backend_cuda_buffer_init_tensor utilizes cudaMemset for zero tensor padding. This runs on the legacy stream which can collide with parallel HIP graph captures on separate streams causing failures.

This switches it to cudaMemsetAsync + sync (matches existing precedent in the file), keeping it off the legacy stream.

Additional information

Resolves the test-backend-ops -j failures below on ROCm.

ROCm 10.2 + gfx1151 running test-backend-ops -j $(nproc),

  • Before
2026-09-11T14:23:16.3067027Z ROCm error: operation failed due to a previous error during capture
2026-09-11T14:23:16.3067555Z   current device: 0, in function ggml_cuda_graph_evaluate_and_capture at /scratch/actions-runner/_work/llama.cpp/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu:4366
2026-09-11T14:23:16.3068105Z ROCm error: operation would make the legacy stream depend on a capturing blocking stream
2026-09-11T14:23:16.3068441Z   hipStreamEndCapture(cuda_ctx->stream(), &graph->graph)
2026-09-11T14:23:16.3068935Z   current device: 0, in function ggml_backend_cuda_buffer_init_tensor at /scratch/actions-runner/_work/llama.cpp/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu:770
  • After
  ...
  Backend 1/2: ROCm0
    ... (all ops OK) ...
  2/2 backends passed
  [exit 0]

Serial case is not affected, passes before and after fix.

Requirements

@harkgill-amd
harkgill-amd requested a review from a team as a code owner September 11, 2026 20:18
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Sep 11, 2026
@harkgill-amd
harkgill-amd force-pushed the harkgill/hip-buffer-init-perthread-stream branch from a177716 to 744a074 Compare September 11, 2026 21:21
@github-actions github-actions Bot added the devops improvements to build systems and github actions label Sep 11, 2026
@JohannesGaessler
JohannesGaessler merged commit 51ce9c1 into ggml-org:master Oct 6, 2026
23 of 24 checks passed
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
…-org#28782)

* ggml-cuda: use per-thread stream for buffer-init padding memset

* ci : re-enable test-backend-ops -j for ROCm

(cherry picked from commit 51ce9c1)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend devops improvements to build systems and github actions ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants