Repository navigation
ggml-cuda: use per-thread stream for buffer-init padding memset - #28782
Merged
JohannesGaessler merged 2 commits intoOct 6, 2026
Merged
JohannesGaessler merged 2 commits into
JohannesGaessler merged 2 commits into
Conversation
harkgill-amd
force-pushed
the
harkgill/hip-buffer-init-perthread-stream
branch
from
September 11, 2026 21:21
a177716 to
744a074
Compare
IMbackK
approved these changes
Sep 13, 2026
JohannesGaessler
approved these changes
Oct 6, 2026
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 8, 2026
…-org#28782) * ggml-cuda: use per-thread stream for buffer-init padding memset * ci : re-enable test-backend-ops -j for ROCm (cherry picked from commit 51ce9c1)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
ggml_backend_cuda_buffer_init_tensorutilizes cudaMemset for zero tensor padding. This runs on the legacy stream which can collide with parallel HIP graph captures on separate streams causing failures.This switches it to cudaMemsetAsync + sync (matches existing precedent in the file), keeping it off the legacy stream.
Additional information
Resolves the
test-backend-ops -jfailures below on ROCm.ROCm 10.2 + gfx1151 running
test-backend-ops -j $(nproc),Serial case is not affected, passes before and after fix.
Requirements