test-llama-archs : toggle causal_attn to catch graph shape changes - #29724
Merged
ggerganov merged 1 commit intoSep 30, 2026
Merged
Conversation
Member
|
Can you rebase to utilize the new |
sihanyu03
force-pushed
the
test-llama-archs-causal-attn
branch
from
September 30, 2026 08:58
3099613 to
730803e
Compare
Contributor
Author
|
Rebased to the most recent master tip |
CISC
approved these changes
Sep 30, 2026
After the device decode, flip causal_attn off, decode n_ubatch/2 then n_ubatch tokens. Both have the same node count, so a shape that depends on the flag makes the second reallocate at an unchanged graph size, which aborts under GGML_SCHED_NO_REALLOC. Skipped for the encode archs.
sihanyu03
force-pushed
the
test-llama-archs-causal-attn
branch
from
September 30, 2026 17:03
730803e to
c353025
Compare
Contributor
Author
|
Done, and tested on CPU and CUDA |
CISC
approved these changes
Sep 30, 2026
pierreguillot
pushed a commit
to Ircam-Partiels/llama.cpp
that referenced
this pull request
Oct 1, 2026
…gml-org#29724) After the device decode, flip causal_attn off, decode n_ubatch/2 then n_ubatch tokens. Both have the same node count, so a shape that depends on the flag makes the second reallocate at an unchanged graph size, which aborts under GGML_SCHED_NO_REALLOC. Skipped for the encode archs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
This PR adds a check to
test-llama-archsthat the graph shape does not depend oncausal_attn.#28751 stopped re-reserving the scheduler when
causal_attnchanges, which is only safe if no model's graph shape depends on the flag. Until this PR, CI only toggledcausal_attnthrough the server's tinygemma3 image tests, so this PR adds the check for every archtest-llama-archscovers.After each row's existing decode, the check sets
causal_attnoff and decodesn_ubatch/2and thenn_ubatchtokens. If the graph changes with the flag, the first decode re-plans the compute buffers for the smaller batch (the graph size changed, so this re-plan counts as expected, withunexpected = false). The second decode ofn_ubatchtokens has the same graph shape as the first but larger tensors, so it needs a reallocation. This time the graph hasn't changed from the previous decode, so the re-plan isunexpected = trueand aborts underGGML_SCHED_NO_REALLOC.Since #28751, every arch the test covers passes.
Additional information
qwen4expwith "unexpected graph reallocation", on CPU and CUDA. The same revert without this check passes.GGML_CUDA_DEVICES=1..4 ./build/bin/test-llama-archs -s 1) on 2xH100, and with all devices virtual on one GPU, and on CPU.Requirements