Skip to content

test-llama-archs : toggle causal_attn to catch graph shape changes - #29724

Merged
ggerganov merged 1 commit into
ggml-org:masterfrom
sihanyu03:test-llama-archs-causal-attn
Sep 30, 2026
Merged

ggerganov merged 1 commit into
ggml-org:masterfrom
sihanyu03:test-llama-archs-causal-attn

Conversation

@sihanyu03

Copy link
Copy Markdown
Contributor

Overview

This PR adds a check to test-llama-archs that the graph shape does not depend on causal_attn.

#28751 stopped re-reserving the scheduler when causal_attn changes, which is only safe if no model's graph shape depends on the flag. Until this PR, CI only toggled causal_attn through the server's tinygemma3 image tests, so this PR adds the check for every arch test-llama-archs covers.

After each row's existing decode, the check sets causal_attn off and decodes n_ubatch/2 and then n_ubatch tokens. If the graph changes with the flag, the first decode re-plans the compute buffers for the smaller batch (the graph size changed, so this re-plan counts as expected, with unexpected = false). The second decode of n_ubatch tokens has the same graph shape as the first but larger tensors, so it needs a reallocation. This time the graph hasn't changed from the previous decode, so the re-plan is unexpected = true and aborts under GGML_SCHED_NO_REALLOC.

Since #28751, every arch the test covers passes.

Additional information

  • Reverting the qwen4exp fix from context : do not re-reserve the scheduler when toggling causal_attn #28751 makes the test abort on qwen4exp with "unexpected graph reallocation", on CPU and CUDA. The same revert without this check passes.
  • Tested with the ggml-ci build flags and command (GGML_CUDA_DEVICES=1..4 ./build/bin/test-llama-archs -s 1) on 2xH100, and with all devices virtual on one GPU, and on CPU.
  • Not run on Metal, ROCm, Vulkan or WebGPU (the CI should cover this though).
  • Not covered: archs the test already skips.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes, AI was used to draft the code from a clear description. The code has since been carefully reviewed and cleaned, and I take full responsibility for the PR

@github-actions github-actions Bot added the testing Everything test related label Sep 30, 2026
@CISC

CISC commented Sep 30, 2026

Copy link
Copy Markdown
Member

Can you rebase to utilize the new models-check CI workflow?

@sihanyu03
sihanyu03 force-pushed the test-llama-archs-causal-attn branch from 3099613 to 730803e Compare September 30, 2026 08:58
@sihanyu03

Copy link
Copy Markdown
Contributor Author

Rebased to the most recent master tip

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice. Just adapt to the changes from #29601 and we can merge.

After the device decode, flip causal_attn off, decode n_ubatch/2 then
n_ubatch tokens. Both have the same node count, so a shape that depends
on the flag makes the second reallocate at an unchanged graph size,
which aborts under GGML_SCHED_NO_REALLOC. Skipped for the encode archs.
@sihanyu03
sihanyu03 force-pushed the test-llama-archs-causal-attn branch from 730803e to c353025 Compare September 30, 2026 17:03
@sihanyu03

Copy link
Copy Markdown
Contributor Author

Done, and tested on CPU and CUDA

@ggerganov
ggerganov merged commit 4f31296 into ggml-org:master Sep 30, 2026
18 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
…gml-org#29724)

After the device decode, flip causal_attn off, decode n_ubatch/2 then
n_ubatch tokens. Both have the same node count, so a shape that depends
on the flag makes the second reallocate at an unchanged graph size,
which aborts under GGML_SCHED_NO_REALLOC. Skipped for the encode archs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants