Skip to content

context : do not re-reserve the scheduler when toggling causal_attn - #28751

Merged
ggerganov merged 4 commits into
ggml-org:masterfrom
sihanyu03:context-causal-attn-no-rereserve
Sep 28, 2026
Merged

ggerganov merged 4 commits into
ggml-org:masterfrom
sihanyu03:context-causal-attn-no-rereserve

Conversation

@sihanyu03

@sihanyu03 sihanyu03 commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Overview

This PR removes the sched_need_reserve = true line from llama_context::set_causal_attn(), which previously marked the scheduler to do a full re-reserve on every change of the flag, and fixes the one architecture (qwen4exp) whose graph shape depended on the flag.

The re-reserve is unnecessary in this case because causal_attn only changes the values written to the KQ mask, not tensor shapes or any other buffer sizes. qwen4exp was the one exception that selected the per block/cell bias path based on the flag, so the causal and non-causal graphs differed in tensor shapes and ops. The second commit selects the block path independent of causal_attn, from the mask shape only and makes the non-causal case follow the reference rule: every visible block competes on score and only unpooled cells are always selected. The causal case works as before.

Additional information

For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma3, Gemma4, and DeepSeek-V4, resulting in two expensive sched_reserve() passes per image. This is especially slow for multi-image or video inputs.

The cost of a re-reserve scales with context and ubatch configurations, so larger settings pay more per image (see table below).

Note: causal_attn is a graph reuse key (llm_graph_params via cparams), so a new graph is built regardless of sched_need_reserve, so this doesn't change the graph rebuilding behaviour.

Speedup table: llama-server with gemma-4-26B-A4B Q4_0 + BF16 mmproj, 130-token images, cache_prompt=false, prompt_ms median of 3 (before -> after):

images config H200 before -> after RTX 4090 before -> after
1 -c 8192 -ub 512 134 -> 105 ms (1.27×) 201 -> 119 ms (1.69×)
24 -c 8192 -ub 512 2278 -> 1562 ms (1.46×) 3559 -> 1748 ms (2.04×)
24 -c 32768 -ub 2048 5379 -> 1584 ms (3.40×) 13377 -> 1759 ms (7.61×)

Generated output remains identical before and after.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes, AI was used to assist in understanding the code, hunting for the cause, implementing initial draft and reviewing. I have subsequently reviewed and tested the code and take full responsibility for it and its maintenance.

@am17an

am17an commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

This seems correct but it is not necessarily true for all models, we would need to check

@sihanyu03

Copy link
Copy Markdown
Contributor Author

This seems correct but it is not necessarily true for all models, we would need to check

Thanks for pointing this out, you're right. I grepped the usage of cparams.causal_attn in src/. The only read that affects tensor sizes is from qwen4exp (Qwen3.8-Flash-Next). In src/models/qwen4exp.cpp:552-569, causal_attn affects blk_bias, which affects the tensor size of qsa->bias. This is now fixed in the PR description.

Two reasons why I think this is still safe:

  1. The flag is part of llm_graph_params, so the next decode rebuilds the graph nonetheless. When the rebuilt graph is allocated, ggml_backend_sched_alloc_splits falls back to ggml_gallocr_reserve_n and grows only the compute buffers that need it. This will happen at most once, so the cost is bounded. llama_context::set_embeddings (src/llama-context.cpp:1162-1169) also has re-reserve commented out despite it changing the graph, relying on the same path.
  2. No code in this repo flips this flag at runtime for qwen4exp models. mtmd only flips it for gemma3, gemma4v, gemma4uv and deepseek4v (tools/mtmd/mtmd.cpp:2170-2186). The only other runtime callers (diffusion example, DFlash draft path) don't apply to qwen4exp.

I also tested Qwen3.8-Flash-Next with the flag flipped via the API after context creation, between prompt and generation, and compared my branch's result vs master with identical results.

If you'd rather not rely on the reallocation at all, I'd be happy to handle qwen4exp differently in a follow-up where the size of qsa->bias doesn't depend on the flag.

@am17an

am17an commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

It's safe but it will likely fail CI when we add tests for that model as we run CI with GGML_SCHED_NO_REALLOC. We should fix that model first.

@ddh0

ddh0 commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

I've tested this change locally, using DeepSeek-V4-Flash-Vision-Exp.

Previously, I was seeing the graph re-reserved before processing any image, between every image, and once to switch back to text generation. Each re-reserve took about 5 second on my machine.

Now, processing of multiple images is significantly faster end-to-end, even though the encode/decode time per-image has not changed.

I've tested with a few different multi-turn, multi-image tasks. As far as I can tell, the quality and content of the model's output has not changed.

I really appreciate this change, and hope it can be worked into mainline at some point. Thanks!

4.09.576.687 I slot create_check: id  0 | task 649 | created context checkpoint 1 of 64 (pos_min = 1006, pos_max = 1006, n_tokens = 1007, size = 17.021 MiB)
4.17.603.589 I slot print_timing: id  0 | task 649 | prompt processing, n_tokens =    418, progress = 0.39, t =   8.03 s / 52.05 tokens per second
4.17.603.591 I slot   operator(): id  0 | task 649 | cached n_tokens = 1425, memory_seq_rm [1425, end)
4.17.603.716 I slot process_mtmd: id  0 | task 649 | encoding mtmd batch from idx = 1425, n_chunks = 1
4.17.809.659 I decoding image batch 1/1, n_tokens_batch = 380
4.31.952.234 I image decoded (batch 1/1) in 14143 ms
4.31.952.519 I slot process_mtmd: id  0 | task 649 | encoding mtmd batch from idx = 1805, n_chunks = 1
4.32.087.821 I decoding image batch 1/1, n_tokens_batch = 380
4.41.292.756 I image decoded (batch 1/1) in 9205 ms
4.41.293.001 I slot process_mtmd: id  0 | task 649 | encoding mtmd batch from idx = 2185, n_chunks = 1
4.41.427.390 I decoding image batch 1/1, n_tokens_batch = 380
4.50.433.660 I image decoded (batch 1/1) in 9006 ms
4.50.433.931 I slot process_mtmd: id  0 | task 649 | encoding mtmd batch from idx = 2565, n_chunks = 1
4.50.567.375 I decoding image batch 1/1, n_tokens_batch = 380
4.58.957.757 I image decoded (batch 1/1) in 8390 ms
4.58.958.001 I slot process_mtmd: id  0 | task 649 | encoding mtmd batch from idx = 2945, n_chunks = 1
4.59.075.577 I decoding image batch 1/1, n_tokens_batch = 328
5.07.485.631 I image decoded (batch 1/1) in 8410 ms
5.07.485.878 I slot process_mtmd: id  0 | task 649 | encoding mtmd batch from idx = 3273, n_chunks = 1
5.07.636.161 I decoding image batch 1/1, n_tokens_batch = 372
5.16.466.875 I image decoded (batch 1/1) in 8830 ms
5.16.940.675 I slot print_timing: id  0 | task 649 | prompt processing, n_tokens =   2648, progress = 1.00, t =  67.37 s / 39.31 tokens per second
5.16.940.681 I slot   operator(): id  0 | task 649 | cached n_tokens = 3655, memory_seq_rm [3655, end)
5.16.941.142 I slot init_sampler: id  0 | task 649 | init sampler, took 0.12 ms, tokens: text = 1439, total = 3659
5.16.943.886 I slot create_check: id  0 | task 649 | created context checkpoint 2 of 64 (pos_min = 3654, pos_max = 3654, n_tokens = 3655, size = 17.021 MiB)
5.28.926.566 I slot print_timing: id  0 | task 649 | n_gen =    100, tg =   8.45 t/s, tg_3s =   8.54 t/s

@sihanyu03
sihanyu03 requested a review from CISC as a code owner September 13, 2026 17:30
@github-actions github-actions Bot added the model Model specific label Sep 13, 2026
@sihanyu03

Copy link
Copy Markdown
Contributor Author

It's safe but it will likely fail CI when we add tests for that model as we run CI with GGML_SCHED_NO_REALLOC. We should fix that model first.

You're right. This has now been fixed in my second commit. qwen4exp now selects the block bias path independently of causal_attn, which was the only place where tensor shapes or ops depended on the flag, and llama_memory_hybrid_idx::set_input_qsa was modified to use the block bias path for non-causal too, leaving causal unchanged.

@sihanyu03

Copy link
Copy Markdown
Contributor Author

I've tested this change locally, using DeepSeek-V4-Flash-Vision-Exp.

Thank you for testing!

@sihanyu03
sihanyu03 force-pushed the context-causal-attn-no-rereserve branch from 5012372 to 421fa4b Compare September 17, 2026 14:33
@sihanyu03

Copy link
Copy Markdown
Contributor Author

Rebased on the current master without conflicts, same three commits still.

@am17an does the second commit address the GGML_SCHED_NO_REALLOC concern? With it, qwen4exp builds the same graph for both causal_attn values, and toggling the flag under that setting no longer reallocates.

@am17an

am17an commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

I'm still no sure. @ngxson can you check if this makes sense?

`llama_context::set_causal_attn()` marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive `sched_reserve()` passes per image. This is especially slow for multi-image or video inputs.

The cost of a re-reserve scales with context and ubatch configurations, so larger settings pay more per image (see table below).

The re-reserve is unnecessary in this case because `causal_attn` only changes the values written to KQ mask, not tensor shapes or any other buffer sizes.

Note: `causal_attn` is a graph reuse key (`llm_graph_params` via `cparams`), so a new graph is built regardless of `sched_need_reserve`, so this doesn't change the graph rebuilding behaviour.

llama-server with gemma-4-26B-A4B Q4_0 + BF16 mmproj, 130-token images,
cache_prompt=false, prompt_ms median of 3 (before -> after):

| images | config | H200 before -> after | RTX 4090 before -> after |
|-|-|-|-|
| 1 | `-c 8192 -ub 512` | 134 -> 105 ms (1.27×) | 201 -> 119 ms (1.69×) |
| 24 | `-c 8192 -ub 512` | 2278 -> 1562 ms (1.46×) | 3559 -> 1748 ms (2.04×) |
| 24 | `-c 32768 -ub 2048` | 5379 -> 1584 ms (3.40×) | 13377 -> 1759 ms (7.61×) |

Generated output remains identical before and after.
The block/cell bias path was selected on cparams.causal_attn, so the
causal and non-causal graphs differed in tensor shapes and ops. With the
re-reserve removed (previous commit), a runtime flip resulted in
reallocating the compute buffers, which would fail under
GGML_SCHED_NO_REALLOC.

This commit selects the block path from the mask shape only, independent
of causal_attn. causal_attn is instead passed to set_input_qsa.
causal_attn is fixed per graph as it's part of the reuse key. Causal
values are unchanged. Non-causal values now follow the reference rule,
where every visible block competes on score and only unpooled cells are
always selected.
@sihanyu03
sihanyu03 force-pushed the context-causal-attn-no-rereserve branch from 421fa4b to 0cbdaaf Compare September 21, 2026 19:12
@sihanyu03

Copy link
Copy Markdown
Contributor Author

@am17an here is a check for the GGML_SCHED_NO_REALLOC concern that anyone can run.

I added a toggle step to test-llama-archs (branch: sihanyu03/llama.cpp@context-causal-attn-no-rereserve...context-causal-attn-no-rereserve-test). After the normal decode it sets causal_attn off, decodes n_ubatch/2 then n_ubatch tokens, sets it back on and decodes once more. Build with -DGGML_SCHED_NO_REALLOC=ON and run test-llama-archs -a qwen4exp:

  • master + only the first commit (no re-reserve on flip): qwen4exp aborts with ggml_backend_sched_alloc_splits: unexpected graph reallocation, just like you said
  • master + the full PR: qwen4exp and the whole sweep pass under the flag (all 366 tests)

I also did a real model check on the rebased head for Qwen3.8-Flash-Next: perplexity on wikitext-2 with -c 4096 -ub 510 is identical to master (~3.6622), and greedy tokens are identical too, with or without causal_attn flips during generation.

I'd be happy to add the test to this PR or send it as a follow-up, whichever you prefer (tested on CUDA and CPU).

Could someone also approve the CI run? The same head is green on my fork's CI.

@sihanyu03

Copy link
Copy Markdown
Contributor Author

@am17an thanks for approving the CI runs, everything is green now.

The PR is ready for review, would you or one of the other reviewers have time to take a look please?

@am17an

am17an commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

@ggerganov @ngxson PTAL. I think it's a good change

@sihanyu03

Copy link
Copy Markdown
Contributor Author

@am17an would you be willing to give a formal review yourself, or should it wait for a core maintainer?

@ggerganov

ggerganov commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

I think disabling the causal_attn reserve per #28927 is OK. Regarding the Qwen4 changes - can't tell if it makes sense. My preference is instead of making more changes to rewrite the Qwen4 inference graph and memory because the current implementation seems very inefficient.

@am17an

am17an commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

We can remove the Qwen4 changes, those are only because it will fail GGML_SCHED_NO_REALLOC tests if we include them. For now we can remove those.

@sihanyu03

sihanyu03 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

@ggerganov thanks for taking a look.

This PR predates #28927 and its first commit is the identical change. Given your preference for a rewrite over more qwen4exp changes, which do you prefer:

  1. take this small shape fix now and leave the rewrite for later
  2. keep the reserve only for qwen4exp until the rewrite
  3. merge the reserve removal alone (the first commit, identical to context : do not re-reserve the scheduler on set_causal_attn #28927) and accept the qwen4exp abort under the flag until then

Happy to do whichever you choose.

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's keep them to not break the CI. We have to be careful to not use the llama-memory-hybrid-idx memory for other models before it is rewritten.

@ggerganov ggerganov added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 28, 2026
@am17an

am17an commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

We have to be careful to not use the llama-memory-hybrid-idx memory for other models before it is rewritten.

It's being used for GLM-5.3 next. I think post that model we should not use it before a refactor.

@ggerganov

Copy link
Copy Markdown
Member

Actually, I am not sure if the inefficiency of Qwen4 is just the graph, or both the graph + memory. Do you have an estimate?

@am17an

am17an commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

The memory/attention is the main problem there. After the GLM 5.3 next PR we can re-use the same ideas from there to make it better.

@ggerganov
ggerganov merged commit ed7ac35 into ggml-org:master Sep 28, 2026
14 checks passed
@sihanyu03
sihanyu03 deleted the context-causal-attn-no-rereserve branch September 28, 2026 10:22
Patt92 pushed a commit to Patt92/llama.cpp that referenced this pull request Sep 28, 2026
Conflict: src/models/qwen4exp.cpp - upstream ed7ac35 (ggml-org#28751) drops
cparams.causal_attn from the blk_bias predicate and carries it into
llm_graph_input_qsa instead (it is part of the graph reuse key); the fork's
!gather gate on the same predicate stays, since the gather path needs the
per-cell bias as its attention mask. Upstream's [TAG_QWEN4_REIMPLEMENT] note
and the fork's shared-MTP helper both kept.
@ggerganov

Copy link
Copy Markdown
Member

@sihanyu03 @am17an Could you take a look at this failing test: https://github.com/ggml-org/llama.cpp/actions/runs/36400625405/job/108857353856#step:11:3174 - seems to happen when pipeline parallelism is enabled (i.e. env GGML_CUDA_DEVICES=2).

@am17an

am17an commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

It looks like a pre-existing bug that has surfaced because we removed the re-reserve which only surfaces during PP or maybe the reserve was actually specifically added for this case. I don't have time to investigate this so we can revert this PR or add this diff.

--- a/src/llama-context.cpp
+++ b/src/llama-context.cpp
@@ -1256,8 +1256,10 @@
 
     cparams.causal_attn = value;
 
-    // no scheduler reserve needed because graph shapes must not depend on causal_attn, a flip only rebuilds the graph
-    //sched_need_reserve = true;
+    // with pipeline parallelism, the embd input of a non-causal image re-plans the compute buffers for that graph only, so reserve again
+    if (cparams.pipeline_parallel) {
+        sched_need_reserve = true;
+    }
 }
 
 bool llama_context::get_causal_attn() const {

@sihanyu03

Copy link
Copy Markdown
Contributor Author

Yes I think it's the same issue as #26873. I will look into it

@sihanyu03

Copy link
Copy Markdown
Contributor Author

Root cause confirmed, and it is indeed the mechanism described on #26873. The image batch with embedding input has a different graph shape, so the allocator re-plans on it and the worst-case plan is lost. Every later graph that is larger in any tensor re-plans again, and with pipeline parallelism each re-plan costs a full barrier across the devices, and under GGML_SCHED_NO_REALLOC the re-plan on the text batch after the image is flagged as unexpected, which is the CI failure. The removed per-flip reserve was masking this for models that toggle causal_attn. Glimmer in #26873 never toggles, which is why this issue appeared before this PR.

@am17an's diff would fix the CI failure but ties the reserve to the flag again and pays two reserves per image under pipeline parallelism, so multi-GPU users wouldn't benefit from this PR.

What I am preparing instead (which should be done before tomorrow) is re-reserve once when a decode switches from embedding inputs back to token inputs, which restores the worst-case plan, costs one reserve per image run on any setup, and covers the #26873 models too. This would fix the failing CIs. However, this doesn't tackle the underlying limitation which would be a much bigger change.

@ggerganov

Copy link
Copy Markdown
Member

The image batch with embedding input has a different graph shape

Hm, that shouldn't be the case. If it were the case, then the previous CI with a single GPU would have also failed. It's something related to the pipeline parallelism.

@am17an

am17an commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

When a graph shape changes re-allocation is allowed but if the graph shape doesn't change and number of outputs change then it will crash. This is what is happening at least in the PP case.

baykalokandemir added a commit to baykalokandemir/llama.cpp-bells-tortured that referenced this pull request Sep 28, 2026
…ggml-org#28118 and ggml-org#28751 do not apply (MTP uses bounded rollback, no causal toggles)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018co7fcJoEQXMxLezg7HpMU
@sihanyu03

Copy link
Copy Markdown
Contributor Author

Yes sorry, I meant the scheduler graph changes, not the llama graph. The worst-case plan is created for a text graph, which does not read inp->embd (inputs are token ids, not embeddings, unlike image inputs), so it is not registered as an input. With pipeline parallelism, 4 copies of each read input are created, so for the image chunk, these 4 copies of inp->embd alter the graph the scheduler builds and the replan is 'expected' (backend_ids_changed becomes true so ggml-backend.cpp:1619 becomes unexpected = false) and hence allowed, so no abort. The text chunk after the image chunk similarly replans as the graph changes, so it's allowed again with no abort. But this plan is not the worst-case plan, and has out_ids at 0 bytes (that batch has no output rows). The final text chunk, which needs outputs and has out_ids != 0 bytes, then fails the check as the tensor no longer fits in the window given by the plan, thus again, needing a replan. However this is now not 'expected', because the graph didn't change (replans are allowed only if the graph changes, as @am17an said), which then triggers the abort under GGML_SCHED_NO_REALLOC. This didn't happen for the single-GPU case because single-GPU has no pipeline parallelism and therefore no copies, so the graph shape never changes and the worst-case plan is never superseded.

My previous suggestion had some issues. A better solution could instead be a dummy read of inp->embd within the text graph, so that graph shapes stay constant across image and text chunks, re-reserve is never needed, and the behaviour of this PR is kept. This increases the compute buffer by 3 extra inp->embd-sized slots for text-only use under pipeline parallelism (image use already pays this after the first image), which should be acceptable. I will test this tomorrow and if it has no issues, open a PR, unless you'd want a quick fix to the CI with @am17an's diff. If for some reason this doesn't work, the fallback is the diff too.

@ggerganov

Copy link
Copy Markdown
Member

@sihanyu03 #29634 should resolve the problem. Please confirm it works as expected for your use case.

pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
…gml-org#28751)

* context : do not re-reserve the scheduler when toggling causal_attn

`llama_context::set_causal_attn()` marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive `sched_reserve()` passes per image. This is especially slow for multi-image or video inputs.

The cost of a re-reserve scales with context and ubatch configurations, so larger settings pay more per image (see table below).

The re-reserve is unnecessary in this case because `causal_attn` only changes the values written to KQ mask, not tensor shapes or any other buffer sizes.

Note: `causal_attn` is a graph reuse key (`llm_graph_params` via `cparams`), so a new graph is built regardless of `sched_need_reserve`, so this doesn't change the graph rebuilding behaviour.

llama-server with gemma-4-26B-A4B Q4_0 + BF16 mmproj, 130-token images,
cache_prompt=false, prompt_ms median of 3 (before -> after):

| images | config | H200 before -> after | RTX 4090 before -> after |
|-|-|-|-|
| 1 | `-c 8192 -ub 512` | 134 -> 105 ms (1.27×) | 201 -> 119 ms (1.69×) |
| 24 | `-c 8192 -ub 512` | 2278 -> 1562 ms (1.46×) | 3559 -> 1748 ms (2.04×) |
| 24 | `-c 32768 -ub 2048` | 5379 -> 1584 ms (3.40×) | 13377 -> 1759 ms (7.61×) |

Generated output remains identical before and after.

* qwen4exp : make the indexer bias shape independent of causal_attn

The block/cell bias path was selected on cparams.causal_attn, so the
causal and non-causal graphs differed in tensor shapes and ops. With the
re-reserve removed (previous commit), a runtime flip resulted in
reallocating the compute buffers, which would fail under
GGML_SCHED_NO_REALLOC.

This commit selects the block path from the mask shape only, independent
of causal_attn. causal_attn is instead passed to set_input_qsa.
causal_attn is fixed per graph as it's part of the reuse key. Causal
values are unchanged. Non-causal values now follow the reference rule,
where every visible block competes on score and only unpooled cells are
always selected.

* context : state the causal_attn shape rule in the comment

* cont : add TODOs

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 1, 2026
…gml-org#28751)

* context : do not re-reserve the scheduler when toggling causal_attn

`llama_context::set_causal_attn()` marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive `sched_reserve()` passes per image. This is especially slow for multi-image or video inputs.

The cost of a re-reserve scales with context and ubatch configurations, so larger settings pay more per image (see table below).

The re-reserve is unnecessary in this case because `causal_attn` only changes the values written to KQ mask, not tensor shapes or any other buffer sizes.

Note: `causal_attn` is a graph reuse key (`llm_graph_params` via `cparams`), so a new graph is built regardless of `sched_need_reserve`, so this doesn't change the graph rebuilding behaviour.

llama-server with gemma-4-26B-A4B Q4_0 + BF16 mmproj, 130-token images,
cache_prompt=false, prompt_ms median of 3 (before -> after):

| images | config | H200 before -> after | RTX 4090 before -> after |
|-|-|-|-|
| 1 | `-c 8192 -ub 512` | 134 -> 105 ms (1.27×) | 201 -> 119 ms (1.69×) |
| 24 | `-c 8192 -ub 512` | 2278 -> 1562 ms (1.46×) | 3559 -> 1748 ms (2.04×) |
| 24 | `-c 32768 -ub 2048` | 5379 -> 1584 ms (3.40×) | 13377 -> 1759 ms (7.61×) |

Generated output remains identical before and after.

* qwen4exp : make the indexer bias shape independent of causal_attn

The block/cell bias path was selected on cparams.causal_attn, so the
causal and non-causal graphs differed in tensor shapes and ops. With the
re-reserve removed (previous commit), a runtime flip resulted in
reallocating the compute buffers, which would fail under
GGML_SCHED_NO_REALLOC.

This commit selects the block path from the mask shape only, independent
of causal_attn. causal_attn is instead passed to set_input_qsa.
causal_attn is fixed per graph as it's part of the reuse key. Causal
values are unchanged. Non-causal values now follow the reference rule,
where every visible block competes on score and only unpooled cells are
always selected.

* context : state the causal_attn shape rule in the comment

* cont : add TODOs

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 2, 2026
…gml-org#28751)

* context : do not re-reserve the scheduler when toggling causal_attn

`llama_context::set_causal_attn()` marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive `sched_reserve()` passes per image. This is especially slow for multi-image or video inputs.

The cost of a re-reserve scales with context and ubatch configurations, so larger settings pay more per image (see table below).

The re-reserve is unnecessary in this case because `causal_attn` only changes the values written to KQ mask, not tensor shapes or any other buffer sizes.

Note: `causal_attn` is a graph reuse key (`llm_graph_params` via `cparams`), so a new graph is built regardless of `sched_need_reserve`, so this doesn't change the graph rebuilding behaviour.

llama-server with gemma-4-26B-A4B Q4_0 + BF16 mmproj, 130-token images,
cache_prompt=false, prompt_ms median of 3 (before -> after):

| images | config | H200 before -> after | RTX 4090 before -> after |
|-|-|-|-|
| 1 | `-c 8192 -ub 512` | 134 -> 105 ms (1.27×) | 201 -> 119 ms (1.69×) |
| 24 | `-c 8192 -ub 512` | 2278 -> 1562 ms (1.46×) | 3559 -> 1748 ms (2.04×) |
| 24 | `-c 32768 -ub 2048` | 5379 -> 1584 ms (3.40×) | 13377 -> 1759 ms (7.61×) |

Generated output remains identical before and after.

* qwen4exp : make the indexer bias shape independent of causal_attn

The block/cell bias path was selected on cparams.causal_attn, so the
causal and non-causal graphs differed in tensor shapes and ops. With the
re-reserve removed (previous commit), a runtime flip resulted in
reallocating the compute buffers, which would fail under
GGML_SCHED_NO_REALLOC.

This commit selects the block path from the mask shape only, independent
of causal_attn. causal_attn is instead passed to set_input_qsa.
causal_attn is fixed per graph as it's part of the reuse key. Causal
values are unchanged. Non-causal values now follow the reference rule,
where every visible block competes on score and only unpooled cells are
always selected.

* context : state the causal_attn shape rule in the comment

* cont : add TODOs

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. model Model specific

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants