Skip to content

llama : keep input embeddings on layer 0 preventing unneeded scheduling thread wakeups - #30152

Closed
deadprogram wants to merge 1 commit into
ggml-org:masterfrom
deadprogram:fix-kv-clear-context
Closed

deadprogram wants to merge 1 commit into
ggml-org:masterfrom
deadprogram:fix-kv-clear-context

Conversation

@deadprogram

@deadprogram deadprogram commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Overview

This PR is to address performance degredation due to unexpected consequences to PR #29622. It resulted in small operations being moved for thread execution unecessarily. This PR addresses it by looking a little more carefully before doing so, and keeping input embeding together whicle still letting mixed models do whatever they need so that #29622 still works as well.

Should address #30018 and #30033

Additional information

Replaces PR #30151 from my personal account.

Should address #30018 and #30033 not sure if any others.

The following shows the automated benchmark results from testing this PR:

RTX 4070 Laptop GPU, CUDA, Q4_K_M models, llama-bench -ngl 99 -p 512 -n 128 -r 5, with -t as the thread count below. Values are tokens per second.

model threads test master master + change b11399
Qwen3-VL-2B 8 tg128 166.2 182.9 183.8
Qwen3-VL-2B 32 tg128 153.8 181.0 183.8
Qwen3-VL-2B 8 pp512 10229 11548 11851
Qwen3-VL-2B 32 pp512 10534 11666 12398
Gemma 4 E2B 8 tg128 121.4 130.9 131.4
Gemma 4 E2B 32 tg128 114.9 130.8 131.4
Gemma 4 E2B 8 pp512 7918 8085 7885
Gemma 4 E2B 32 pp512 8063 7709 7988

master is upstream 03aa006. b11399 is the release before PR 29622. pp512 varies about 3 percent from run to run.

Requirements

Generative tools used to diagnose the problem and for producing the benchmarks to determine if this PR actually did fix the problem.

UPDATE: reedited again to try to clarify.

This PR is address performance degredation due to unexpected consequences to
PR ggml-org#29622. It results in moving moving small operations for thread execution unecessarily,
which this PR addresses by looking a little more carefully before doing so, so that the
mixed models that ggml-org#29622 is intended to support still work as well.

Should address ggml-org#30018 and ggml-org#30033

Signed-off-by: deadprogram <ron@hybridgroup.com>
@deadprogram
deadprogram requested a review from ggerganov as a code owner October 8, 2026 10:59
@deadprogram
deadprogram marked this pull request as draft October 8, 2026 13:06
@deadprogram
deadprogram marked this pull request as ready for review October 8, 2026 13:08
@deadprogram deadprogram changed the title llama : fix KV being cleared during context shift llama : keep input embeddings on layer 0 preventing unneeded scheduling thread wakeups Oct 8, 2026
@deadprogram

deadprogram commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

I had to change the title for this PR to correctly match what I was trying to do here. Thanks for looking!

@deadprogram
deadprogram marked this pull request as draft October 8, 2026 14:09
@ggerganov

Copy link
Copy Markdown
Member

Thanks for the report, but it's not the correct fix. Please test #30160

@deadprogram

Copy link
Copy Markdown
Contributor Author

Testing #30160 now thanks for quick response @ggerganov now closing this PR.

@deadprogram deadprogram closed this Oct 8, 2026
@deadprogram
deadprogram deleted the fix-kv-clear-context branch October 8, 2026 14:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants