Skip to content

ggml-webgpu: Fix some binding alias issues to support all archs, fix recurrent-state-rollback test - #25931

Merged
ggerganov merged 22 commits into
ggml-org:masterfrom
reeselevine:webgpu-all-archs
Jul 28, 2026
Merged

ggerganov merged 22 commits into
ggml-org:masterfrom
reeselevine:webgpu-all-archs

Conversation

@reeselevine

@reeselevine reeselevine commented Jul 20, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Enable more WebGPU architecture coverage by removing the blanket test skip, while keeping deepseek32 disabled for now. Fix binding aliasing issues in GLU and SSM scan, and handle zero-sized concat inputs.

Fixes the WebGPU CI failures as well.

Requirements

@reeselevine
reeselevine requested review from a team and JohannesGaessler as code owners July 20, 2026 17:19
@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning WebGPU labels Jul 20, 2026
@reeselevine
reeselevine marked this pull request as draft July 20, 2026 18:01
@reeselevine
reeselevine marked this pull request as ready for review July 22, 2026 00:56
@reeselevine reeselevine changed the title ggml-webgpu: Add overlap glu variant to support all archs, fix recurrent-state-rollback test ggml-webgpu: Fix some binding alias issues to support all archs, fix recurrent-state-rollback test Jul 22, 2026
Comment thread ggml/src/ggml-webgpu/ggml-webgpu.cpp Outdated
@yomaytk

yomaytk commented Jul 26, 2026

Copy link
Copy Markdown
Member

The test-llama-archs CI test is failing. Does it still need to be addressed?
https://github.com/ggml-org/llama.cpp/actions/runs/30134708680/job/89616184248?pr=25931

@reeselevine

Copy link
Copy Markdown
Contributor Author

It looks like one other model has issues, so I added that to the skip list for now.

@fairydreaming

Copy link
Copy Markdown
Contributor

@reeselevine Looks like I doubled this work in #26146 yesterday not knowing that this PR exists. Can you summarize your findings about DeepSeek V3.2 failures on macos (I see that it's still disabled in your PR)?

@yomaytk yomaytk left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The CI is green, nice! The skipped model can be addressed in the future.
@fairydreaming, thanks for also working on this.

@reeselevine reeselevine added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Jul 28, 2026
@ggerganov

Copy link
Copy Markdown
Member

@fairydreaming Ok to merge, or would you like to discuss the DS32 furhter?

@fairydreaming

Copy link
Copy Markdown
Contributor

@fairydreaming Ok to merge, or would you like to discuss the DS32 furhter?

@ggerganov OK to merge, I was just curious. DeepSeek V3.2 is currently still disabled in test-llama-archs on WebGPU so shouldn't cause problems.

@reeselevine

Copy link
Copy Markdown
Contributor Author

Oh yeah sorry for not responding @fairydreaming. From my understanding it was a combination of the CI node not supporting WebGPU flash attention but using the lightning indexer from that model. It seemed like potentially a bug in the tensor scheduling side of llama.cpp but I didn’t have time to investigate more deeply yet.

@ggerganov

ggerganov commented Jul 28, 2026 •

Copy link
Copy Markdown
Member

From my understanding it was a combination of the CI node not supporting WebGPU flash attention but using the lightning indexer from that model.

There is some chance that this has been fixed with #25832. (though unlikely)

@ggerganov
ggerganov merged commit bc71c24 into ggml-org:master Jul 28, 2026
27 of 29 checks passed
huaxel pushed a commit to huaxel/CachyLLama that referenced this pull request Aug 2, 2026
…recurrent-state-rollback test (ggml-org#25931)

* Add overlap glu variant to support all archs, fix recurrent-state-rollback test

* format

* Fix all arch overlapped ranges

* format

* diagnose bus error on apple ci

* More testing

* more testing

* more targeted testing

* Fix bug in alignment for > 4gb buffer offsets

* Fix bug in view offsets

* Try avoiding multi_buffers

* not fixed yet, more logging :(

* Handle edge case in set_rows

* Try looking at view source

* Skip deepseek32 for now and clean up trace infrastructure

* simplify skipping

* last cleanup

* actually final cleanup

* update handling of overlap

* format

* try skipping other failing model
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
…recurrent-state-rollback test (ggml-org#25931)

* Add overlap glu variant to support all archs, fix recurrent-state-rollback test

* format

* Fix all arch overlapped ranges

* format

* diagnose bus error on apple ci

* More testing

* more testing

* more targeted testing

* Fix bug in alignment for > 4gb buffer offsets

* Fix bug in view offsets

* Try avoiding multi_buffers

* not fixed yet, more logging :(

* Handle edge case in set_rows

* Try looking at view source

* Skip deepseek32 for now and clean up trace infrastructure

* simplify skipping

* last cleanup

* actually final cleanup

* update handling of overlap

* format

* try skipping other failing model
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
…recurrent-state-rollback test (ggml-org#25931)

* Add overlap glu variant to support all archs, fix recurrent-state-rollback test

* format

* Fix all arch overlapped ranges

* format

* diagnose bus error on apple ci

* More testing

* more testing

* more targeted testing

* Fix bug in alignment for > 4gb buffer offsets

* Fix bug in view offsets

* Try avoiding multi_buffers

* not fixed yet, more logging :(

* Handle edge case in set_rows

* Try looking at view source

* Skip deepseek32 for now and clean up trace infrastructure

* simplify skipping

* last cleanup

* actually final cleanup

* update handling of overlap

* format

* try skipping other failing model
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
…recurrent-state-rollback test (ggml-org#25931)

* Add overlap glu variant to support all archs, fix recurrent-state-rollback test

* format

* Fix all arch overlapped ranges

* format

* diagnose bus error on apple ci

* More testing

* more testing

* more targeted testing

* Fix bug in alignment for > 4gb buffer offsets

* Fix bug in view offsets

* Try avoiding multi_buffers

* not fixed yet, more logging :(

* Handle edge case in set_rows

* Try looking at view source

* Skip deepseek32 for now and clean up trace infrastructure

* simplify skipping

* last cleanup

* actually final cleanup

* update handling of overlap

* format

* try skipping other failing model
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…recurrent-state-rollback test (ggml-org#25931)

* Add overlap glu variant to support all archs, fix recurrent-state-rollback test

* format

* Fix all arch overlapped ranges

* format

* diagnose bus error on apple ci

* More testing

* more testing

* more targeted testing

* Fix bug in alignment for > 4gb buffer offsets

* Fix bug in view offsets

* Try avoiding multi_buffers

* not fixed yet, more logging :(

* Handle edge case in set_rows

* Try looking at view source

* Skip deepseek32 for now and clean up trace infrastructure

* simplify skipping

* last cleanup

* actually final cleanup

* update handling of overlap

* format

* try skipping other failing model
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…recurrent-state-rollback test (ggml-org#25931)

* Add overlap glu variant to support all archs, fix recurrent-state-rollback test

* format

* Fix all arch overlapped ranges

* format

* diagnose bus error on apple ci

* More testing

* more testing

* more targeted testing

* Fix bug in alignment for > 4gb buffer offsets

* Fix bug in view offsets

* Try avoiding multi_buffers

* not fixed yet, more logging :(

* Handle edge case in set_rows

* Try looking at view source

* Skip deepseek32 for now and clean up trace infrastructure

* simplify skipping

* last cleanup

* actually final cleanup

* update handling of overlap

* format

* try skipping other failing model
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. testing Everything test related WebGPU

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants