Skip to content

ggml-openvino: fix gemma-3 on GPU with stateful execution - #340

Merged
ravi9 merged 2 commits into
dev_backend_openvinofrom
ov-fix-eltwise-align-gemma3
Oct 6, 2026
Merged

ravi9 merged 2 commits into
dev_backend_openvinofrom
ov-fix-eltwise-align-gemma3

Conversation

@ravi9

@ravi9 ravi9 commented Oct 6, 2026

Copy link
Copy Markdown
Owner

#337 added AlignEltwiseOperandRanks, which unsqueezes the lower-rank operand of an Add/Multiply/Subtract whose operand ranks differ. In gemma-3 that operand is the RMS norm output in the post-attention residual add, and the GPU plugin then computes the layer wrongly: gemma-3 returns empty answers on GPU with stateful execution.

The pass now skips the rewrite when the lower-rank operand is an RMS norm output. Qwen3.5 and gemma-4, which the pass fixes, are not affected.

Also updates the validated models table: Qwen3.5 and gemma-4-E2B now pass on GPU with stateful execution.

ravi9 added 2 commits October 6, 2026 06:56
…randRanks

The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract
whose operand ranks differ. In gemma-3 the lower-rank operand of the
post-attention residual add is the norm output, and unsqueezing it makes
the GPU plugin compute the layer wrongly: gemma-3 returns empty answers
on GPU with stateful execution.

Skip the rewrite when the lower-rank operand is an RMS norm output.
@ravi9
ravi9 merged commit 654b580 into dev_backend_openvino Oct 6, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant