Skip to content

ggml-openvino: fix CI tests; fix GPU regressions. - #30037

Merged
ggerganov merged 10 commits into
ggml-org:masterfrom
ravi9:pr-2026-10-05
Oct 6, 2026
Merged

ggerganov merged 10 commits into
ggml-org:masterfrom
ravi9:pr-2026-10-05

Conversation

@ravi9

@ravi9 ravi9 commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Overview

Fix CI tests and resolve GPU stateful-execution issues affecting Gemma and Qwen3.5.

  • Skip unselected forward-graph branches, support DUP, and make inp_scale_rows token dimension dynamic. Handles the recent mixed token/embd input .
  • Fix Qwen3.5 stateful execution after the build_rs change.
  • Update backend op tests: create FILL constants in the output type, reject quantized CONCAT, and skip GPU MUL_MAT cases with unbound Q4_1/Q4_K weights that fail GPU compilation with u4 weights and an f16 zero point.
  • Add GPU-only rank-alignment and MoE fusion fixes for stateful graphs. The rank-alignment pass skips an operand that is an RMS norm output (for Gemma 3).
  • Update docs with validated models.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes. For debugging, testing, reviewing, and fixing bugs.

ravi9 and others added 10 commits October 6, 2026 08:19
Upstream ggml-org#29622 adds a mixed token/embd branch to every input
embedding graph through ggml_build_forward_select(). Its nodes are
not flagged for compute, but the backend translated them anyway,
and the DUP in that branch was unsupported, so the scheduler split
the graph and passed the embeddings across the split with a fixed
token count. The first single-token decode then failed
(test-thread-safety on CPU and GPU).

Build the OV model from the compute nodes only, and translate a
same-type contiguous DUP like CONT so the graph stays on one backend.
ggml-org#29622 also moves the per-token embedding scale (gemma3, gemma3n,
gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim
and pad it per chunk on the static (NPU) path.
Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU
plugin fails to compile that form for some row counts with "clFinish,
error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on
the MUL_MAT cases added in ggml-org#29869 (e.g. m=1000, n=2, k=1024). Model
weights use a u4 zero point and are not affected.

Report these cases as unsupported on GPU until the plugin is fixed.
Op tests check support before allocating, so the check matches unbound
weights only; model loading probes with a dummy buffer and keeps its
weights on the GPU.
translate_fill always built an f32 constant, so an f16 FILL produced
f32 data and the copy back overran the f16 output buffer. Use the
output type for the constant.
Quantized inputs are dequantized when translated, so the backend cannot
write a quantized CONCAT output. Report it as unsupported, as for CPY
to a quantized type.
ggml-org#29856 changed build_rs to gather all recurrent states with one GET_ROWS
on the s_copy leaf and take the ubatch and extra states as views of it.
The stateful path matched only the previous form, a GET_ROWS per view of
s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU
("is_axis_valid(axis, r)" in a Concat).

For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the
active-state gather, keep the rank-4 layout of reshapes that read a view
of it, and map the copy of the empty extra-state view to the single-slot
remainder writeback. Do not warn about the dynamic dim of empty views.
…randRanks

The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract
whose operand ranks differ. In gemma-3 the lower-rank operand of the
post-attention residual add is the norm output, and unsqueezing it makes
the GPU plugin compute the layer wrongly: gemma-3 returns empty answers
on GPU with stateful execution.

Skip the rewrite when the lower-rank operand is an RMS norm output.
@ravi9
ravi9 requested a review from wine99 as a code owner October 6, 2026 03:20
@github-actions github-actions Bot added documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning OpenVINO labels Oct 6, 2026
@ggerganov
ggerganov merged commit b9a5a00 into ggml-org:master Oct 6, 2026
19 of 21 checks passed
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
* ggml-openvino: skip unselected graph branches and support DUP

Upstream ggml-org#29622 adds a mixed token/embd branch to every input
embedding graph through ggml_build_forward_select(). Its nodes are
not flagged for compute, but the backend translated them anyway,
and the DUP in that branch was unsupported, so the scheduler split
the graph and passed the embeddings across the split with a fixed
token count. The first single-token decode then failed
(test-thread-safety on CPU and GPU).

Build the OV model from the compute nodes only, and translate a
same-type contiguous DUP like CONT so the graph stays on one backend.

* ggml-openvino: make inp_scale_rows token dim dynamic

ggml-org#29622 also moves the per-token embedding scale (gemma3, gemma3n,
gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim
and pad it per chunk on the static (NPU) path.

* ggml-openvino: skip GPU MUL_MAT op tests with unbound Q4_1/Q4_K weights

Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU
plugin fails to compile that form for some row counts with "clFinish,
error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on
the MUL_MAT cases added in ggml-org#29869 (e.g. m=1000, n=2, k=1024). Model
weights use a u4 zero point and are not affected.

Report these cases as unsupported on GPU until the plugin is fixed.
Op tests check support before allocating, so the check matches unbound
weights only; model loading probes with a dummy buffer and keeps its
weights on the GPU.

* ggml-openvino: create FILL in the output type

translate_fill always built an f32 constant, so an f16 FILL produced
f32 data and the copy back overran the f16 output buffer. Use the
output type for the constant.

* ggml-openvino: reject CONCAT with a quantized type

Quantized inputs are dequantized when translated, so the backend cannot
write a quantized CONCAT output. Report it as unsupported, as for CPY
to a quantized type.

* ggml-openvino: handle the single recurrent state gather of build_rs

ggml-org#29856 changed build_rs to gather all recurrent states with one GET_ROWS
on the s_copy leaf and take the ubatch and extra states as views of it.
The stateful path matched only the previous form, a GET_ROWS per view of
s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU
("is_axis_valid(axis, r)" in a Concat).

For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the
active-state gather, keep the rank-4 layout of reshapes that read a view
of it, and map the copy of the empty extra-state view to the single-slot
remainder writeback. Do not warn about the dynamic dim of empty views.

* openvino: align eltwise operand ranks to work around a GPU-plugin defect

* openvino: match the MoE fusion on the rank-3 stateful graph

* ggml-openvino: do not unsqueeze an RMS norm output in AlignEltwiseOperandRanks

The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract
whose operand ranks differ. In gemma-3 the lower-rank operand of the
post-attention residual add is the norm output, and unsqueezing it makes
the GPU plugin compute the layer wrongly: gemma-3 returns empty answers
on GPU with stateful execution.

Skip the rewrite when the lower-rank operand is an RMS norm output.

* docs : update OpenVINO validated models

---------

Co-authored-by: Mustafa Cavus <mustafa.cavus@intel.com>
(cherry picked from commit b9a5a00)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning OpenVINO

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants