Repository navigation
ggml-openvino: fix CI tests; fix GPU regressions. - #30037
Merged
Merged
Conversation
Upstream ggml-org#29622 adds a mixed token/embd branch to every input embedding graph through ggml_build_forward_select(). Its nodes are not flagged for compute, but the backend translated them anyway, and the DUP in that branch was unsupported, so the scheduler split the graph and passed the embeddings across the split with a fixed token count. The first single-token decode then failed (test-thread-safety on CPU and GPU). Build the OV model from the compute nodes only, and translate a same-type contiguous DUP like CONT so the graph stays on one backend.
ggml-org#29622 also moves the per-token embedding scale (gemma3, gemma3n, gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim and pad it per chunk on the static (NPU) path.
Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU plugin fails to compile that form for some row counts with "clFinish, error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on the MUL_MAT cases added in ggml-org#29869 (e.g. m=1000, n=2, k=1024). Model weights use a u4 zero point and are not affected. Report these cases as unsupported on GPU until the plugin is fixed. Op tests check support before allocating, so the check matches unbound weights only; model loading probes with a dummy buffer and keeps its weights on the GPU.
translate_fill always built an f32 constant, so an f16 FILL produced f32 data and the copy back overran the f16 output buffer. Use the output type for the constant.
Quantized inputs are dequantized when translated, so the backend cannot write a quantized CONCAT output. Report it as unsupported, as for CPY to a quantized type.
ggml-org#29856 changed build_rs to gather all recurrent states with one GET_ROWS on the s_copy leaf and take the ubatch and extra states as views of it. The stateful path matched only the previous form, a GET_ROWS per view of s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU ("is_axis_valid(axis, r)" in a Concat). For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the active-state gather, keep the rank-4 layout of reshapes that read a view of it, and map the copy of the empty extra-state view to the single-slot remainder writeback. Do not warn about the dynamic dim of empty views.
…randRanks The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract whose operand ranks differ. In gemma-3 the lower-rank operand of the post-attention residual add is the norm output, and unsqueezing it makes the GPU plugin compute the layer wrongly: gemma-3 returns empty answers on GPU with stateful execution. Skip the rewrite when the lower-rank operand is an RMS norm output.
CISC
approved these changes
Oct 6, 2026
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 8, 2026
* ggml-openvino: skip unselected graph branches and support DUP Upstream ggml-org#29622 adds a mixed token/embd branch to every input embedding graph through ggml_build_forward_select(). Its nodes are not flagged for compute, but the backend translated them anyway, and the DUP in that branch was unsupported, so the scheduler split the graph and passed the embeddings across the split with a fixed token count. The first single-token decode then failed (test-thread-safety on CPU and GPU). Build the OV model from the compute nodes only, and translate a same-type contiguous DUP like CONT so the graph stays on one backend. * ggml-openvino: make inp_scale_rows token dim dynamic ggml-org#29622 also moves the per-token embedding scale (gemma3, gemma3n, gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim and pad it per chunk on the static (NPU) path. * ggml-openvino: skip GPU MUL_MAT op tests with unbound Q4_1/Q4_K weights Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU plugin fails to compile that form for some row counts with "clFinish, error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on the MUL_MAT cases added in ggml-org#29869 (e.g. m=1000, n=2, k=1024). Model weights use a u4 zero point and are not affected. Report these cases as unsupported on GPU until the plugin is fixed. Op tests check support before allocating, so the check matches unbound weights only; model loading probes with a dummy buffer and keeps its weights on the GPU. * ggml-openvino: create FILL in the output type translate_fill always built an f32 constant, so an f16 FILL produced f32 data and the copy back overran the f16 output buffer. Use the output type for the constant. * ggml-openvino: reject CONCAT with a quantized type Quantized inputs are dequantized when translated, so the backend cannot write a quantized CONCAT output. Report it as unsupported, as for CPY to a quantized type. * ggml-openvino: handle the single recurrent state gather of build_rs ggml-org#29856 changed build_rs to gather all recurrent states with one GET_ROWS on the s_copy leaf and take the ubatch and extra states as views of it. The stateful path matched only the previous form, a GET_ROWS per view of s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU ("is_axis_valid(axis, r)" in a Concat). For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the active-state gather, keep the rank-4 layout of reshapes that read a view of it, and map the copy of the empty extra-state view to the single-slot remainder writeback. Do not warn about the dynamic dim of empty views. * openvino: align eltwise operand ranks to work around a GPU-plugin defect * openvino: match the MoE fusion on the rank-3 stateful graph * ggml-openvino: do not unsqueeze an RMS norm output in AlignEltwiseOperandRanks The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract whose operand ranks differ. In gemma-3 the lower-rank operand of the post-attention residual add is the norm output, and unsqueezing it makes the GPU plugin compute the layer wrongly: gemma-3 returns empty answers on GPU with stateful execution. Skip the rewrite when the lower-rank operand is an RMS norm output. * docs : update OpenVINO validated models --------- Co-authored-by: Mustafa Cavus <mustafa.cavus@intel.com> (cherry picked from commit b9a5a00)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Fix CI tests and resolve GPU stateful-execution issues affecting Gemma and Qwen3.5.
DUP, and makeinp_scale_rowstoken dimension dynamic. Handles the recent mixed token/embd input .build_rschange.FILLconstants in the output type, reject quantizedCONCAT, and skip GPUMUL_MATcases with unbound Q4_1/Q4_K weights that fail GPU compilation with u4 weights and an f16 zero point.Requirements