Repository navigation
ggml-openvino: fix test-thread-safety, gemma, Qwen3.5 stateful and test-backend-ops failures - #338
Merged
Merged
Conversation
Upstream ggml-org#29622 adds a mixed token/embd branch to every input embedding graph through ggml_build_forward_select(). Its nodes are not flagged for compute, but the backend translated them anyway, and the DUP in that branch was unsupported, so the scheduler split the graph and passed the embeddings across the split with a fixed token count. The first single-token decode then failed (test-thread-safety on CPU and GPU). Build the OV model from the compute nodes only, and translate a same-type contiguous DUP like CONT so the graph stays on one backend.
ggml-org#29622 also moves the per-token embedding scale (gemma3, gemma3n, gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim and pad it per chunk on the static (NPU) path.
Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU plugin fails to compile that form for some row counts with "clFinish, error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on the MUL_MAT cases added in ggml-org#29869 (e.g. m=1000, n=2, k=1024). Model weights use a u4 zero point and are not affected. Report these cases as unsupported on GPU until the plugin is fixed. Op tests check support before allocating, so the check matches unbound weights only; model loading probes with a dummy buffer and keeps its weights on the GPU.
translate_fill always built an f32 constant, so an f16 FILL produced f32 data and the copy back overran the f16 output buffer. Use the output type for the constant.
Quantized inputs are dequantized when translated, so the backend cannot write a quantized CONCAT output. Report it as unsupported, as for CPY to a quantized type.
ggml-org#29856 changed build_rs to gather all recurrent states with one GET_ROWS on the s_copy leaf and take the ubatch and extra states as views of it. The stateful path matched only the previous form, a GET_ROWS per view of s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU ("is_axis_valid(axis, r)" in a Concat). For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the active-state gather, keep the rank-4 layout of reshapes that read a view of it, and map the copy of the empty extra-state view to the single-slot remainder writeback. Do not warn about the dynamic dim of empty views.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes OpenVINO backend failures caused by recent upstream changes:
0bb496dbd)a3a1c4747)11fe02151)build_rsfrom graph: fix CI realloc abort by gathering the recurrent states once ggml-org/llama.cpp#29856 (436f6f89e)Changes
ggml_build_forward_select(). Its nodes are not flagged for compute, but the backend translated them, and the DUP in that branch was unsupported. The scheduler split the graph and passed the embeddings across the split with a fixed token count, so the first single-token decode failed:Can't set the input tensor with index: 15, because the model input (shape=[1,1,6,288]) and the tensor (shape=(1.1.1.288)) are incompatible.The OV model is now built from the compute nodes only, and a same-type contiguous DUP is translated like CONT. Both parts are needed: either one alone still fails.
inp_scale_rowstoken dim dynamic. llama: support both embd + raw tokens in batch ggml-org/llama.cpp#29622 also moves the per-token embedding scale (gemma3, gemma3n, gemma4) into a new[1, n_tokens]input. It had a fixed shape, which broke gemma-3 on every device. It now has a dynamic token dim and is padded per chunk on the static (NPU) path.[GPU] clFinish, error code: -5 CL_OUT_OF_RESOURCES). That aborts test-backend-ops on the MUL_MAT cases added in metal : few-row MMA mat-mul and batched copies for speculative decoding ggml-org/llama.cpp#29869. Model weights use a u4 zero point and are not affected. These cases are reported as unsupported on GPU until the plugin is fixed. The check matches only unbound weights, so model loading keeps Q4_1/Q4_K weights on the GPU.build_rs. graph: fix CI realloc abort by gathering the recurrent states once ggml-org/llama.cpp#29856 changedbuild_rsto gather all recurrent states with one GET_ROWS on the s_copy leaf and take the ubatch and extra states as views of it. The stateful path matched only the previous form, a GET_ROWS per view of s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU (is_axis_valid(axis, r)). For a single-slot cache, the single gather and its views now map to the existing single-slot cases.Assisted-by: Claude
Requirements