Skip to content

ggml-openvino: fix test-thread-safety, gemma, Qwen3.5 stateful and test-backend-ops failures - #338

Merged
cavusmustafa merged 6 commits into
dev_backend_openvinofrom
ov-fix-thread-safety
Oct 5, 2026
Merged

cavusmustafa merged 6 commits into
dev_backend_openvinofrom
ov-fix-thread-safety

Conversation

@ravi9

@ravi9 ravi9 commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Fixes OpenVINO backend failures caused by recent upstream changes:

Changes

  • Skip unselected graph branches and support DUP. llama: support both embd + raw tokens in batch ggml-org/llama.cpp#29622 adds a mixed token/embd branch to every input embedding graph through ggml_build_forward_select(). Its nodes are not flagged for compute, but the backend translated them, and the DUP in that branch was unsupported. The scheduler split the graph and passed the embeddings across the split with a fixed token count, so the first single-token decode failed:
    Can't set the input tensor with index: 15, because the model input (shape=[1,1,6,288]) and the tensor (shape=(1.1.1.288)) are incompatible.
    The OV model is now built from the compute nodes only, and a same-type contiguous DUP is translated like CONT. Both parts are needed: either one alone still fails.
  • Make inp_scale_rows token dim dynamic. llama: support both embd + raw tokens in batch ggml-org/llama.cpp#29622 also moves the per-token embedding scale (gemma3, gemma3n, gemma4) into a new [1, n_tokens] input. It had a fixed shape, which broke gemma-3 on every device. It now has a dynamic token dim and is padded per chunk on the static (NPU) path.
  • Skip GPU MUL_MAT op tests with unbound Q4_1/Q4_K weights. Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU plugin fails to compile that form for some row counts ([GPU] clFinish, error code: -5 CL_OUT_OF_RESOURCES). That aborts test-backend-ops on the MUL_MAT cases added in metal : few-row MMA mat-mul and batched copies for speculative decoding ggml-org/llama.cpp#29869. Model weights use a u4 zero point and are not affected. These cases are reported as unsupported on GPU until the plugin is fixed. The check matches only unbound weights, so model loading keeps Q4_1/Q4_K weights on the GPU.
  • Handle the single recurrent state gather of build_rs. graph: fix CI realloc abort by gathering the recurrent states once ggml-org/llama.cpp#29856 changed build_rs to gather all recurrent states with one GET_ROWS on the s_copy leaf and take the ubatch and extra states as views of it. The stateful path matched only the previous form, a GET_ROWS per view of s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU (is_axis_valid(axis, r)). For a single-slot cache, the single gather and its views now map to the existing single-slot cases.
  • Create FILL in the output type. The new f16 FILL cases from webgpu: add f16 support to fill/set_rows ggml-org/llama.cpp#29897 produced f32 data, which overran the f16 output buffer.
  • Reject CONCAT with a quantized type. The new q8_0 CONCAT case from metal : few-row MMA mat-mul and batched copies for speculative decoding ggml-org/llama.cpp#29869 was translated with dequantized inputs, so the quantized output was wrong.

Assisted-by: Claude

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. For debugging, review, fixes and testing.

ravi9 added 2 commits October 5, 2026 22:27
Upstream ggml-org#29622 adds a mixed token/embd branch to every input
embedding graph through ggml_build_forward_select(). Its nodes are
not flagged for compute, but the backend translated them anyway,
and the DUP in that branch was unsupported, so the scheduler split
the graph and passed the embeddings across the split with a fixed
token count. The first single-token decode then failed
(test-thread-safety on CPU and GPU).

Build the OV model from the compute nodes only, and translate a
same-type contiguous DUP like CONT so the graph stays on one backend.
ggml-org#29622 also moves the per-token embedding scale (gemma3, gemma3n,
gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim
and pad it per chunk on the static (NPU) path.
@ravi9 ravi9 changed the title ggml-openvino: handle the mixed token/embd input path from upstream #29622 ggml-openvino: fix CI, handle the mixed token/embd input path Oct 5, 2026
ravi9 added 3 commits October 5, 2026 23:56
Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU
plugin fails to compile that form for some row counts with "clFinish,
error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on
the MUL_MAT cases added in ggml-org#29869 (e.g. m=1000, n=2, k=1024). Model
weights use a u4 zero point and are not affected.

Report these cases as unsupported on GPU until the plugin is fixed.
Op tests check support before allocating, so the check matches unbound
weights only; model loading probes with a dummy buffer and keeps its
weights on the GPU.
translate_fill always built an f32 constant, so an f16 FILL produced
f32 data and the copy back overran the f16 output buffer. Use the
output type for the constant.
Quantized inputs are dequantized when translated, so the backend cannot
write a quantized CONCAT output. Report it as unsupported, as for CPY
to a quantized type.
@ravi9 ravi9 changed the title ggml-openvino: fix CI, handle the mixed token/embd input path ggml-openvino: fix failures from recent upstream graph and op-test changes Oct 5, 2026
@ravi9 ravi9 changed the title ggml-openvino: fix failures from recent upstream graph and op-test changes ggml-openvino: fix test-thread-safety, gemma and test-backend-ops failures Oct 5, 2026
ggml-org#29856 changed build_rs to gather all recurrent states with one GET_ROWS
on the s_copy leaf and take the ubatch and extra states as views of it.
The stateful path matched only the previous form, a GET_ROWS per view of
s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU
("is_axis_valid(axis, r)" in a Concat).

For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the
active-state gather, keep the rank-4 layout of reshapes that read a view
of it, and map the copy of the empty extra-state view to the single-slot
remainder writeback. Do not warn about the dynamic dim of empty views.
@ravi9 ravi9 changed the title ggml-openvino: fix test-thread-safety, gemma and test-backend-ops failures ggml-openvino: fix test-thread-safety, gemma, Qwen3.5 stateful and test-backend-ops failures Oct 5, 2026
@cavusmustafa
cavusmustafa merged commit 95af3f3 into dev_backend_openvino Oct 5, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants