Repository navigation
ggml-openvino : make op support a property of the registry - #307
Draft
cavusmustafa wants to merge 19 commits into
Draft
cavusmustafa wants to merge 19 commits into
cavusmustafa wants to merge 19 commits into
Conversation
cavusmustafa
force-pushed
the
ov-supports-op-robust
branch
from
September 2, 2026 23:43
b800043 to
abf6caf
Compare
ravi9
force-pushed
the
dev_backend_openvino
branch
2 times, most recently
from
September 3, 2026 21:10
fd9bc04 to
33237ab
Compare
cavusmustafa
force-pushed
the
ov-supports-op-robust
branch
from
September 4, 2026 23:47
abf6caf to
d50cef8
Compare
ravi9
force-pushed
the
dev_backend_openvino
branch
from
September 14, 2026 17:49
37b4e1d to
926ef68
Compare
wine99
force-pushed
the
dev_backend_openvino
branch
from
September 16, 2026 08:07
0365b00 to
37b53fd
Compare
ravi9
force-pushed
the
dev_backend_openvino
branch
from
September 17, 2026 17:37
40a2a96 to
972d231
Compare
ravi9
force-pushed
the
dev_backend_openvino
branch
2 times, most recently
from
September 30, 2026 22:54
6b08f8a to
ab07546
Compare
cavusmustafa
force-pushed
the
ov-supports-op-robust
branch
2 times, most recently
from
October 1, 2026 00:44
5908ba4 to
c85beef
Compare
Squash of ravi9#312: - ggml-openvino: add detailed inference profiling (Yu, Zijun) - ggml-openvino: use remote output tensors by default (Yu, Zijun) - ggml-openvino: optimize single-sequence recurrent state (Yu, Zijun) - opt1: remove recurrent reset for single sequence, opt2: direct gdn outputs (break parallel sequence) (Yu, Zijun) - fix parallel sequences (Yu, Zijun) - ggml-openvino: simplify graph cache key (ynimmaga) - enable stateful for qwen35 single sequence (Yu, Zijun) - Fix after rebasing (Yu, Zijun) - Add k-requant option q4_asym64 (Yu, Zijun) - Fix qwen35 llama-bench -p 0 (Yu, Zijun) - Simplify RESHAPE translation (Yu, Zijun) - openvino: fuse MoE routing (Yu, Zijun) - openvino: fuse GDN qk normalization (Yu, Zijun) - openvino: enable GPU MoE fusion by default (Yu, Zijun) - ggml-openvino: add cache_only mode to import cached compiled model on disk directly (Yu, Zijun) - openvino : report the device allocation limit to ggml (Łukasz Ślusarczyk) - Fix windows build (Yu, Zijun) Co-authored-by: ynimmaga <ynimmaga@users.noreply.github.com> Co-authored-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com>
- Only the device selected by GGML_OPENVINO_DEVICE reports as GPU; the other OpenVINO devices report as IGPU so llama.cpp does not offload to them. Initializing a non-selected device logs a warning. - Name devices OPENVINO<i> again and show the OpenVINO id in the description. Raw "CPU" names shadowed the ggml CPU backend. - Support GPU.N: create the OpenCL queue on OpenVINO's own context for the selected device, and replace "GPU"/"NPU" string comparisons with ggml_openvino_is_gpu()/ggml_openvino_is_npu(). - An unavailable GGML_OPENVINO_DEVICE is now an error that lists the available devices, instead of silently falling back to CPU. - Memory: cap iGPU/NPU free memory at system available memory, fall back to system memory instead of 0/0 when the plugin lacks memory properties, and ignore host USM allocations in GPU usage. - Initialize the device config once under a lock, even if OpenCL setup fails. - Fix supports_op return type for non-selected devices (build error).
clGetExtensionFunctionAddressForPlatform was called on the first platform returned by clGetPlatformIDs. The address it returns is only valid for the platform it was queried on, and the first platform is not always the one that holds the device OpenVINO selected. On a host whose first platform comes from another vendor the lookup returns null, and then every read, write and memset on a GPU buffer fails with "clEnqueueMemcpyINTEL not available". Look both entry points up in init(), on the platform of the device OpenVINO picked, and keep them in the device config next to the command queue. Assisted-by: Claude Opus 5
FuseMoeCompressed only matches models whose gate and up projections are separate GatherMatmul ops. gemma-4 packs both into one expert weight and splits the result after the GEMM, so its MoE block stayed unfused and ran the expert GEMMs as per-token GEMVs. Add FuseMoeCompressedFusedGateUp, which matches that shape (one GatherMatmul -> Slice/Slice -> Gelu(ERF) -> Multiply) and folds it into the same MOECompressed op, using GEMM3_SWIGLU with GEGLU_ERF. The fused weight, scale and zero point are split into gate/up halves by copying raw bytes, since a graph Slice would be rewritten to StridedSlice and constant folded, whose reference evaluator crashes on sub-byte types. gemma-4 also applies a per-expert output scale to the down projection before the router weights. MOECompressed takes only one per-expert weight, so that scale is folded into the routing weights, which is exact. The op reads the zero point straight off a weight port and needs an integer Constant there, so the matcher requires one and leaves natively quantized experts (exact f16 zp) to the unfused path. gemma-4-26B-A4B on Arc B390, GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all, llama-bench -p 512 -n 128 -r 2, against a GGML_OPENVINO_MOE_OP=0 baseline: pp512 66.16 -> 1608.73 t/s, tg128 25.94 -> 26.46 t/s. Perplexity over 12 chunks is unchanged (1451.3 +/- 177.9 unfused vs 1427.6 +/- 175.1 fused). No effect without that requant option, on models with separate gate/up weights, or on CPU. test-backend-ops -b OPENVINO0 is unchanged by this commit: two MUL_MAT_ID m_v cases fail, the same two on the unmodified base.
Stateful execution drops the leading size-1 batch dim, so OV tensors are rank
3 while GgmlOvDecoder::get_shape/get_stride still report GGML_MAX_DIMS=4
reversed entries. Several MoE ops derive OV axis indices straight from that
metadata, so they picked the wrong axis. A MoE model with
GGML_OPENVINO_STATEFUL_EXECUTION=1 aborts while building the graph:
Check 'is_axis_valid(axis, r)' failed at src/core/src/validation_util.cpp:336
While validating node 'opset11::TopK ... _ffn_moe_probs ...'
Axis 3 out of the tensor rank range [-3, 2].
Fix idiom throughout: take the axis from the real OV rank, or shift a
metadata-derived axis down by metadata_rank - actual_rank.
argsort.cpp the router top-k axis is 2 on rank 3, not 3. This is the
abort quoted above.
add.cpp the MoE expert-sum bypass collapses the 8-ADD chain into one
ReduceSum on hardcoded axis 2, which on rank 3 reduces n_embd
instead of the expert axis. Now rank-2, with the following
Unsqueeze at rank-3.
get_rows.cpp squeezing a hardcoded {0,1} also strips the batch dim
whenever it is 1, which is every decode step. Squeeze down to
the trailing two dims instead.
mul_mat_id.cpp pick the reshape dims by actual rank, and skip the trailing
Unsqueeze that re-adds the batch dim.
view.cpp the expert-plane slice had the Slice axis, dst_ov_axis, the
ShapeOf+Gather index and the Reshape target all rank-4.
utils.cpp process_view_input_new's "translate_view already resolved
this VIEW, skip re-slicing" shortcut required equal ranks. 4
vs 3 never matched, so every resolved expert plane got
re-sliced. Now compares the common trailing dims. Same axis
shift for the Slice in the view-chain walker.
Stateless is unchanged by construction: every edit is gated on the actual
rank, so axis_shift == 0 reproduces the previous code exactly. Checked on
OV-CPU by diffing greedy output against the unmodified base for dense
gemma-4-E2B, granite-1b-a400m and gemma-4-26B-A4B; all identical.
granite-1b-a400m on OV-CPU aborts with the error above before this change;
after it, it generates and is byte-identical to stateless. Dense gemma-4-E2B
is identical stateless vs stateful both before and after. test-backend-ops
-b OPENVINO0 is unchanged: two pre-existing MUL_MAT_ID m_v cases fail, the
same two on the unmodified base.
gemma-4-26B-A4B is a poor correctness vehicle here. On OV it already drifts
into degenerate repetition a few tokens in, in stateless as much as stateful,
and the two modes diverge somewhere inside that degenerate region instead of
matching token for token. Each mode is self-reproducible across runs.
Known limitation: FuseMoeCompressedFusedGateUp does not match the rank-3
graph, so a MoE model run with GGML_OPENVINO_STATEFUL_EXECUTION=1 loses the
prefill fusion while gaining decode. gemma-4-26B-A4B on Arc B390,
GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all, llama-bench -p 512 -n 128 -r 2:
unfused (GGML_OPENVINO_MOE_OP=0) pp512 66.16 tg128 25.94
fused, stateless (default) pp512 1608.73 tg128 26.46
fused, stateful pp512 66.18 tg128 29.91
Stateful is opt-in and off by default, and MoE did not run there at all
before this, so nothing that previously worked regresses. Making the pass
match rank 3 is the follow-up.
init() logged the error and returned, which left the device name a GPU but remote_context empty. The remote buffer and tensor paths assert only on the device being a GPU and then dereference that empty optional. Those paths have no host fallback, and a device that OpenVINO listed should have a working OpenCL context, so stop instead of continuing. An OpenCL stack that is broken as a whole is still caught earlier by the device availability check, which falls back to CPU. Assisted-by: Claude Opus 5
The single-argument form of the OpenVINO RTTI macros is the intended one, but their selector macro leaves __VA_ARGS__ empty, which -Wpedantic reports on every pass and op header. Turn that warning off for this backend only, the way ggml-cuda and ggml-sycl already do for their own third-party warnings. Also drop a break and a dead assignment around a GGML_ABORT, which is noreturn. Assisted-by: Claude Opus 5
HARDSIGMOID used a 1/6 constant in the input type, which is not exact in bf16, and EXPM1 lost precision for small inputs in f16. Both now compute in f32 and convert back, except on NPU where the f32 path gives wrong results. Fixes the HARDSIGMOID/EXPM1 test-backend-ops failures on GPU.
Show the selecting GGML_OPENVINO_DEVICE value and active device in --list-devices, startup logs, and backend tests. Clarify OpenVINO selection uses GGML_OPENVINO_DEVICE, not -dev.
A remote buffer exists only on a GPU device, and init() aborts there if the queue cannot be created, so the queue is never null at these call sites. Assisted-by: Claude Opus 5
ravi9
force-pushed
the
dev_backend_openvino
branch
from
October 1, 2026 15:52
ff61a56 to
5014571
Compare
cavusmustafa
force-pushed
the
ov-supports-op-robust
branch
from
October 1, 2026 17:38
c85beef to
3c2cba9
Compare
ravi9
force-pushed
the
dev_backend_openvino
branch
3 times, most recently
from
October 3, 2026 19:48
308ccd8 to
836d571
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Created this PR as an initial draft for suports_op changes. CIs and supported models should be tested. Below is the AI generated description:
supports_op() could accept a node the backend then failed to translate. The conditions lived in a switch in ggml-openvino.cpp that 19 of the 54 registered ops never reached, so those ops were accepted unchecked and failed later inside the translator - ARGSORT on an unknown sort order, TRI on an out-of-range triangle type, PAD on an empty input.
Move the per-op conditions into openvino/op_support.{h,cpp} and give every registry entry a support rule beside its translator:
The constructor is deliberate. Without it OpEntry is an aggregate and {translate_foo} compiles with supports silently null, which would defeat the only guarantee the type exists to provide. A registry entry now cannot be added without stating when the op may be used, and that is checked by the compiler rather than by review.
The 20 case groups of is_op_supported_case() are migrated unchanged, so behaviour is identical: test-backend-ops -b OPENVINO0 gives 3386/3386, 17783 not supported, 0 FAIL, exit 0, matching the base exactly. New rules for ARGSORT, PAD and TRI turn three translator throws into gate declines.
Two categories in op_support.cpp are marked for reviewers because they are not statements about what an op means: declines that depend on the device name, which work around specific plugin defects and should be deleted when those are fixed, and declines keyed on a tensor name or an exact test shape.
Also: ggml-openvino.cpp loses 427 lines and gains 36; the empty ops_not_support_view_input set was dead code and is removed.
Note GGML_OP_SSM_CONV's rule has its "return true" commented out, so its stated intent of keeping the op on CPU is not implemented. Migrated as-is rather than changed.
Assisted-by: Claude Opus 5