Skip to content

ggml-openvino: derive attention sizes from the KQ mask shape - #339

Merged
cavusmustafa merged 4 commits into
ravi9:dev_backend_openvinofrom
osabnis:ov-dynamic-decode
Oct 9, 2026
Merged

cavusmustafa merged 4 commits into
ravi9:dev_backend_openvinofrom
osabnis:ov-dynamic-decode

Conversation

@osabnis

@osabnis osabnis commented Oct 5, 2026 •

Copy link
Copy Markdown

Overview

Faster decode for dynamic (stateless) graphs in the OpenVINO backend, primarily on GPU.

Previously, n_seq_active and attention_size(_swa) were passed as separate graph inputs. This PR derives them from the KQ mask shape using ShapeOf + Gather inside the model. The GPU plugin can then infer the Q reshape and KV slice shapes through shape inference instead of runtime input values. Q's head size also becomes static, allowing the plugin to select its fused SDPA kernel.

Two companion changes are required for the fused kernel to work correctly:

  • Clamp the GPU mask to -30000. A key block that is entirely -inf for a row produces NaNs in the fused kernel; this was observed with Gemma-4 prompts at context length 1024.
  • Reshape the KV slice to explicitly restate the head count and head size. A slice with a runtime end loses its dimensions in the GPU plugin, causing SDPA to fall back to the reference kernel. This resulted in incorrect output and CL_OUT_OF_RESOURCES errors for SmolLM2 and Phi-3.5.

Additional changes:

  • Skip the identity reshape in SET_ROWS.
  • Use a shared Split for one-element view slices in dynamic GPU graphs, such as Gemma-4's per-layer embeddings, instead of one StridedSlice per layer.

Both new settings are documented in docs/backend/OPENVINO.md and included in the graph and model cache keys:

  • GGML_OPENVINO_DISABLE_SHAPE_FROM_MASK=1: restore the previous separate-input behavior.

Additional information

Decode performance

llama-bench -fa 1 -p 0 -n 128 -d <depth>, Q4_K_M, Arc B390 (Panther Lake) GPU, OpenVINO master. Runs were performed at the console with a mean of two interleaved rounds.

Base: fork + #336. Test: base + this PR. Both on the same OpenVINO build.

Model Depth 512 1024 2048 4096
Gemma-4 E2B +15% +8% +13% +2%
SmolLM2 1.7B +24% +24% +24% +27%
Phi-3.5 mini +13% +15% +11% +7%
Llama-3.2 1B +13% +18% +14% +13%
Phi-4 mini +8% +9% +11% +14%
Gemma-3 4B +9% +9% +8% +10%
Qwen3.5 4B +2% +3% +3% +6%

Gemma-4 at depth 0: +12%. The Gemma-4 baseline varied by approximately 10% between rounds at depths 512 and 2048.

Accuracy

GPU results against CPU references, #336 alone vs #336 + this PR, both on the same OpenVINO build:

Check #336 #336 + this PR
Gemma-4 E2B KLD, decode, ctx 1024 0.363583 0.363413
Gemma-4 E2B KLD, decode, ctx 2048 0.357108 0.356962
Gemma-4 E2B KLD, prompt, ctx 1024 0.367275 0.367396
Gemma-4 E2B KLD, 4 sequences, ctx 512 0.324305 0.324295
PPL, 4 sequences x 512: Gemma-4 E2B 119.9494 119.9567
SmolLM2 1.7B 10.4481 10.4482
Phi-3.5 mini 5.9942 5.9946
Llama-3.2 1B 10.4543 10.4531
Phi-4 mini 7.5936 7.5859
Qwen3.5 4B 11.2862 11.2828

The differences are within run-to-run variation on this GPU. Gemma-3 4B fails the perplexity run on both builds
(a known f16 overflow, also on the branch today).

The speed and full accuracy runs used the fork at 308ccd80a with #336. This branch is based on 836d57176, where it builds and runs successfully and gives the same perplexity on SmolLM2, Llama-3.2, Phi-3.5, and Gemma-4 within run-to-run variation.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. AI ran the unit testing and analysis while I did the rest.

Dynamic CPU/GPU graphs previously passed n_seq_active and
attention_size(_swa) as separate inputs. Derive these values directly
from the KQ mask shape using ShapeOf + Gather inside the model. This
allows the GPU plugin to infer the Q reshape and KV slice shapes from
shape inference rather than runtime input values. Q's head size also
becomes static, enabling the fused SDPA kernel.

GGML_OPENVINO_DISABLE_SHAPE_FROM_MASK=1 restores the previous
separate-input behavior.

The fused SDPA path also requires two supporting changes:

- Clamp the GPU mask to -30000. A key block that is entirely -inf for
  a row produces NaNs in the fused kernel.
- Restate the head count and head size after the KV slice. The plugin
  loses all slice dimensions when the end index is dynamic, causing
  SDPA to fall back to the reference kernel.

Additional graph and quantization improvements:

- Skip the identity reshape of rows in SET_ROWS.
- Share a single Split for one-element view slices in dynamic GPU
  graphs (e.g. Gemma-4 per-layer embeddings), replacing one
  StridedSlice per layer.
- Add GGML_OPENVINO_REQUANT_EMBD=native (off by default) to preserve
  the file's native quantization for token_embd/output instead of
  requantizing to Q8_0_C.
- Include both environment settings in the graph and model cache keys
  and document them.
The dynamic GPU path reused a Split output for any single-element view
by deriving the chunk index from offset / stride. This produced
incorrect results for views that were not chunk-aligned or whose
remaining dimensions did not preserve the source strides, as seen in
test-backend-ops CONT with use_view_slice.

Restrict Split reuse to cases where the view exactly matches a split
chunk, meaning the offset is aligned to the split stride and the
remaining non-unit dimensions preserve the original source strides.
Gemma-4 per-layer embedding slices continue to use the shared Split
path.
@osabnis
osabnis force-pushed the ov-dynamic-decode branch from 9a78c56 to 0aa29c7 Compare October 6, 2026 21:23
Initialize the newly added `shape_source` and `shape_axis` fields in
the `ModelExtraInputInfo` brace initializer used by
`add_extra_inputs()`.

Without initializers for these fields, GCC emits
`-Wmissing-field-initializers`, causing CI to fail with
`LLAMA_FATAL_WARNINGS`.
@osabnis

osabnis commented Oct 9, 2026

Copy link
Copy Markdown
Author

Fixed the CI failure in the self-hosted OpenVINO job.

The GCC build stopped on ggml-decoder.cpp: the brace initializer in add_extra_inputs() did not set the two fields this PR adds to ModelExtraInputInfo (shape_source, shape_axis), and CI builds with -Werror, so -Wmissing-field-initializers became an error. MSVC does not report this warning, which is why the Windows builds passed.
The new commit sets both fields explicitly. No behavior change: they get the same values as before (empty and -1).

Checked with clang and the CI warning flags on all files this PR changes: the only warning was this one, and it is gone with the fix.

Comment thread ggml/src/ggml-openvino/ggml-openvino-extra.cpp Outdated
The option identifies embedding and output weights by their tensor
names, which are defined by core llama.cpp. Remove it from this PR;
preserving the file quantization of these weights will be addressed in
a separate change that identifies them by their roles in the graph.
@cavusmustafa
cavusmustafa merged commit ab759cf into ravi9:dev_backend_openvino Oct 9, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants