Repository navigation
ggml-openvino: derive attention sizes from the KQ mask shape - #339
Merged
cavusmustafa merged 4 commits intoOct 9, 2026
Merged
Conversation
Dynamic CPU/GPU graphs previously passed n_seq_active and attention_size(_swa) as separate inputs. Derive these values directly from the KQ mask shape using ShapeOf + Gather inside the model. This allows the GPU plugin to infer the Q reshape and KV slice shapes from shape inference rather than runtime input values. Q's head size also becomes static, enabling the fused SDPA kernel. GGML_OPENVINO_DISABLE_SHAPE_FROM_MASK=1 restores the previous separate-input behavior. The fused SDPA path also requires two supporting changes: - Clamp the GPU mask to -30000. A key block that is entirely -inf for a row produces NaNs in the fused kernel. - Restate the head count and head size after the KV slice. The plugin loses all slice dimensions when the end index is dynamic, causing SDPA to fall back to the reference kernel. Additional graph and quantization improvements: - Skip the identity reshape of rows in SET_ROWS. - Share a single Split for one-element view slices in dynamic GPU graphs (e.g. Gemma-4 per-layer embeddings), replacing one StridedSlice per layer. - Add GGML_OPENVINO_REQUANT_EMBD=native (off by default) to preserve the file's native quantization for token_embd/output instead of requantizing to Q8_0_C. - Include both environment settings in the graph and model cache keys and document them.
The dynamic GPU path reused a Split output for any single-element view by deriving the chunk index from offset / stride. This produced incorrect results for views that were not chunk-aligned or whose remaining dimensions did not preserve the source strides, as seen in test-backend-ops CONT with use_view_slice. Restrict Split reuse to cases where the view exactly matches a split chunk, meaning the offset is aligned to the split stride and the remaining non-unit dimensions preserve the original source strides. Gemma-4 per-layer embedding slices continue to use the shared Split path.
osabnis
force-pushed
the
ov-dynamic-decode
branch
from
October 6, 2026 21:23
9a78c56 to
0aa29c7
Compare
Initialize the newly added `shape_source` and `shape_axis` fields in the `ModelExtraInputInfo` brace initializer used by `add_extra_inputs()`. Without initializers for these fields, GCC emits `-Wmissing-field-initializers`, causing CI to fail with `LLAMA_FATAL_WARNINGS`.
Author
|
Fixed the CI failure in the self-hosted OpenVINO job. The GCC build stopped on Checked with clang and the CI warning flags on all files this PR changes: the only warning was this one, and it is gone with the fix. |
cavusmustafa
requested changes
Oct 9, 2026
The option identifies embedding and output weights by their tensor names, which are defined by core llama.cpp. Remove it from this PR; preserving the file quantization of these weights will be addressed in a separate change that identifies them by their roles in the graph.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Faster decode for dynamic (stateless) graphs in the OpenVINO backend, primarily on GPU.
Previously,
n_seq_activeandattention_size(_swa)were passed as separate graph inputs. This PR derives them from the KQ mask shape usingShapeOf + Gatherinside the model. The GPU plugin can then infer the Q reshape and KV slice shapes through shape inference instead of runtime input values. Q's head size also becomes static, allowing the plugin to select its fused SDPA kernel.Two companion changes are required for the fused kernel to work correctly:
-30000. A key block that is entirely-inffor a row produces NaNs in the fused kernel; this was observed with Gemma-4 prompts at context length 1024.CL_OUT_OF_RESOURCESerrors for SmolLM2 and Phi-3.5.Additional changes:
SET_ROWS.Splitfor one-element view slices in dynamic GPU graphs, such as Gemma-4's per-layer embeddings, instead of oneStridedSliceper layer.Both new settings are documented in
docs/backend/OPENVINO.mdand included in the graph and model cache keys:GGML_OPENVINO_DISABLE_SHAPE_FROM_MASK=1: restore the previous separate-input behavior.Additional information
Decode performance
llama-bench -fa 1 -p 0 -n 128 -d <depth>, Q4_K_M, Arc B390 (Panther Lake) GPU, OpenVINO master. Runs were performed at the console with a mean of two interleaved rounds.Base: fork + #336. Test: base + this PR. Both on the same OpenVINO build.
Gemma-4 at depth 0: +12%. The Gemma-4 baseline varied by approximately 10% between rounds at depths 512 and 2048.
Accuracy
GPU results against CPU references, #336 alone vs #336 + this PR, both on the same OpenVINO build:
The differences are within run-to-run variation on this GPU. Gemma-3 4B fails the perplexity run on both builds
(a known f16 overflow, also on the branch today).
The speed and full accuracy runs used the fork at
308ccd80awith #336. This branch is based on836d57176, where it builds and runs successfully and gives the same perplexity on SmolLM2, Llama-3.2, Phi-3.5, and Gemma-4 within run-to-run variation.Requirements