Skip to content

ggml-openvino : modularize backend, op registry support, buffer management, and quantization - #324

Open
mostafafaheem wants to merge 33 commits into
ravi9:dev_backend_openvinofrom
mostafafaheem:refactor-backend-openvino-consolidated
Open

mostafafaheem wants to merge 33 commits into
ravi9:dev_backend_openvinofrom
mostafafaheem:refactor-backend-openvino-consolidated

Conversation

@mostafafaheem

@mostafafaheem mostafafaheem commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator

Overview

This PR merges multiple previous refactoring PRs into one, namely #240, #278, #279, #284, #307, and resolves all merge conflicts. This is the AI generated description for the combined changes:

  1. Op Registry and Op Support Policy:

    • Bind op support verification directly to the op registry via OpEntry constructor (CreatorFunction + SupportsFunction), enforcing support rules at compile-time for all registered operations.
    • Extract per-op device support logic from monolithic switch in ggml-openvino.cpp into openvino/op_table.{h,cpp}.
    • Convert runtime translator exceptions (ARGSORT, PAD, TRI) into explicit frontend gate declines.
    • Centralize dynamic dimension inference and parameter creation helpers.
  2. Buffer Management and RAII:

    • Extract buffer allocation, tensor registration, and context tracking into ggml-openvino-buffer.{h,cpp}.
    • Adopt RAII for tensor extra lifecycle management with std::unique_ptr.
    • Extract weight buffer RSS memory release logic (GGML_OPENVINO_RELEASE_WEIGHTS) into ggml-openvino-weight-buffer-release.{h,cpp}.
    • Remove duplicated buffer storage state and clean up buffer context tagging.
  3. Quantization Subsystem Decomposition:

    • Split monolithic ggml-quants.{h,cpp} into modular components under ggml/src/ggml-openvino/quant/:
      • weights.{h,cpp}: weight orchestration, layout helpers, shape centralization, and zero-point modes.
      • requant.cpp: quantized weight requantization.
      • extract.cpp: weight data extraction routines.
      • graph.cpp: integer and quantized weight OpenVINO subgraph builders.
      • quantization.{h,cpp}: quantization storage planning and types.
    • Decouple quantization routines from ggml-openvino-extra.{h,cpp}.
  4. API Documentation and Code Hygiene:

    • Document public OpenVINO backend C API in ggml-openvino.h and buffer interface with Doxygen comments.
    • Slim down ggml-openvino.cpp into a lightweight orchestration layer.
    • Clean up unused headers, dead code, and apply linter/formatting fixes.

Additional information

Requirements

zhaixuejun1993 and others added 30 commits September 20, 2026 12:44
Move OpenVINO op support / unsupported-case policy logic out of ggml-openvino.cpp into a dedicated implementation pair: ggml-openvino-op-support.cpp and ggml-openvino-op-support.h.

Keep behavior unchanged by retaining the original device callback shape and delegating through ggml_openvino_device_supports_op_impl().

This reduces ggml-openvino.cpp size and keeps support-policy code isolated for easier maintenance and future policy updates.
Expand the developer-facing OP/Limitation comment block to include all operators currently registered in openvino/op_table.cpp.

For each registered op, document either a concrete runtime policy gate or explicitly mark that no extra restriction exists beyond the global type/rank gates.
Move OpenVINO backend buffer allocation into a small storage abstraction that owns either host memory or GPU remote USM tensors. The buffer context now keeps a unique_ptr to this storage and only exposes the raw data pointer and ov::Tensor wrapper for ggml/OpenVINO integration.

Store tensor extras in unique_ptrs owned by the OpenVINO buffer context instead of manually deleting raw pointers at each replacement and during context teardown. This makes tensor->extra a non-owning view while the context remains responsible for lifetime, reducing leak and double-delete risks when extras are replaced or buffers are destroyed.

Remove the obsolete raw-pointer tensor extra factory and keep the unique_ptr-returning factory as the only creation API so new call sites cannot accidentally reintroduce ambiguous ownership.
Move the GGML_OPENVINO_RELEASE_WEIGHTS host weight-buffer registry out of ggml-openvino.cpp and into a dedicated ggml-openvino-weight-buffer-release module.

Keep ggml-openvino.cpp focused on backend buffer/device glue while the new helper owns registration, release state, and MADV_DONTNEED handling for host weight pages. Include the helper explicitly from the backend and utils call sites instead of exposing these declarations through ggml-openvino-extra.h.
Keep ggml_backend_openvino_buffer_context from caching data, size, and ov_buffer fields that are already owned by ggml_openvino_buffer_storage.

Expose small data() and size() accessors on the context so callers continue to read the buffer base and allocation size through the storage owner. This leaves storage as the single source of truth for host and remote buffer state.

Also compute the KV-cache buffer offset before replacing the old context during host-to-remote migration, so the offset calculation no longer depends on a pointer after its owning storage has been destroyed.
Remove the unused name field from ggml_backend_openvino_buffer_context. Buffer type and device contexts still keep their names for get_name callbacks, but concrete buffer instances only need device, id, remote state, storage, and tensor extras.
Move the OpenVINO backend buffer context and ggml_backend_buffer_i callbacks out of ggml-openvino.cpp into ggml-openvino-buffer.cpp/.h. Keep ggml-openvino.cpp focused on buffer types, backend/device registration, and high-level OpenVINO backend entry points.

Preserve the existing public C ABI for ggml_backend_buffer_is_openvino and ggml_backend_openvino_buffer_get_ctx_id by defining them with GGML_BACKEND_API from the new buffer implementation file. This avoids hidden or C++-mangled symbols when ggml-openvino is built as a shared backend.

Fold the host/remote buffer storage helpers into the new buffer implementation file because they are only used by the concrete buffer context. This removes the separate buffer-storage files while keeping host aligned allocation and GPU USM remote allocation behavior unchanged.

Add Doxygen-style descriptions for the new internal buffer helpers and the host weight-buffer release helpers.
Add Doxygen-style descriptions for the public OpenVINO backend API declarations, including backend initialization, buffer type queries, device count, and registry access.
Add explicit host/device metadata to OpenVINO buffer type contexts and use it when querying OpenVINO buffer type kinds. This avoids identifying buffer types by comparing get_name callback pointers while still checking that the buffer type belongs to the OpenVINO registry before reading its context.
Move the GGML_OPENVINO_STATEFUL_EXECUTION environment check into a shared ggml_openvino_is_stateful_enabled helper. This keeps the backend runtime context and buffer initialization paths aligned on the same stateful execution policy.
Move GGML_UNUSED markers ahead of their return statements so they are reachable and keep the intent clear to readers and compilers.
Factor the duplicated device and host buffer type initialization loops into a shared helper with separate static state for each buffer type kind. This keeps returned buffer type pointers and context storage stable while reducing repeated setup logic.
Remove OpenVINO runtime, quantization, and C library includes from ggml-openvino.cpp that are no longer needed after moving the concrete buffer implementation into its own source file.
Document that the OpenVINO backend currently exposes one logical ggml device and selects the actual OpenVINO plugin/device through configuration.
Clarify that OpenVINO registry and device contexts are intentionally owned by the process-lifetime backend registry singleton.
Assisted-by: GitHub Copilot
Assisted-by: GitHub Copilot
Assisted-by: GitHub Copilot

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Quantization races, invalid Q8 ranges, unsafe weight-buffer tracking, incorrect platform handling, and unrelated policy-file deletion remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 2 High severity · 3 Medium severity

Open (5)
What changed in this PR

Refactors the OpenVINO backend into dedicated operation, buffer, and quantization modules.

Changes:

  • Associates operation support checks with translator registry entries.
  • Extracts buffer lifecycle, weight release, and quantization logic.
  • Adds API documentation and removes AGENTS.md.
File Description
AGENTS.md Deletes repository guidance.
ggml/​include/​ggml-openvino.h Documents public APIs.
ggml/​src/​ggml-openvino/​utils.h Declares stateful-mode helper.
ggml/​src/​ggml-openvino/​utils.cpp Implements stateful-mode helper.
ggml/​src/​ggml-openvino/​quant/​weights.h Declares weight processing APIs.
ggml/​src/​ggml-openvino/​quant/​weights.cpp Orchestrates weight conversion.
ggml/​src/​ggml-openvino/​quant/​requant.cpp Implements requantization.
ggml/​src/​ggml-openvino/​quant/​quantization.h Defines quantization layouts.
ggml/​src/​ggml-openvino/​quant/​quantization.cpp Plans quantized storage.
ggml/​src/​ggml-openvino/​quant/​graph.cpp Builds dequantization graphs.
ggml/​src/​ggml-openvino/​quant/​extract.cpp Extracts quantized data.
ggml/​src/​ggml-openvino/​openvino/​utils.h Updates utility declarations.
ggml/​src/​ggml-openvino/​openvino/​utils.cpp Updates shared-pointer parameter handling.
ggml/​src/​ggml-openvino/​openvino/​translate_session.h Uses registry entries.
ggml/​src/​ggml-openvino/​openvino/​translate_session.cpp Translates through registry entries.
ggml/​src/​ggml-openvino/​openvino/​op/​mul_mat_id.cpp Updates includes and references.
ggml/​src/​ggml-openvino/​openvino/​op_table.h Defines support-policy registry.
ggml/​src/​ggml-openvino/​openvino/​op_table.cpp Implements operation support rules.
ggml/​src/​ggml-openvino/​ggml-quants.h Removes legacy declarations.
ggml/​src/​ggml-openvino/​ggml-openvino-weight-buffer-release.h Declares weight release APIs.
ggml/​src/​ggml-openvino/​ggml-openvino-weight-buffer-release.cpp Implements RSS reclamation.
ggml/​src/​ggml-openvino/​ggml-openvino-extra.h Removes relocated APIs.
ggml/​src/​ggml-openvino/​ggml-openvino-extra.cpp Retains tensor-extra creation.
ggml/​src/​ggml-openvino/​ggml-openvino-buffer.h Declares buffer APIs.
ggml/​src/​ggml-openvino/​ggml-openvino-buffer.cpp Implements buffer management.
ggml/​src/​ggml-openvino/​ggml-decoder.cpp Uses extracted modules.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +44 to +49
for (const auto & buffer : reg.buffers) {
if (buffer.first == data) {
return;
}
}
reg.buffers.emplace_back(data, size);
Comment on lines +68 to +77
auto * zp = static_cast<uint8_t *>(zp_arr.data());
ov::parallel_for(scales_arr.get_size(), [&](size_t i) {
scales[i] = ov::float16::from_bits(*((uint16_t *) (data + i * bytes_per_block)));
if (i % 2 == 0) {
zp[i / 2] = 8;
} else {
zp[i / 2] |= (8 << 4);
}
unpack_32_4(data + i * bytes_per_block + 2, weights + i * 16);
});
}
}
#endif
reg.released = true;
Comment on lines +404 to +406
if (fabsf(scale - 1.0f) < 1e-6f && is_gemma3n_flash_attn_pattern(op)) {
return {false, "FLASH_ATTN_EXT gemma3n pattern on GPU is not supported"};
}
max = std::max(v, max);
}

const float d = (max - min) / ((1 << 8) - 1);
@ravi9
ravi9 force-pushed the dev_backend_openvino branch 6 times, most recently from 308ccd8 to 836d571 Compare October 3, 2026 19:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants