Skip to content

ggml-meta: resolve multi buffer views - #29266

Merged
0cc4m merged 2 commits into
masterfrom
0cc4m/meta-backend-multi-buffer-view-fix
Sep 23, 2026
Merged

0cc4m merged 2 commits into
masterfrom
0cc4m/meta-backend-multi-buffer-view-fix

Conversation

@0cc4m

@0cc4m 0cc4m commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

Overview

This is the most minimal change I can come up with to unblock Vulkan tensor parallel. It resolves multi-buffers in views, which unblocks most cases with split model or kv tensors.

@JohannesGaessler I know you want a proper allocator change, but for the short time it would help us a lot to unblock the most common multi-buffer cases. I don't think this will interfere with a long term proper refactor.

Requirements

@github-actions github-actions Bot added the ggml changes relating to the ggml tensor library for machine learning label Sep 22, 2026

@JohannesGaessler JohannesGaessler left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These changes are I think okay and I probably would have done equivalent ones if I had known that it could be fixed like this. Please add a TODO to revisit the logic in the future if and when the memory allocation has been refactored.

@pwilkin pwilkin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this can actually unlock most cases, then I think it's worth it.

@0cc4m
0cc4m merged commit 9425611 into master Sep 23, 2026
23 of 26 checks passed
@0cc4m
0cc4m deleted the 0cc4m/meta-backend-multi-buffer-view-fix branch September 23, 2026 05:35
@BrewTestBot BrewTestBot mentioned this pull request Sep 23, 2026
1 task done
LadislavSopko pushed a commit to 0ics-srls/llama.cpp that referenced this pull request Oct 5, 2026
* ggml-meta: resolve multi buffer views

* add TODO to revisit if graph allocator gets refactored
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* ggml-meta: resolve multi buffer views

* add TODO to revisit if graph allocator gets refactored
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
* ggml-meta: resolve multi buffer views

* add TODO to revisit if graph allocator gets refactored

(cherry picked from commit 9425611)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants