Skip to content

sched: fix backend assignment for ops with view-backed outputs - #28075

Open
yomaytk wants to merge 1 commit into
ggml-org:masterfrom
yomaytk:fix-backend-sched
Open

yomaytk wants to merge 1 commit into
ggml-org:masterfrom
yomaytk:fix-backend-sched

Conversation

@yomaytk

@yomaytk yomaytk commented Aug 31, 2026 •

Copy link
Copy Markdown
Member

Overview

This PR fixes a scheduler bug and enables the WebGPU backend to pass all the
currently skipped models in test-llama-archs (deepseek32, glm-dsa, dots3note,
qwen4exp).

When a node writes into a view of another tensor, it must be assigned to a
backend that can access the view_src buffer, since the scheduler copies inputs
across backends but not outputs.

For the skipped models on the WebGPU backend, the following code in the DSA
variant of build_attn causes the error:

ggml_tensor * kq_mask_all = ggml_fill(ctx0, kq_mask, -INFINITY);
…
ggml_tensor * kq_mask_top_k = ggml_set_rows(ctx0, kq_mask_all, …);

The WebGPU FILL op does not support the F16 kq_mask, so the FILL is assigned to the CPU but the SET_ROWS is assigned to WebGPU. The SET_ROWS output is a view of the FILL result, which causes the error.

So this PR adds logic to reassign the writer node to a backend that can access
the view_src buffer.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes — used for root cause analysis and for discussing/reviewing the changes.

@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning labels Aug 31, 2026
@yomaytk
yomaytk force-pushed the fix-backend-sched branch 2 times, most recently from 975d5b1 to 7d333cd Compare August 31, 2026 15:08
@yomaytk

yomaytk commented Sep 1, 2026

Copy link
Copy Markdown
Member Author

test-llama-archs passes on CI: https://github.com/ggml-org/llama.cpp/actions/runs/33372854756

Comment thread ggml/src/ggml-backend.cpp Outdated
*cur_backend_id = vsrc_backend_id;
GGML_ASSERT(vsrc_backend_id == -1 || ggml_backend_supports_op(sched->backends[*cur_backend_id], node));
SET_CAUSE(node, "4.vsrc");
} else if (!ggml_is_view_op(node->op) && !ggml_backend_sched_buffer_supported(sched, node->view_src, *cur_backend_id) && vsrc_backend_id != -1) {

@yomaytk yomaytk Sep 1, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This change doesn’t cover the case where the view_src has no backend yet (vsrc_backend_id == -1) because I am not sure such a case actually happen.
If you are aware of such a case, I would appreciate it if you could let me know.

@yomaytk
yomaytk marked this pull request as ready for review September 1, 2026 19:27
@yomaytk
yomaytk requested a review from ggerganov September 1, 2026 19:29
@yomaytk
yomaytk force-pushed the fix-backend-sched branch 2 times, most recently from cb85f71 to e77f40a Compare October 5, 2026 15:19
@yomaytk yomaytk mentioned this pull request Oct 5, 2026
Comment thread ggml/src/ggml-backend.cpp Outdated
*cur_backend_id = vsrc_backend_id;
GGML_ASSERT(vsrc_backend_id == -1 || ggml_backend_supports_op(sched->backends[*cur_backend_id], node));
SET_CAUSE(node, "4.vsrc");
} else if (!ggml_is_view_op(node->op) && !ggml_backend_sched_buffer_supported(sched, node->view_src, *cur_backend_id) && vsrc_backend_id != -1) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch can only ever be executed if *cur_backend_id != -1, meaning that a backend ID has already been assigned. The code then proceeds to re-assign the backend. I think that if at all possible we should be assigning backends exactly once per node and then not override that assignment later on. Did you check which of the previous passes does the assignment? Would it be viable to intervene there instead?

@yomaytk yomaytk Oct 8, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Did you check which of the previous passes does the assignment?

No, I didn't when I opened this PR. So I checked which pass does the assignment on the WebGPU backend for the recent CI failure reported here, and found that pass 2 does it.
In general, the backend of node->view_src is not checked at all when a node is assigned before pass 4, so similar failures could also occur in pass 3.

I think that if at all possible we should be assigning backends exactly once per node
Would it be viable to intervene there instead?

Yes, I agree with that. So instead of overriding the assignment in pass 4, I changed passes 2 and 3 so that each node is assigned only once. Pass 2 assigns a backend to nodes only when node->view_src == NULL, and pass 3 considers only backends that can access the view_src buffer for nodes with node->view_src != NULL.

I propose this new change in the following commit instead, and I confirmed that test-llama-archs passes on the WebGPU backend on my M5. Could you take a look at this?

master...yomaytk:llama.cpp:fix-backend-sched-again

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you please force-push that here so that we have a record of how we decided to do things in a single place?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Got it, I've force-pushed it.

@yomaytk yomaytk Oct 8, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, this change caused a new CI failure in the Meta backend on nvidia: https://github.com/ggml-org/llama.cpp/actions/runs/37757525194/job/113245637763?pr=28075#step:9:131
I'm looking into it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants