Skip to content

opencl: refine a8x bin kernel loading condition - #29503

Merged
ggerganov merged 1 commit into
ggml-org:masterfrom
qualcomm:lh/fix-bin-kernels-load-condition
Sep 27, 2026
Merged

ggerganov merged 1 commit into
ggml-org:masterfrom
qualcomm:lh/fix-bin-kernels-load-condition

Conversation

@lhez

@lhez lhez commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

Overview

The conditions for loading some A8x binary kernels are too restrictive and prevent from the kernels being loaded on compatible GPUs other than X2. The kernel lib itself checks GPU model to make proper kernels are returned. This PR refines the condition and allows the lib to make the decision.

Additional information

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes. I made the changes, asked Codex to check.

@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend labels Sep 26, 2026
@lhez
lhez marked this pull request as ready for review September 26, 2026 22:43
@lhez
lhez requested a review from a team as a code owner September 26, 2026 22:43
@max-krasnyansky max-krasnyansky added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 26, 2026
@ggerganov
ggerganov merged commit c9064dd into ggml-org:master Sep 27, 2026
15 checks passed
wanghqc added a commit to qualcomm/llama.cpp that referenced this pull request Sep 29, 2026
ggml-opencl.cpp: this branch's layout kept, the three upstream OpenCL changes
(ggml-org#29401 q5_K bin kernels, ggml-org#29439 q8_0 dp4a bin kernel, ggml-org#29503 bin kernel
loading condition) replayed onto it. The q5_K bin layout and the q8_0 dp4a bin
GEMM are opt-in here (GGML_OPENCL_Q5_K_BIN=1, GGML_OPENCL_Q8_0_BIN_DP4A=1):
by default they would take the spec/MTP verify widths from the cooperative-K
and narrow kernels. q5_K bin is limited to 2-D weights, q8_0 bin to N > 16.

common/speculative, server-context: this branch's llama_batch implementation
kept (tree drafting puts several seq_ids on one token, which common_batch
cannot hold); a common_batch overload of common_speculative_process converts
for the new callers, the mtmd post-decode callback takes the new embd batch,
and ggml-org#28876, ggml-org#29556 and ggml-org#29648 are applied to server-context.
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. OpenCL Issues specific to the OpenCL backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants