Skip to content

opencl: add bin kernel kernel_gemm_noshuffle_q6_k_q8_1_dp4a_ila_a8_bin - #29057

Merged
lhez merged 1 commit into
ggml-org:masterfrom
qualcomm:q6_k-a8-dp4a-bin-kernel
Sep 23, 2026
Merged

lhez merged 1 commit into
ggml-org:masterfrom
qualcomm:q6_k-a8-dp4a-bin-kernel

Conversation

@shaofeiqi

Copy link
Copy Markdown
Contributor

Overview

Add an optimized DP4A bin kernel for Q6_K non-MoE GEMM.

Requirements

@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend labels Sep 17, 2026
@shaofeiqi
shaofeiqi force-pushed the q6_k-a8-dp4a-bin-kernel branch from 95cfada to 7aafc3e Compare September 22, 2026 20:32
@shaofeiqi
shaofeiqi marked this pull request as ready for review September 22, 2026 20:34
@shaofeiqi
shaofeiqi requested a review from a team as a code owner September 22, 2026 20:34
@lhez lhez added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 23, 2026
@lhez
lhez merged commit fee39dd into ggml-org:master Sep 23, 2026
14 of 15 checks passed
@BrewTestBot BrewTestBot mentioned this pull request Sep 24, 2026
1 task done
wanghqc added a commit to qualcomm/llama.cpp that referenced this pull request Sep 25, 2026
217 upstream commits since ad6c668, 8 of them in ggml-opencl. Two files
conflicted, 16 hunks.

- ssm_scan: upstream's generic kernel (ggml-org#28881) is added beside the
  specialised Mamba-2 kernels as the fallback for element-wise A and any
  other power-of-two d_state. The specialised d128/d256 kernels keep their
  row-folded variants and snapshot support; all of them are released when
  the device cannot give them a 64-lane subgroup, as upstream does.
  supports_op keeps the snapshot-width restriction and widens d_state to
  upstream's rule.
- FA bin kernel (ggml-org#29046), q6_K ILA GEMM (ggml-org#28678, ggml-org#29057), q4_0/q4_K dp4a
  ILA GEMMs (ggml-org#29055, ggml-org#29056): taken as upstream wrote them. The capability
  check reads has_integer_dot_product, this branch's name for the field.
- q4_0 MoE dp4a/ILA arbitration: kept this branch's version, which
  upstream adopted with the same routing threshold.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. OpenCL Issues specific to the OpenCL backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants