Skip to content

hexagon: HMX matmul with F16 activation and F16/F32 weights of any row count - #29626

Closed
njsyw1997 wants to merge 2 commits into
ggml-org:masterfrom
aizip:hex-hmx-fp-rag-32
Closed

njsyw1997 wants to merge 2 commits into
ggml-org:masterfrom
aizip:hex-hmx-fp-rag-32

Conversation

@njsyw1997

Copy link
Copy Markdown
Contributor

Overview

This PR lets the HMX matmul run on F16/F32 weights whose row count is not a
multiple of 32, and on F16 activations. Both show up on tensors that only
exist at runtime, so they cannot be padded or repacked ahead of time the way
quantized weights are; they have to be handled inside the kernels. The typical
case is ggml_conv_2d in the Qwen3-VL vision encoder: it passes the im2col
result as src0 (its row count is the number of patches) and the kernel weight
as src1 (F16 for an F16 mmproj).

It also fixed a typo which will cause the second weight chunk in the batched matmul prologue wrong.

Additional information

git clang-format produced unexpected results. The code is just manually formatted.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes. With help of Claude Farble. All the code has already been reviewed by me.

@njsyw1997
njsyw1997 requested a review from a team as a code owner September 29, 2026 02:49
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Hexagon labels Sep 29, 2026
@njsyw1997

Copy link
Copy Markdown
Contributor Author

Close this PR. It's already included in #29974

@njsyw1997 njsyw1997 closed this Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Hexagon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant