Skip to content

hexagon: faster converting CPY and permuted CONT/CPY - #30132

Open
njsyw1997 wants to merge 2 commits into
ggml-org:masterfrom
aizip:hex-cpy-transpose
Open

njsyw1997 wants to merge 2 commits into
ggml-org:masterfrom
aizip:hex-cpy-transpose

Conversation

@njsyw1997

Copy link
Copy Markdown
Contributor

Overview

#30067 moved every CPY/CONT onto the DMA → VTCM path, which regressed two cases:

  1. Converting CPY (e.g. f32 → f16)
    The kernel moves one row per DMA descriptor (dma_queue_push is called with nrows = 1), with only two rows in flight, so short rows are bound by DMA latency. This PR moves a block of rows (up to 32 KB) per 2D descriptor and converts them row by row in VTCM.

  2. Permuted CONT/CPY
    A transposing copy falls back to the reshape path, which issues DMA rows of a single element. Following the idea of hexagon: copy short rows through VTCM with vgather (CONCAT dim 0, CPY) #29739, this PR copies whole tiles into VTCM and permutes them there with HVX gathers. Leading dims that stay contiguous on both sides are merged into one element (up to 128 B).
    The code before hexagon: CPY/CONCAT/CONT/DUP overhaul to use DMA/HVX for all cases #30067 did not show the problem might because it read DDR through the L2 cache, and most small cases reused cached lines efficiently. Now that every CPY is staged through VTCM by DMA, there is no cache to reuse.

Note on the bank-conflict avoidance: the input tile pitch is padded so that the lanes of each gather fall into different VTCM banks. On my test device this does not change the op time, because DMA dominates the cost. The gathers themselves do get about 3× faster (theoretically, from HTP programmer guide), so I kept the padding as headroom for devices where the gathers are not fully hidden behind DMA.

Additional information

SM8850 / HTP v81, test-backend-ops perf -b HTP0:

case master this PR
CPY f32→f16 [512,3072] 180.5 µs 154.9 µs
CONT f32 [1024,64,64] perm(2,1,0,3) 51.2 ms 0.96 ms

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes. With the help of Claude Farble. All the code has already reviewed by me.

@njsyw1997
njsyw1997 requested a review from a team as a code owner October 8, 2026 04:24
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Hexagon labels Oct 8, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Hexagon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant