Repository navigation
Conversation
… so short rows are not DMA-latency bound
…rs, the dense leading dims moving as one element
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
#30067 moved every CPY/CONT onto the DMA → VTCM path, which regressed two cases:
Converting CPY (e.g. f32 → f16)
The kernel moves one row per DMA descriptor (
dma_queue_pushis called withnrows = 1), with only two rows in flight, so short rows are bound by DMA latency. This PR moves a block of rows (up to 32 KB) per 2D descriptor and converts them row by row in VTCM.Permuted CONT/CPY
A transposing copy falls back to the reshape path, which issues DMA rows of a single element. Following the idea of hexagon: copy short rows through VTCM with vgather (CONCAT dim 0, CPY) #29739, this PR copies whole tiles into VTCM and permutes them there with HVX gathers. Leading dims that stay contiguous on both sides are merged into one element (up to 128 B).
The code before hexagon: CPY/CONCAT/CONT/DUP overhaul to use DMA/HVX for all cases #30067 did not show the problem might because it read DDR through the L2 cache, and most small cases reused cached lines efficiently. Now that every CPY is staged through VTCM by DMA, there is no cache to reuse.
Note on the bank-conflict avoidance: the input tile pitch is padded so that the lanes of each gather fall into different VTCM banks. On my test device this does not change the op time, because DMA dominates the cost. The gathers themselves do get about 3× faster (theoretically, from HTP programmer guide), so I kept the padding as headroom for devices where the gathers are not fully hidden behind DMA.
Additional information
SM8850 / HTP v81,
test-backend-ops perf -b HTP0:CPY f32→f16 [512,3072]CONT f32 [1024,64,64] perm(2,1,0,3)Requirements