Skip to content

hexagon: optimize concat op - #29673

Merged
max-krasnyansky merged 3 commits into
ggml-org:masterfrom
qualcomm:tr/hex-concat-opt2
Sep 29, 2026
Merged

max-krasnyansky merged 3 commits into
ggml-org:masterfrom
qualcomm:tr/hex-concat-opt2

Conversation

@trivikram-reddy1

@trivikram-reddy1 trivikram-reddy1 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Overview

This PR contains following changes to speed-up concat op

  • Reduce pkts in transpose hot loop, gather directly into dst buffer, use special instruction to synchronize gather operations
  • Optimize DMA-HVX pipeline, transpose operation not in critical path any more
  • Add transpose helpers, these can be reused in other kernels (gdn, ssm_conv etc.,)
  • Use fastdiv

Additional information

Results from 8 elite Gen5

| model          |     size | params | ngl | threads | n_ubatch | dev  |   test |   (Before)  t/s |    (After)  t/s |   Diff |
| ---------------| -------: | -----: | --: | ------: | -------: | ---- | -----: | --------------: | --------------: | -----: |
| qwen35 2B Q4_0 | 1.12 GiB | 1.88 B |  -1 |       6 |     1024 | HTP0 | pp1024 |  2640.12 ± 0.49 |  2720.66 ± 4.08 | +3.05% |
| qwen35 2B Q4_0 | 1.12 GiB | 1.88 B |  -1 |       6 |     1024 | HTP0 | pp2048 |  2638.95 ± 1.85 |  2718.55 ± 1.71 | +3.02% |
| qwen35 2B Q4_0 | 1.12 GiB | 1.88 B |  -1 |       6 |     1024 | HTP0 |  tg128 |    40.44 ± 0.11 |    41.42 ± 0.13 | +2.42% |

| Op     | Dims                            | DTypes           | (Before) Avg usec |  (After) Avg usec |    Diff |
| ------ | --------------------------------| ---------------- | ----------------- | ----------------- | ------- |
| CONCAT | 3:6144 x 1024:6144 -> 1027:6144 | f32 x f32 -> f32 |           1762.00 |           1137.49 | -35.44% |
| CONCAT |       3:6144 x 1:6144 -> 4:6144 | f32 x f32 -> f32 |             60.92 |             27.58 | -54.73% |

Requirements

gather directly into dst buffer, use special instruction for gather sync
replace calls to sw divide with fastpath
@trivikram-reddy1
trivikram-reddy1 requested a review from a team as a code owner September 29, 2026 17:31
@trivikram-reddy1

Copy link
Copy Markdown
Contributor Author

@max-krasnyansky and @lhez, could you please review when you get a chance.

@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Hexagon labels Sep 29, 2026
@max-krasnyansky
max-krasnyansky merged commit 7fee178 into ggml-org:master Sep 29, 2026
15 of 18 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
* hex-concat: reduce pkts in gather/transpose hot loop

gather directly into dst buffer, use special instruction for gather sync

* hex-concat: use fastdiv

replace calls to sw divide with fastpath

* hex-concat: optimize DMA-HVX pipeline and add transpose helpers
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* hex-concat: reduce pkts in gather/transpose hot loop

gather directly into dst buffer, use special instruction for gather sync

* hex-concat: use fastdiv

replace calls to sw divide with fastpath

* hex-concat: optimize DMA-HVX pipeline and add transpose helpers
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
* hex-concat: reduce pkts in gather/transpose hot loop

gather directly into dst buffer, use special instruction for gather sync

* hex-concat: use fastdiv

replace calls to sw divide with fastpath

* hex-concat: optimize DMA-HVX pipeline and add transpose helpers

(cherry picked from commit 7fee178)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Hexagon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants