Skip to content

hexagon: ssm-conv updates - #29971

Merged
max-krasnyansky merged 3 commits into
ggml-org:masterfrom
qualcomm:hexagon-ssm-conv-updates
Oct 6, 2026
Merged

max-krasnyansky merged 3 commits into
ggml-org:masterfrom
qualcomm:hexagon-ssm-conv-updates

Conversation

@tboinovski1

Copy link
Copy Markdown
Contributor

Overview

  • Transpose via VTCM gather - replaced the old transpose_src0_block that did a 32x32 fp32 transpose with standard 5-stage HVX vshuff butterfly with hvx_ssm_conv_transpose_block: one Q6_vgather_ARMVw per output row.
  • Register-resident convolution - inverted the inner loop from token-outer/channel-inner to channel-block-outer/token-inner and added a d_conv == 4 (Mamba2) path holding the four taps and the sliding window in registers, so each new output token costs one load plus four qf32 multiply-accumulates with a single convertion at the end. The general d_conv path is unchanged as a fallback.
  • Double-buffered DMA in prefill
  • Decode restructuring - the weight and first input fetches are now queued back to back so the two DDR reads overlap, and hvx_ssm_conv_decode_4 multiplies directly in the raw {channel, tap} DMA layout, then deinterleaves the products with two levels of Q6_W_vdeal_VVR plus three adds.

Additional information

Verification - all 45 SSM_CONV test-backend-ops correctness tests still pass on HTP0

Test perf from my Kaanapali QRD:

test-backend-ops
before

  SSM_CONV(type=f32,ne_a=[515,3328,1,1],ne_b=[4,3328,1,1]):                     2504 runs -   519.67 us/run -    13403 kB/run -   24.60 GB/s
  SSM_CONV(type=f32,ne_a=[937,8192,1,1],ne_b=[4,8192,1,1]):                      560 runs -  1999.06 us/run -    60000 kB/run -   28.62 GB/s
  SSM_CONV(type=f32,ne_a=[4,3328,1,1],ne_b=[4,3328,1,1]):             221184 runs -     4.64 us/run -      117 kB/run -   24.06 GB/s
  SSM_CONV(type=f32,ne_a=[515,6144,1,1],ne_b=[4,6144,1,1]):                     1357 runs -   877.01 us/run -    24744 kB/run -   26.91 GB/s
  SSM_CONV(type=f32,ne_a=[4,6144,1,1],ne_b=[4,6144,1,1]):             172032 runs -     5.84 us/run -      216 kB/run -   35.25 GB/s
  SSM_CONV(type=f32,ne_a=[515,10240,1,1],ne_b=[4,10240,1,1]):                    814 runs -  1491.12 us/run -    41240 kB/run -   26.38 GB/s
  SSM_CONV(type=f32,ne_a=[4,10240,1,1],ne_b=[4,10240,1,1]):                   122880 runs -     8.22 us/run -      360 kB/run -   41.77 GB/s

after

  SSM_CONV(type=f32,ne_a=[515,3328,1,1],ne_b=[4,3328,1,1]):                     5008 runs -   252.23 us/run -    13403 kB/run -   50.68 GB/s
  SSM_CONV(type=f32,ne_a=[937,8192,1,1],ne_b=[4,8192,1,1]):                     1120 runs -  1307.48 us/run -    60000 kB/run -   43.76 GB/s
  SSM_CONV(type=f32,ne_a=[4,3328,1,1],ne_b=[4,3328,1,1]):             237568 runs -     4.26 us/run -      117 kB/run -   26.22 GB/s
  SSM_CONV(type=f32,ne_a=[515,6144,1,1],ne_b=[4,6144,1,1]):                     2714 runs -   481.60 us/run -    24744 kB/run -   49.00 GB/s
  SSM_CONV(type=f32,ne_a=[4,6144,1,1],ne_b=[4,6144,1,1]):             180224 runs -     5.56 us/run -      216 kB/run -   37.06 GB/s
  SSM_CONV(type=f32,ne_a=[515,10240,1,1],ne_b=[4,10240,1,1]):                   1628 runs -   881.91 us/run -    41240 kB/run -   44.60 GB/s
  SSM_CONV(type=f32,ne_a=[4,10240,1,1],ne_b=[4,10240,1,1]):                   131072 runs -     7.95 us/run -      360 kB/run -   43.19 GB/s

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes, to reason about the DMA flow and also to test and verify the implementation.

@tboinovski1 tboinovski1 changed the title hexagon: ssm-conv double-buffered DMA for prefill and decode restruct… hexagon: ssm-conv updates Oct 5, 2026
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Hexagon labels Oct 5, 2026
@tboinovski1
tboinovski1 force-pushed the hexagon-ssm-conv-updates branch from 1a2a90e to e4e110f Compare October 5, 2026 17:01
@tboinovski1
tboinovski1 marked this pull request as ready for review October 5, 2026 17:37
@tboinovski1
tboinovski1 requested a review from a team as a code owner October 5, 2026 17:37
@max-krasnyansky
max-krasnyansky force-pushed the hexagon-ssm-conv-updates branch from e4e110f to 97100bf Compare October 6, 2026 00:33
@max-krasnyansky

Copy link
Copy Markdown
Member

@lhez for another review/ack.

@max-krasnyansky
max-krasnyansky merged commit 5e03bdd into ggml-org:master Oct 6, 2026
22 of 23 checks passed
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
* hexagon: ssm-conv double-buffered DMA for prefill and decode restructuring

* hex-ssm-conv: remove divs from loops and fix trace events

* hex-dma: improved SSM_CONV dma pipeline and streamlined dma_queue

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
(cherry picked from commit 5e03bdd)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Hexagon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants