Skip to content

hexagon: copy short rows through VTCM with vgather (CONCAT dim 0, CPY) - #29739

Closed
karusrus wants to merge 1 commit into
ggml-org:masterfrom
karusrus:hexagon-short-rows-vtcm
Closed

karusrus wants to merge 1 commit into
ggml-org:masterfrom
karusrus:hexagon-short-rows-vtcm

Conversation

@karusrus

Copy link
Copy Markdown

Overview

Qwen3.5 decode concatenates the conv state with one token along dim 0 (rows of 3 + 1 f32, 8192 rows per layer) and copies the conv state view back (rows of 3 f32 at stride 4 into a contiguous dst). Both are short-row copies: the existing paths move a few bytes per row and wait on memory for every row.

This PR sends such copies through VTCM: one DMA brings the source span in, vgather places every word at its output position, and one DMA writes the rows out. The gather offsets are periodic (they repeat every ne / gcd(ne, 32) vectors), so they are built once per call and then advanced with a vector add.

It is used for f32 CONCAT on dim 0 and f32 same-type reshape CPY, when rows have at most 16 elements, dst rows are contiguous, tensors are 2D and the session runs on a single device. Everything else takes the existing paths. New file: htp/hvx-gather-rows.h.

Additional information

This complements #29673, which sped up the same CONCAT. SM7750 (Snapdragon 7 Gen 4), Qwen3.5-4B Q4_0 decode, current master vs this PR:

per op, usec master this PR
CONCAT 3:8192 x 1:8192 -> 4:8192 143.6 27.3
CPY conv state, 3:8192 at stride 4 -> contiguous 107.7 19.9

tg64: 8.53 -> 8.93 t/s (+4.7%). test-backend-ops: CONCAT 48/48, CPY 136/136.

An earlier version of this change (before #29673) was also checked on a chip with FP16 HMX: the tests pass and decode was faster there too.

Found while working on #29473. It doesn't use HMX and doesn't depend on that work.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. The code was written with Claude Code (Claude Opus 5.5). All numbers are from runs on the device.

…, CPY)

Decode of Qwen3.5 concatenates the conv state with one token (rows of 4 f32)
and copies the conv state view back (rows of 3 f32 at stride 4). Both paths
move these 8192 short rows a few bytes at a time and wait on memory for
every row.

Short rows now go through VTCM: one DMA brings the source span in, vgather
places every word at its output position, one DMA writes the rows out.
Single-device sessions only; everything else keeps the existing paths.

SM7750, Qwen3.5-4B Q4_0 decode, per op: CONCAT 143.6 -> 27.3 us,
CPY 107.7 -> 19.9 us; tg64 8.56 -> 8.95 t/s. test-backend-ops CONCAT 48/48,
CPY 136/136.

Assisted-by: Claude Opus 5.5
@karusrus
karusrus requested a review from a team as a code owner September 30, 2026 11:21
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Hexagon labels Sep 30, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 30, 2026

Copy link
Copy Markdown

Hi @karusrus, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@max-krasnyansky

Copy link
Copy Markdown
Member

@karusrus can you please check if #29685 covers your use-case.
It was a bit more generic and should cover this one too.

@karusrus

karusrus commented Oct 1, 2026 •

Copy link
Copy Markdown
Author

@max-krasnyansky thanks, #29685 is the more general version of this, so I'm closing this one in its favour.

The case I was after is the Qwen3.5 conv-state CONCAT (dim 0, rows of a few elements plus one token) and CPY of the same short rows. I'll re-measure it on SM7750 with current master and come back with numbers only if short rows are still noticeably slower.

That frees my one open-PR slot, so I'm reopening #29740 (integer HMX for SoCs without FP16 HMX). It doesn't touch the CONCAT/CPY files.

@karusrus karusrus closed this Oct 1, 2026
@karusrus

karusrus commented Oct 1, 2026

Copy link
Copy Markdown
Author

@max-krasnyansky I re-measured on SM7750 with current master (ec7630a, includes #29685). #29685 doesn't take these shapes, so the short-row ops are unchanged:

per op, usec (avg over decode) master before #29685 master ec7630a this PR (on its 30.09 base)
CONCAT 3:8192 x 1:8192 -> 4:8192 143.1 145.5 27.1
CPY conv state, 3:8192 at stride 4 -> contiguous 105.9 106.1 20.1

Qwen3.5-4B Q4_0 tg64, two runs each in the same session: 8.53 / 8.54, 8.52 / 8.51, 8.96 / 8.92 t/s.

Both ops miss the new DMA paths: in the CONCAT, src1 is the transposed qkv view (nb0 = 32768), and the CPY is a reshape from rows of 3 at stride 4, so the source isn't contiguous.

As a new contributor I can keep one PR open, so I'd leave this closed until #29740 is through, then rebase it on top of #29685 and reopen. If you'd rather look at this one first, say so and I'll swap them.

Disclosure: written with Claude Code; the numbers are from runs on the device.

@max-krasnyansky

max-krasnyansky commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

#30067 converts everything in CPU/CONCAT to DMA/VTCM and vectorized type converstions.
Tested on v73 and should work for your case as well

@karusrus

karusrus commented Oct 7, 2026

Copy link
Copy Markdown
Author

@max-krasnyansky thanks for the pointer. I measured #30067 on SM7750 (master 36a7391), Qwen3.5-4B Q4_0 decode, both builds in the same session, per-op averages:

per op, usec master 01.10 (ec7630a) master now (36a7391, with #30067) short-row path (#29739)
CONCAT 3:8192 x 1:8192 -> 4:8192 145.5 142.5 26.1
CPY conv state, 3:8192 at stride 4 -> contiguous 106.1 257.5 19.7

So on this chip the conv-state CPY got slower with #30067 (106 -> 257 us), and the CONCAT is unchanged (src1 is still the transposed qkv view, nb0 = 32768). CONCAT 52/52 and CPY 140/140 pass on master. tg64: 8.46 / 8.54 t/s on master vs 9.16 / 9.15 with the short-row path; that build is on an older base, so the per-op numbers are the cleaner comparison. I can share the full op profiles if that helps.

I'm happy to rebase the short-row path on top of #30067 as a small follow-up once there's room for it.

P.S. Separately, when you have a moment: #29740 (integer HMX for SoCs without FP16 HMX, up to 3.3x prefill on SM7750) is rebased on master and waiting for a look. I'm happy to adjust it to whatever detection or mode you prefer.

@max-krasnyansky

Copy link
Copy Markdown
Member

Oh, interesting. I didn't expect any perf bump but didn't expect the regressions either.
I'll take another look at yours asap, definitely don't want to miss any optimization oportunity.

And yes, sorry for the delay on the INT HMX thingy. Kind of a long backlog right now but will definitely get back to you.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Hexagon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants