Repository navigation
Conversation
…, CPY) Decode of Qwen3.5 concatenates the conv state with one token (rows of 4 f32) and copies the conv state view back (rows of 3 f32 at stride 4). Both paths move these 8192 short rows a few bytes at a time and wait on memory for every row. Short rows now go through VTCM: one DMA brings the source span in, vgather places every word at its output position, one DMA writes the rows out. Single-device sessions only; everything else keeps the existing paths. SM7750, Qwen3.5-4B Q4_0 decode, per op: CONCAT 143.6 -> 27.3 us, CPY 107.7 -> 19.9 us; tg64 8.56 -> 8.95 t/s. test-backend-ops CONCAT 48/48, CPY 136/136. Assisted-by: Claude Opus 5.5
|
Hi @karusrus, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
@max-krasnyansky thanks, #29685 is the more general version of this, so I'm closing this one in its favour. The case I was after is the Qwen3.5 conv-state CONCAT (dim 0, rows of a few elements plus one token) and CPY of the same short rows. I'll re-measure it on SM7750 with current master and come back with numbers only if short rows are still noticeably slower. That frees my one open-PR slot, so I'm reopening #29740 (integer HMX for SoCs without FP16 HMX). It doesn't touch the CONCAT/CPY files. |
|
@max-krasnyansky I re-measured on SM7750 with current master (ec7630a, includes #29685). #29685 doesn't take these shapes, so the short-row ops are unchanged:
Qwen3.5-4B Q4_0 tg64, two runs each in the same session: 8.53 / 8.54, 8.52 / 8.51, 8.96 / 8.92 t/s. Both ops miss the new DMA paths: in the CONCAT, src1 is the transposed qkv view (nb0 = 32768), and the CPY is a reshape from rows of 3 at stride 4, so the source isn't contiguous. As a new contributor I can keep one PR open, so I'd leave this closed until #29740 is through, then rebase it on top of #29685 and reopen. If you'd rather look at this one first, say so and I'll swap them. Disclosure: written with Claude Code; the numbers are from runs on the device. |
|
#30067 converts everything in CPU/CONCAT to DMA/VTCM and vectorized type converstions. |
|
@max-krasnyansky thanks for the pointer. I measured #30067 on SM7750 (master 36a7391), Qwen3.5-4B Q4_0 decode, both builds in the same session, per-op averages:
So on this chip the conv-state CPY got slower with #30067 (106 -> 257 us), and the CONCAT is unchanged (src1 is still the transposed qkv view, nb0 = 32768). I'm happy to rebase the short-row path on top of #30067 as a small follow-up once there's room for it. P.S. Separately, when you have a moment: #29740 (integer HMX for SoCs without FP16 HMX, up to 3.3x prefill on SM7750) is rebased on master and waiting for a look. I'm happy to adjust it to whatever detection or mode you prefer. |
|
Oh, interesting. I didn't expect any perf bump but didn't expect the regressions either. And yes, sorry for the delay on the INT HMX thingy. Kind of a long backlog right now but will definitely get back to you. |
Overview
Qwen3.5 decode concatenates the conv state with one token along dim 0 (rows of 3 + 1 f32, 8192 rows per layer) and copies the conv state view back (rows of 3 f32 at stride 4 into a contiguous dst). Both are short-row copies: the existing paths move a few bytes per row and wait on memory for every row.
This PR sends such copies through VTCM: one DMA brings the source span in,
vgatherplaces every word at its output position, and one DMA writes the rows out. The gather offsets are periodic (they repeat everyne / gcd(ne, 32)vectors), so they are built once per call and then advanced with a vector add.It is used for f32 CONCAT on dim 0 and f32 same-type reshape CPY, when rows have at most 16 elements, dst rows are contiguous, tensors are 2D and the session runs on a single device. Everything else takes the existing paths. New file:
htp/hvx-gather-rows.h.Additional information
This complements #29673, which sped up the same CONCAT. SM7750 (Snapdragon 7 Gen 4), Qwen3.5-4B Q4_0 decode, current master vs this PR:
3:8192 x 1:8192 -> 4:81923:8192at stride 4 -> contiguoustg64: 8.53 -> 8.93 t/s (+4.7%).
test-backend-ops:CONCAT48/48,CPY136/136.An earlier version of this change (before #29673) was also checked on a chip with FP16 HMX: the tests pass and decode was faster there too.
Found while working on #29473. It doesn't use HMX and doesn't depend on that work.
Requirements