Skip to content

hexagon: flatten matmul into 2d to use HMX in multi-sequence - #29779

Closed
jhen0409 wants to merge 1 commit into
ggml-org:masterfrom
jhen0409:jhen/hexagon-mm-flatten-3d
Closed

jhen0409 wants to merge 1 commit into
ggml-org:masterfrom
jhen0409:jhen/hexagon-mm-flatten-3d

Conversation

@jhen0409

@jhen0409 jhen0409 commented Oct 1, 2026

Copy link
Copy Markdown
Member

Overview

Cont #29344 (mentioned about ssm_out MUL_MAT). Flatten 3d src1/dst (original use HVX MATMUL_4D_REPACKED_IMPL) into 2d to use HMX, which is faster on n_seqs > 1. If the rows of dims 1..3 can be walked with one stride and src0 not batched, the flattening is possible.

Based on my investigation, models like lfm2.cpp, mamba-base.cpp (mamba, jamba build_mamba_layer), qwen35|qwen3next|qwen35moe|qwen4exp.cpp (ssm_out), plamo2.cpp will cover the path. (or maybe minimax-01.cpp and more)

Additional information

Test on IQ-9075, run two models (that cover the path) with llama-batched-bench --device HTP0:0 -ngl 99 -fa on -t 6 --cpu-mask 0xfc -c 4096 -b 512 -ub 256 -npp 128 -ntg 32 -npl 1,2,4:

Model npl test master (t/s) patch (t/s) change
Qwen3.5-2B Q4_0 1 pp128 673.17 673.00 -0.0%
Qwen3.5-2B Q4_0 1 tg32 20.68 20.66 -0.1%
Qwen3.5-2B Q4_0 2 pp128 x 2 452.70 719.49 +58.9%
Qwen3.5-2B Q4_0 2 tg32 x 2 29.53 30.19 +2.2%
Qwen3.5-2B Q4_0 4 pp128 x 4 445.93 703.51 +57.8%
Qwen3.5-2B Q4_0 4 tg32 x 4 33.58 34.18 +1.8%
LFM2-2.6B Q4_0 1 pp128 685.60 683.94 -0.2%
LFM2-2.6B Q4_0 1 tg32 22.32 22.64 +1.4%
LFM2-2.6B Q4_0 2 pp128 x 2 209.92 687.22 +227.4%
LFM2-2.6B Q4_0 2 tg32 x 2 27.62 28.13 +1.9%
LFM2-2.6B Q4_0 4 pp128 x 4 209.26 679.47 +224.7%
LFM2-2.6B Q4_0 4 tg32 x 4 30.74 31.47 +2.4%
Per-call MUL_MAT time with a batched src1 (GGML_HEXAGON_PROFILE=1, npl 2, one run)
Model npl weight MUL_MAT shape (src0 x src1 -> dst) calls master (ms) patch (ms) change
Qwen3.5-2B Q4_0 2 ssm_out 2048:2048 x 2048:128:2:1 -> 2048:128:2:1 18 12.203 0.545 -95.5%
Qwen3.5-2B Q4_0 2 ssm_out 2048:2048 x 2048:1:2:1 -> 2048:1:2:1 72 0.174 0.112 -35.9%
LFM2-2.6B Q4_0 2 shortconv.in_proj 2048:6144 x 2048:128:2:1 -> 6144:128:2:1 22 30.077 1.276 -95.8%
LFM2-2.6B Q4_0 2 shortconv.out_proj 2048:2048 x 2048:128:2:1 -> 2048:128:2:1 22 10.254 0.543 -94.7%
LFM2-2.6B Q4_0 2 shortconv.in_proj 2048:6144 x 2048:1:2:1 -> 6144:1:2:1 88 0.303 0.252 -16.9%
LFM2-2.6B Q4_0 2 shortconv.out_proj 2048:2048 x 2048:1:2:1 -> 2048:1:2:1 88 0.109 0.092 -15.6%

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, Claude Code, I did the validation and review.

@jhen0409
jhen0409 requested a review from a team as a code owner October 1, 2026 00:42
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Hexagon labels Oct 1, 2026
@jhen0409 jhen0409 closed this Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Hexagon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant