Skip to content

About

Integer HMX for llama.cpp on Snapdragon 7 Gen 4 (NPU without FP16): 2.4–2.9x prompt processing vs CPU

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Integer HMX for llama.cpp on Snapdragon 7 Gen 4

On Snapdragon 7 Gen 4 (SM7750, Hexagon v73), llama.cpp's Hexagon backend produces garbage: the HMX matrix unit on this chip has no FP16 support, while ggml-hexagon enables FP16 HMX based on the architecture version (v73+). Qualcomm's own runtime confirms it — QNN 2.50 on SM7750 logs Hexagon arch=73 ... fp16=false.

This work moves the math to integer HMX, which the chip does support. Result: prompt processing on the NPU is 2.2–3.3× faster than the best CPU configuration and 1.4–2.3× faster than the Adreno GPU, with the same perplexity as the CPU.

Status (7 October 2026)

  • Integer HMX is in review upstream as ggml-org/llama.cpp#29740 (draft), branch hexagon-int-hmx-nofp16: four commits on current master, integer HMX for Q4_0 matmul and attention prefill, chunked HVX gated delta net, HMX FP16 auto-detection, cleanup. The Hexagon maintainer confirmed the root cause in #29473: some lower-end v73 chips have no FP16 HMX.
  • v1.4 short-row copies were proposed as #29739 and closed in favour of the upstream copy rewrites (#29685, #30067). Neither takes these shapes: on SM7750 with #30067 the conv-state CPY takes 257 us per op (106 us before it, 20 us on the short-row path), CONCAT 142 us (26 us). The re-measure is posted in #29739.
  • Side contribution: NEON + i8mm vec_dot for PrismML's ternary PTQ1_0 models, merged as PrismML-Eng/llama.cpp#290.

Older branches: hexagon-int-hmx-pr (five commits on a97cce8, includes v1.4) and hexagon-int-hmx (the first version, linked in the issue).

Results

Motorola Edge 70 (Snapdragon 7 Gen 4, 12 GB), llama-bench pp512, tokens/s. CPU = best of 2/4/6/8 threads. GPU = stock OpenCL backend (Adreno 722).

Model NPU (this patch) CPU GPU vs CPU vs GPU
Llama-3.2-1B-Instruct Q4_0 550.6 226.8 307.2 2.43× 1.79×
Qwen3.5-4B Q4_0 (all layers q4_0) 171.9 51.7 73.9 3.33× 2.33×
Qwen3.5-4B Q4_0 (unsloth, mixed quants) 105.9 47.1 76.2 2.25× 1.39×

Quality, Qwen3.5-4B pure Q4_0, wikitext-2, 20 × 512 tokens:

Perplexity Time
CPU 10.8533 ± 0.424 4:02
NPU 10.8355 ± 0.423 1:08

On the test prompt, greedy outputs on the NPU were byte-identical to the CPU for all three models. On a longer, 300-token wikitext prompt the greedy text follows the CPU for about 10 tokens and then picks a different word; perplexity stays the same. test-backend-ops: MUL_MAT q4_0 38/38, FLASH_ATTN_EXT 2583/2588 (5 long-context cases, kv 8K–16K, still fail).

How it works

  • Weights. q4_0 blocks are re-quantized to int8 per column against the maximum block scale of the current K chunk. That scale goes into the HMX output conversion, shifted so the 16-bit store cannot overflow.
  • Activations. Stored as 16-bit and fed as two 8-bit planes (the uh:2x1 mode). A second bias word cancels the unsigned offset exactly.
  • Accumulation. HMX runs the integer matmul. Two stores per K chunk (coarse + fine) are stitched on HVX into the exact 32-bit sum, then rescaled per row and column.
  • Pipeline. HMX computes tile i while HVX reduces tile i−1 and converts weights for tile i+1.
  • Gated DeltaNet (Qwen3.5). The stock chunked GDN kernel needs FP16 HMX, so on this chip it falls back to a token-by-token HVX path. The branch adds a chunked HVX path in f32: 8 tokens per pass over the state, state kept transposed so every product is a vector AXPY, decay ratios in the log domain. GDN time 1.39 s -> 0.85 s per 512 tokens.
  • Attention. Q·Kᵀ and P·V both run on integer HMX. K is smoothed per channel, with the factor folded back into Q. Softmax runs on HVX in fp16, and the 1/sum factor is folded into the row scale.

The integer HMX behaviour was mapped by testing on the Hexagon SDK's x86 emulator, then verified on the phone with the same inputs: 12/12 output hashes match.

Which chips need this

Qualcomm AI Hub (QNN 2.50) reports HMX FP16 per chipset:

FP16 Chipsets
no Snapdragon 888 (v68), 778G (v68), QCS6490 (v68), 7 Gen 4 (v73), QCM6690 (v73)
yes 8 Gen 1 (v69), 8 Gen 2 (v73), 8 Gen 3 (v75), 8 Elite (v79), 8 Elite Gen 5 (v81), X Elite / X Plus (v73), X2 Elite (v81), QCS8550 / IQ-9075 (v73), QCS8275 (v75), SA8295P (v68), SA8775P (v73), SA7255P (v75)

On v73 the backend enables FP16 HMX, so 7 Gen 4 and QCM6690 are exposed. 7 Gen 4 is confirmed on a real device; QCM6690 is untested.

Decode (v1.4)

Decode on this phone is limited by memory bandwidth, not compute. NPU matmuls read weights at ~27 GB/s, the CPU at ~23 GB/s. Running CPU and NPU decode at the same time gives only ~15% more total throughput than either alone, so splitting decode between them does not pay off.

On the NPU, Qwen3.5 decode also lost ~8 ms per token in two small ops: CONCAT of the conv state with the new token and the CPY of the conv state back. They moved 8192 rows of 3-4 floats one row at a time and waited on memory for every row. v1.4 moves such short rows through VTCM: one DMA in, vgather to place every word, one DMA out.

Qwen3.5-4B Q4_0, per token before v1.4
CONCAT (conv state) 5.6 ms 0.66 ms
CPY (conv state) 2.6 ms 0.52 ms
tg64, NPU 8.31 t/s 8.99 t/s

The CPU is still a bit faster at decode: 9.4-9.7 t/s (Qwen3.5-4B), 33.5 t/s (Llama-3.2-1B, NPU 30.1). Tests: CONCAT 48/48, CPY 136/136, GATED_DELTA_NET 36/36, MUL_MAT 759/759; output unchanged; pp512 unchanged (171.5).

Limits

  • HMX runs prefill only. Decode stays on HVX and the CPU is still slightly faster (see above).
  • Splitting one model's layers across NPU + GPU is slower than the NPU alone: 364–397 vs 456 t/s (same build) on Llama-3.2-1B pp512. Layers run one after another, so the speed lands between the two devices.
  • Small batches are slower than the CPU: 8 tokens take 336 ms on the NPU vs 203 ms on the CPU (Qwen3.5-4B). The weight conversion costs the same for 8 tokens as for 512.
  • Only q4_0 weights run on HMX. Other types stay on CPU/HVX, which is why mixed-quant models gain less.
  • Tested on one device.
  • Headroom: HMX is idle most of the time; HVX weight conversion and output reduction are the bottleneck. Qualcomm's QNN reaches ~9.6 TFLOPS (w4a16) on the same chip.

Build and run

Same as upstream llama.cpp for Snapdragon (docs/backend/snapdragon), on the PR branch (the v1.4 copies are only on hexagon-int-hmx-pr):

git clone -b hexagon-int-hmx-nofp16 https://github.com/karusrus/llama.cpp
cd llama.cpp
docker run -it --rm -u $(id -u):$(id -g) --volume $(pwd):/workspace --platform linux/amd64 \
  ghcr.io/snapdragon-toolchain/arm64-android:v0.7
# inside the container
cp docs/backend/snapdragon/CMakeUserPresets.json .
cmake --preset arm64-android-snapdragon-release -B build-snapdragon
cmake --build build-snapdragon
cmake --install build-snapdragon --prefix pkg-android/llama.cpp

On the phone no settings are needed:

./bin/llama-bench -m model.gguf -p 512 -n 0 -ngl 99 --device HTP0
  • At session start the backend runs one FP16 HMX tile and checks the result. On SM7750 it logs HMX FP16: no (integer HMX for Q4_0 matmul and attention) and picks: integer HMX for Q4_0 matmuls and attention prefill, HVX for other weight types and for gated delta net. On FP16-capable chips nothing changes.
  • GGML_HEXAGON_HMX_FP16=0|1 forces the choice (1 on SM7750 reproduces the original garbage).
  • Use a model where all matmul weights are Q4_0 (llama-quantize --pure ... Q4_0) for the full speedup.

Repo layout

patches/ — the five commits of the branch on top of a97cce8, as a patch (MIT, like llama.cpp) · src/ — the integer HMX kernels · aihub/ — FP16 chip map and QNN ceiling scripts · scripts/ — benchmark scripts run on the phone · data/ — raw logs.

About

Integer HMX for llama.cpp on Snapdragon 7 Gen 4 (NPU without FP16): 2.4–2.9x prompt processing vs CPU

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages