On Snapdragon 7 Gen 4 (SM7750, Hexagon v73), llama.cpp's Hexagon backend produces garbage:
the HMX matrix unit on this chip has no FP16 support, while ggml-hexagon enables FP16 HMX
based on the architecture version (v73+). Qualcomm's own runtime confirms it — QNN 2.50 on
SM7750 logs Hexagon arch=73 ... fp16=false.
This work moves the math to integer HMX, which the chip does support. Result: prompt processing on the NPU is 2.2–3.3× faster than the best CPU configuration and 1.4–2.3× faster than the Adreno GPU, with the same perplexity as the CPU.
- Integer HMX is in review upstream as ggml-org/llama.cpp#29740 (draft),
branch
hexagon-int-hmx-nofp16: four commits on current master, integer HMX for Q4_0 matmul and attention prefill, chunked HVX gated delta net, HMX FP16 auto-detection, cleanup. The Hexagon maintainer confirmed the root cause in #29473: some lower-end v73 chips have no FP16 HMX. - v1.4 short-row copies were proposed as #29739 and closed in favour of the upstream copy rewrites (#29685, #30067). Neither takes these shapes: on SM7750 with #30067 the conv-state CPY takes 257 us per op (106 us before it, 20 us on the short-row path), CONCAT 142 us (26 us). The re-measure is posted in #29739.
- Side contribution: NEON + i8mm
vec_dotfor PrismML's ternary PTQ1_0 models, merged as PrismML-Eng/llama.cpp#290.
Older branches: hexagon-int-hmx-pr (five commits on a97cce8,
includes v1.4) and hexagon-int-hmx (the first version, linked in the issue).
Motorola Edge 70 (Snapdragon 7 Gen 4, 12 GB), llama-bench pp512, tokens/s.
CPU = best of 2/4/6/8 threads. GPU = stock OpenCL backend (Adreno 722).
| Model | NPU (this patch) | CPU | GPU | vs CPU | vs GPU |
|---|---|---|---|---|---|
| Llama-3.2-1B-Instruct Q4_0 | 550.6 | 226.8 | 307.2 | 2.43× | 1.79× |
| Qwen3.5-4B Q4_0 (all layers q4_0) | 171.9 | 51.7 | 73.9 | 3.33× | 2.33× |
| Qwen3.5-4B Q4_0 (unsloth, mixed quants) | 105.9 | 47.1 | 76.2 | 2.25× | 1.39× |
Quality, Qwen3.5-4B pure Q4_0, wikitext-2, 20 × 512 tokens:
| Perplexity | Time | |
|---|---|---|
| CPU | 10.8533 ± 0.424 | 4:02 |
| NPU | 10.8355 ± 0.423 | 1:08 |
On the test prompt, greedy outputs on the NPU were byte-identical to the CPU for all three models. On a longer,
300-token wikitext prompt the greedy text follows the CPU for about 10 tokens and then picks a different word;
perplexity stays the same.
test-backend-ops: MUL_MAT q4_0 38/38, FLASH_ATTN_EXT 2583/2588 (5 long-context cases, kv 8K–16K, still fail).
- Weights. q4_0 blocks are re-quantized to int8 per column against the maximum block scale of the current K chunk. That scale goes into the HMX output conversion, shifted so the 16-bit store cannot overflow.
- Activations. Stored as 16-bit and fed as two 8-bit planes (the
uh:2x1mode). A second bias word cancels the unsigned offset exactly. - Accumulation. HMX runs the integer matmul. Two stores per K chunk (coarse + fine) are stitched on HVX into the exact 32-bit sum, then rescaled per row and column.
- Pipeline. HMX computes tile i while HVX reduces tile i−1 and converts weights for tile i+1.
- Gated DeltaNet (Qwen3.5). The stock chunked GDN kernel needs FP16 HMX, so on this chip it falls back to a token-by-token HVX path. The branch adds a chunked HVX path in f32: 8 tokens per pass over the state, state kept transposed so every product is a vector AXPY, decay ratios in the log domain. GDN time 1.39 s -> 0.85 s per 512 tokens.
- Attention. Q·Kᵀ and P·V both run on integer HMX. K is smoothed per channel, with the factor folded back into Q. Softmax runs on HVX in fp16, and the 1/sum factor is folded into the row scale.
The integer HMX behaviour was mapped by testing on the Hexagon SDK's x86 emulator, then verified on the phone with the same inputs: 12/12 output hashes match.
Qualcomm AI Hub (QNN 2.50) reports HMX FP16 per chipset:
| FP16 | Chipsets |
|---|---|
| no | Snapdragon 888 (v68), 778G (v68), QCS6490 (v68), 7 Gen 4 (v73), QCM6690 (v73) |
| yes | 8 Gen 1 (v69), 8 Gen 2 (v73), 8 Gen 3 (v75), 8 Elite (v79), 8 Elite Gen 5 (v81), X Elite / X Plus (v73), X2 Elite (v81), QCS8550 / IQ-9075 (v73), QCS8275 (v75), SA8295P (v68), SA8775P (v73), SA7255P (v75) |
On v73 the backend enables FP16 HMX, so 7 Gen 4 and QCM6690 are exposed. 7 Gen 4 is confirmed on a real device; QCM6690 is untested.
Decode on this phone is limited by memory bandwidth, not compute. NPU matmuls read weights at ~27 GB/s, the CPU at ~23 GB/s. Running CPU and NPU decode at the same time gives only ~15% more total throughput than either alone, so splitting decode between them does not pay off.
On the NPU, Qwen3.5 decode also lost ~8 ms per token in two small ops: CONCAT of the conv state with the new token and
the CPY of the conv state back. They moved 8192 rows of 3-4 floats one row at a time and waited on memory for every row.
v1.4 moves such short rows through VTCM: one DMA in, vgather to place every word, one DMA out.
| Qwen3.5-4B Q4_0, per token | before | v1.4 |
|---|---|---|
| CONCAT (conv state) | 5.6 ms | 0.66 ms |
| CPY (conv state) | 2.6 ms | 0.52 ms |
| tg64, NPU | 8.31 t/s | 8.99 t/s |
The CPU is still a bit faster at decode: 9.4-9.7 t/s (Qwen3.5-4B), 33.5 t/s (Llama-3.2-1B, NPU 30.1). Tests: CONCAT 48/48, CPY 136/136, GATED_DELTA_NET 36/36, MUL_MAT 759/759; output unchanged; pp512 unchanged (171.5).
- HMX runs prefill only. Decode stays on HVX and the CPU is still slightly faster (see above).
- Splitting one model's layers across NPU + GPU is slower than the NPU alone: 364–397 vs 456 t/s (same build) on Llama-3.2-1B pp512. Layers run one after another, so the speed lands between the two devices.
- Small batches are slower than the CPU: 8 tokens take 336 ms on the NPU vs 203 ms on the CPU (Qwen3.5-4B). The weight conversion costs the same for 8 tokens as for 512.
- Only q4_0 weights run on HMX. Other types stay on CPU/HVX, which is why mixed-quant models gain less.
- Tested on one device.
- Headroom: HMX is idle most of the time; HVX weight conversion and output reduction are the bottleneck. Qualcomm's QNN reaches ~9.6 TFLOPS (w4a16) on the same chip.
Same as upstream llama.cpp for Snapdragon (docs/backend/snapdragon), on the PR branch
(the v1.4 copies are only on hexagon-int-hmx-pr):
git clone -b hexagon-int-hmx-nofp16 https://github.com/karusrus/llama.cpp
cd llama.cpp
docker run -it --rm -u $(id -u):$(id -g) --volume $(pwd):/workspace --platform linux/amd64 \
ghcr.io/snapdragon-toolchain/arm64-android:v0.7
# inside the container
cp docs/backend/snapdragon/CMakeUserPresets.json .
cmake --preset arm64-android-snapdragon-release -B build-snapdragon
cmake --build build-snapdragon
cmake --install build-snapdragon --prefix pkg-android/llama.cpp
On the phone no settings are needed:
./bin/llama-bench -m model.gguf -p 512 -n 0 -ngl 99 --device HTP0
- At session start the backend runs one FP16 HMX tile and checks the result. On SM7750 it logs
HMX FP16: no (integer HMX for Q4_0 matmul and attention)and picks: integer HMX for Q4_0 matmuls and attention prefill, HVX for other weight types and for gated delta net. On FP16-capable chips nothing changes. GGML_HEXAGON_HMX_FP16=0|1forces the choice (1 on SM7750 reproduces the original garbage).- Use a model where all matmul weights are Q4_0 (
llama-quantize --pure ... Q4_0) for the full speedup.
patches/ — the five commits of the branch on top of a97cce8, as a patch (MIT, like llama.cpp) · src/ — the integer HMX kernels ·
aihub/ — FP16 chip map and QNN ceiling scripts · scripts/ — benchmark scripts run on the phone · data/ — raw logs.