Regression: --lazy-mode auto halves prompt-processing (pp512) for qwen4exp on Vulkan (AMD iGPU)
Summary
After commit 257813839 ("llama: improve TENSOR_READ_LAZY handling (#27837)"), the default
--lazy-mode auto roughly halves the prefill throughput of the qwen4exp architecture
(Qwen3.8-Flash-Next, IQ4_XS) on the Vulkan backend of an AMD Strix Halo iGPU
(Radeon 8060S / gfx1151). Setting --lazy-mode off restores the previous performance.
- Reported build:
0.3.0-dev (build 10723, commit 010be9683)
- bisected first-bad commit:
257813839
- Good baseline (before the commit, build 10662 /
18443257a): pp512 ≈ 429 t/s
How to reproduce
Test with llama-bench, model Qwen3.8-Flash-Next IQ4_XS (qwen4exp arch, ~87 GiB),
Vulkan backend, full offload:
llama-bench -m Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf -ngl 99 -p 512 -n 128 # lazy-mode auto (default)
llama-bench -m Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf -ngl 99 -p 512 -n 128 --lazy-mode off
Results (pp512 t/s)
| Configuration |
pp512 (t/s) |
tg128 (t/s) |
--lazy-mode auto (default, HEAD) |
216 |
24.1 |
--lazy-mode off (HEAD) |
406 |
24.2 |
pre-commit 257813839 (build 10662) |
429 |
25.6 |
Same model, same build, same backend — only the --lazy-mode value differs. tg128
is unaffected; only prompt processing regresses.
Environment
- CPU/GPU: AMD Ryzen AI MAX+ 395 "Strix Halo" (gfx1151), Radeon 8060S iGPU
- Memory: 128 GB LPDDR5X unified
- Backend: Vulkan (
radv, mesa 26.0); device detected as
Radeon 8060S Graphics (RADV STRIX_HALO), uma: 1
- Kernel:
7.2.0-070200-generic
- Build: Release,
build-vulkan (CPU native + Vulkan), commit 010be9683
Bisect log (first-bad)
good: 0b5be7e4a hip: tune rdna 3 mmq config (#26284) -> pp512 ≈ 410
bad : 257813839 llama: improve TENSOR_READ_LAZY handling (#27837) -> pp512 ≈ 218
Note: I also tested reverting the GDN/LID commit 8663224810 ("context: disable
non-fused GDN and LID ops") as a suspected cause — reverting it had no effect, so the
lazy-read commit is the confirmed culprit.
Suspected cause
llama_model_loader::lazy_read now lazily maps tensors during load even for the
auto mode. On Vulkan with this iGPU/UMA setup, the lazy-read path likely produces
suboptimal buffer placement (or host-visible staging) for the large embedding/SSM
tensors of the qwen4exp architecture, making the prefill read weights from a slower path.
Expected behavior
--lazy-mode auto should not degrade prompt processing relative to pre-257813839
behavior. At minimum, the auto heuristic should avoid lazy-reading tensors when the
backend is Vulkan (or when doing so forces a non-device-local buffer on UMA iGPUs).
Workaround
Add --lazy-mode off to the llama-server/llama-bench invocation. This restores pp512
to ~406 t/s (vs 216 with auto). Applied in our start_llamacpp.sh for all models.
Regression:
--lazy-mode autohalves prompt-processing (pp512) for qwen4exp on Vulkan (AMD iGPU)Summary
After commit
257813839("llama: improve TENSOR_READ_LAZY handling (#27837)"), the default--lazy-mode autoroughly halves the prefill throughput of the qwen4exp architecture(
Qwen3.8-Flash-Next, IQ4_XS) on the Vulkan backend of an AMD Strix Halo iGPU(Radeon 8060S / gfx1151). Setting
--lazy-mode offrestores the previous performance.0.3.0-dev (build 10723, commit 010be9683)25781383918443257a): pp512 ≈ 429 t/sHow to reproduce
Test with
llama-bench, modelQwen3.8-Flash-NextIQ4_XS (qwen4exp arch, ~87 GiB),Vulkan backend, full offload:
llama-bench -m Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf -ngl 99 -p 512 -n 128 # lazy-mode auto (default) llama-bench -m Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf -ngl 99 -p 512 -n 128 --lazy-mode offResults (pp512 t/s)
--lazy-mode auto(default, HEAD)--lazy-mode off(HEAD)257813839(build 10662)Same model, same build, same backend — only the
--lazy-modevalue differs.tg128is unaffected; only prompt processing regresses.
Environment
radv, mesa 26.0); device detected asRadeon 8060S Graphics (RADV STRIX_HALO), uma: 17.2.0-070200-genericbuild-vulkan(CPU native + Vulkan), commit010be9683Bisect log (first-bad)
Note: I also tested reverting the GDN/LID commit
8663224810("context: disablenon-fused GDN and LID ops") as a suspected cause — reverting it had no effect, so the
lazy-read commit is the confirmed culprit.
Suspected cause
llama_model_loader::lazy_readnow lazily maps tensors during load even for theautomode. On Vulkan with this iGPU/UMA setup, the lazy-read path likely producessuboptimal buffer placement (or host-visible staging) for the large embedding/SSM
tensors of the qwen4exp architecture, making the prefill read weights from a slower path.
Expected behavior
--lazy-mode autoshould not degrade prompt processing relative to pre-257813839behavior. At minimum, the
autoheuristic should avoid lazy-reading tensors when thebackend is Vulkan (or when doing so forces a non-device-local buffer on UMA iGPUs).
Workaround
Add
--lazy-mode offto the llama-server/llama-bench invocation. This restores pp512to ~406 t/s (vs 216 with
auto). Applied in ourstart_llamacpp.shfor all models.