Skip to content

Regression: --lazy-mode auto halves pp512 for qwen4exp on Vulkan (AMD iGPU) #28160

Description

@ilariofebi

Regression: --lazy-mode auto halves prompt-processing (pp512) for qwen4exp on Vulkan (AMD iGPU)

Summary

After commit 257813839 ("llama: improve TENSOR_READ_LAZY handling (#27837)"), the default
--lazy-mode auto roughly halves the prefill throughput of the qwen4exp architecture
(Qwen3.8-Flash-Next, IQ4_XS) on the Vulkan backend of an AMD Strix Halo iGPU
(Radeon 8060S / gfx1151). Setting --lazy-mode off restores the previous performance.

  • Reported build: 0.3.0-dev (build 10723, commit 010be9683)
  • bisected first-bad commit: 257813839
  • Good baseline (before the commit, build 10662 / 18443257a): pp512 ≈ 429 t/s

How to reproduce

Test with llama-bench, model Qwen3.8-Flash-Next IQ4_XS (qwen4exp arch, ~87 GiB),
Vulkan backend, full offload:

llama-bench -m Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf -ngl 99 -p 512 -n 128          # lazy-mode auto (default)
llama-bench -m Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf -ngl 99 -p 512 -n 128 --lazy-mode off

Results (pp512 t/s)

Configuration pp512 (t/s) tg128 (t/s)
--lazy-mode auto (default, HEAD) 216 24.1
--lazy-mode off (HEAD) 406 24.2
pre-commit 257813839 (build 10662) 429 25.6

Same model, same build, same backend — only the --lazy-mode value differs. tg128
is unaffected; only prompt processing regresses.

Environment

  • CPU/GPU: AMD Ryzen AI MAX+ 395 "Strix Halo" (gfx1151), Radeon 8060S iGPU
  • Memory: 128 GB LPDDR5X unified
  • Backend: Vulkan (radv, mesa 26.0); device detected as
    Radeon 8060S Graphics (RADV STRIX_HALO), uma: 1
  • Kernel: 7.2.0-070200-generic
  • Build: Release, build-vulkan (CPU native + Vulkan), commit 010be9683

Bisect log (first-bad)

good: 0b5be7e4a hip: tune rdna 3 mmq config (#26284)     -> pp512 ≈ 410
bad : 257813839 llama: improve TENSOR_READ_LAZY handling (#27837) -> pp512 ≈ 218

Note: I also tested reverting the GDN/LID commit 8663224810 ("context: disable
non-fused GDN and LID ops") as a suspected cause — reverting it had no effect, so the
lazy-read commit is the confirmed culprit.

Suspected cause

llama_model_loader::lazy_read now lazily maps tensors during load even for the
auto mode. On Vulkan with this iGPU/UMA setup, the lazy-read path likely produces
suboptimal buffer placement (or host-visible staging) for the large embedding/SSM
tensors of the qwen4exp architecture, making the prefill read weights from a slower path.

Expected behavior

--lazy-mode auto should not degrade prompt processing relative to pre-257813839
behavior. At minimum, the auto heuristic should avoid lazy-reading tensors when the
backend is Vulkan (or when doing so forces a non-device-local buffer on UMA iGPUs).

Workaround

Add --lazy-mode off to the llama-server/llama-bench invocation. This restores pp512
to ~406 t/s (vs 216 with auto). Applied in our start_llamacpp.sh for all models.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions