Skip to content

CUDA: qwen35 on RTX 5090 (sm_120) decodes at 76% of roofline on native Linux with MTP ~1.7x; Windows/Ollama path runs 1.5-1.6x slower at every draft depth #28196

Description

@Majesty401

Summary

Token generation for the qwen35 architecture (hybrid: 48/64 Gated DeltaNet linear-attention layers + 16 full-attention) runs at ~28% of the memory-bandwidth bound on an RTX 5090 (sm_120, Windows), while the same GGUF reaches ~86% of the bound on an RTX 4090 (Ada, Linux). A MoE model on the identical rig/build decodes fast, so the GPU, driver, and Windows stack appear healthy — the deficit looks specific to the qwen35 decode path on consumer Blackwell.

Environment

  • GPU: NVIDIA GeForce RTX 5090 (32 GB GDDR7, driver 616.56)
  • OS: Windows 11 (build 26220)
  • llama.cpp: b10741 (build 9d817213a), official release binaries llama-b10741-bin-win-cuda-13.3-x64.zip + matching cudart bundle
  • Model: Qwen3.5-27B Q4_K_M GGUF, 15.65 GiB, arch reported as qwen35 27B Q4_K - Medium (pulled via the Ollama registry, tag qwen3.8:27b)

Measurements

llama-bench -m <model> -ngl 99 -fa 1 -p 1024 -n 400 (defaults otherwise, 5 reps):

model test t/s
qwen35 27B Q4_K_M pp1024 1515.83 ± 47.03
qwen35 27B Q4_K_M tg400 30.09 ± 0.44
qwen3moe 30B.A3B Q4_K_M (control, same rig) pp1024 4081.20 ± 89.86
qwen3moe 30B.A3B Q4_K_M (control, same rig) tg400 134.18 ± 1.84

Expected

Single-token decode on a dense(-per-token) 16.8 GB model is memory-bandwidth bound: 1792 GB/s ÷ 16.8 GB ≈ ~107 t/s ceiling. Measured 30.09 ≈ 28% of that.

Reference on Ada: an RTX 4090 (Linux, Ollama 0.32.x engine, byte-identical GGUF, speculative decoding disabled) decodes the same model at 48.8 t/s ≈ 86% of its own ~57 t/s ceiling. So the qwen35 path reaches near-roofline on sm_89 but not on sm_120.

Ruled out on this rig

  • GPU state: P0 throughout decode, core ~3030–3050 MHz, memory clock at full 14014 MHz (28 Gbps), ~300 W, 93–98% utilization (sampled live once per second during tg).
  • Host link: the rig happened to run at PCIe Gen5 x1 and later Gen5 x4 (marginal slot, being fixed) — tg was 27.95 vs 30.09 t/s respectively, so decode is not host-link-bound.
  • Rig/stack health: the MoE control above (134 t/s) rules out any global Windows/driver per-token overhead large enough to explain the qwen35 number.
  • Flash attention / KV cache type: FA on/off and KV q8_0 vs f16 were A/B-tested on the same GGUF (via Ollama's runner) — no material change (FA off is slightly worse).

Possibly related

Under Ollama's runner (which enables its multi-token-prediction path for this model), llama-server on this machine crashed twice at different draft_num_predict values with:

exit status 0xc0000409: stack-based buffer overrun
: CUDA error: shared object initialization failed

Mentioning it in case the sm_120 code path is shared; I can open a separate issue with debug logs if useful.

Happy to run traces (nsys/ncu), patched builds, or additional A/Bs on this hardware.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions