Summary
Token generation for the qwen35 architecture (hybrid: 48/64 Gated DeltaNet linear-attention layers + 16 full-attention) runs at ~28% of the memory-bandwidth bound on an RTX 5090 (sm_120, Windows), while the same GGUF reaches ~86% of the bound on an RTX 4090 (Ada, Linux). A MoE model on the identical rig/build decodes fast, so the GPU, driver, and Windows stack appear healthy — the deficit looks specific to the qwen35 decode path on consumer Blackwell.
Environment
- GPU: NVIDIA GeForce RTX 5090 (32 GB GDDR7, driver 616.56)
- OS: Windows 11 (build 26220)
- llama.cpp: b10741 (build
9d817213a), official release binaries llama-b10741-bin-win-cuda-13.3-x64.zip + matching cudart bundle
- Model: Qwen3.5-27B Q4_K_M GGUF, 15.65 GiB, arch reported as
qwen35 27B Q4_K - Medium (pulled via the Ollama registry, tag qwen3.8:27b)
Measurements
llama-bench -m <model> -ngl 99 -fa 1 -p 1024 -n 400 (defaults otherwise, 5 reps):
| model |
test |
t/s |
| qwen35 27B Q4_K_M |
pp1024 |
1515.83 ± 47.03 |
| qwen35 27B Q4_K_M |
tg400 |
30.09 ± 0.44 |
| qwen3moe 30B.A3B Q4_K_M (control, same rig) |
pp1024 |
4081.20 ± 89.86 |
| qwen3moe 30B.A3B Q4_K_M (control, same rig) |
tg400 |
134.18 ± 1.84 |
Expected
Single-token decode on a dense(-per-token) 16.8 GB model is memory-bandwidth bound: 1792 GB/s ÷ 16.8 GB ≈ ~107 t/s ceiling. Measured 30.09 ≈ 28% of that.
Reference on Ada: an RTX 4090 (Linux, Ollama 0.32.x engine, byte-identical GGUF, speculative decoding disabled) decodes the same model at 48.8 t/s ≈ 86% of its own ~57 t/s ceiling. So the qwen35 path reaches near-roofline on sm_89 but not on sm_120.
Ruled out on this rig
- GPU state: P0 throughout decode, core ~3030–3050 MHz, memory clock at full 14014 MHz (28 Gbps), ~300 W, 93–98% utilization (sampled live once per second during tg).
- Host link: the rig happened to run at PCIe Gen5 x1 and later Gen5 x4 (marginal slot, being fixed) — tg was 27.95 vs 30.09 t/s respectively, so decode is not host-link-bound.
- Rig/stack health: the MoE control above (134 t/s) rules out any global Windows/driver per-token overhead large enough to explain the qwen35 number.
- Flash attention / KV cache type: FA on/off and KV q8_0 vs f16 were A/B-tested on the same GGUF (via Ollama's runner) — no material change (FA off is slightly worse).
Possibly related
Under Ollama's runner (which enables its multi-token-prediction path for this model), llama-server on this machine crashed twice at different draft_num_predict values with:
exit status 0xc0000409: stack-based buffer overrun
: CUDA error: shared object initialization failed
Mentioning it in case the sm_120 code path is shared; I can open a separate issue with debug logs if useful.
Happy to run traces (nsys/ncu), patched builds, or additional A/Bs on this hardware.
Summary
Token generation for the
qwen35architecture (hybrid: 48/64 Gated DeltaNet linear-attention layers + 16 full-attention) runs at ~28% of the memory-bandwidth bound on an RTX 5090 (sm_120, Windows), while the same GGUF reaches ~86% of the bound on an RTX 4090 (Ada, Linux). A MoE model on the identical rig/build decodes fast, so the GPU, driver, and Windows stack appear healthy — the deficit looks specific to the qwen35 decode path on consumer Blackwell.Environment
9d817213a), official release binariesllama-b10741-bin-win-cuda-13.3-x64.zip+ matching cudart bundleqwen35 27B Q4_K - Medium(pulled via the Ollama registry, tagqwen3.8:27b)Measurements
llama-bench -m <model> -ngl 99 -fa 1 -p 1024 -n 400(defaults otherwise, 5 reps):Expected
Single-token decode on a dense(-per-token) 16.8 GB model is memory-bandwidth bound: 1792 GB/s ÷ 16.8 GB ≈ ~107 t/s ceiling. Measured 30.09 ≈ 28% of that.
Reference on Ada: an RTX 4090 (Linux, Ollama 0.32.x engine, byte-identical GGUF, speculative decoding disabled) decodes the same model at 48.8 t/s ≈ 86% of its own ~57 t/s ceiling. So the qwen35 path reaches near-roofline on sm_89 but not on sm_120.
Ruled out on this rig
Possibly related
Under Ollama's runner (which enables its multi-token-prediction path for this model),
llama-serveron this machine crashed twice at differentdraft_num_predictvalues with:Mentioning it in case the sm_120 code path is shared; I can open a separate issue with debug logs if useful.
Happy to run traces (nsys/ncu), patched builds, or additional A/Bs on this hardware.