Repository navigation
SYCL MTP on Intel Arc: correct output but no speed gain over baseline #23533
Description
Activity
@R-SITES
Yes, it's known issue.
It will take a little more time to check.We will check if SYCL graph can help it since oneAPI 2026.0 has support it officially.
Welcome more suggestion!Thank you!
Reacted by R1SITES, karavayev and Thicc-fil-a@arthw — thanks for confirming this is on your radar. Here is additional data from thorough testing on Intel Arc Pro B70 (Battlemage, PCI 8086:e223) that may help narrow the investigation:
Test matrix (b9292, oneAPI 2025.3.3, Qwen3.6-35B-A3B MoE Q4_K_M):
Config Tok/s Draft Accept Output No MTP 73.3 — ✅ MTP n_max=2 57.8 26/26 (100%) ✅ MTP n_max=3 ~48 ~80% ✅ Vulkan MTP n_max=2 53 ~81% ✅ What we tried for SYCL MTP speed (none moved the needle):
GGML_SYCL_MMV_Y=4— register spill, dropped to 25 tok/s. MMV_Y=1 is optimal for Arc EU register fileGGML_SYCL_DISABLE_GRAPH=0re-enabled graphs — MUL_MAT_ID blockingstream->wait()prevents full pipeline recording. The fused MMVQ path dispatches 8 experts in one kernel, but the MoE routing copy callsstream->wait()for CPU-side expert selection, which cannot be captured- Manual fusion attempt (sigmoid+mul→single kernel) — compiled correctly but did not change throughput; saving 1 kernel launch out of ~960 per token is negligible
- oneAPI 2026.0 DPC++ compiler — ggml-sycl backend does not compile with the 2026.0 compiler. Additionally, the 2026.0 runtime requires
urDeviceWaitExpfrom Level Zero v2 API not yet in the public driver (latest: 26.18.38308.1) - MoE prefill PR [SYCL] improve MoE prefill throughput (+70% with Qwen3.6-35B) #23142 (applied) — improved prefill, +12% on MTP generation
- Multi-column MMVQ PR sycl : port multi-column MMVQ from CUDA backend (~45% speculative decoding speedup on Intel Arc) #21845 — broke output on Qwen3.6 MoE
Root cause analysis:
Each token requires ~960 expert matmul dispatches. With MTP, verification adds another batch. The kernel launch overhead (100-500µs vs ~5µs CUDA) dominates. 100% draft acceptance confirms this is dispatch overhead, not compute.
What would help:
- Record the full MoE layer graph as a single SYCL graph replayable with zero dispatch overhead
- The MUL_MAT_ID blocking
wait()needs to become event-based - Alternatively, pre-compute expert mappings on GPU and remove CPU-in-the-loop routing
Additional note on oneAPI 2026.0 / Ubuntu 26.04 compatibility:
The oneAPI 2026.0 compiler and runtime are installed but cannot currently function on Ubuntu 26.04 because Intel GPU driver repo (https://apt.repos.intel.com/gpu/ubuntu) does not yet list the "resolute" codename (returns 404). Once the matching Level Zero v2 driver lands in either the Intel GPU repo or on the compute-runtime GitHub releases page, we can immediately rebuild with 2026.0 and test SYCL graph support.
Happy to test any experimental branch — hardware is available for immediate testing.
@R-SITES
Thank you for your check and sharing!Looks like SYCL Graph is still out of work in SYCL backend.
I will check other issue.
Ubuntu 26.04 is not recommended to normal user:
- it can't be upgrade from old Ubuntu, need to install from scatch.
- It's still unstable for unknown reason.
Thank you!
Reacted by Matt Kucia@arthw — t
Additional note on oneAPI 2026.0 / Ubuntu 26.04 compatibility:oneAPI 2026.0 DPC++ compiler — ggml-sycl backend does not compile with the 2026.0 compiler. Additionally, the 2026.0
runtime requires urDeviceWaitExp from Level Zero v2 API not yet in the public driver (latest: 26.18.38308.1)The oneAPI 2026.0 compiler and runtime are installed but cannot currently function on Ubuntu 26.04 because Intel GPU driver repo (https://apt.repos.intel.com/gpu/ubuntu) does not yet list the "resolute" codename (returns 404). Once the matching Level Zero v2 driver lands in either the Intel GPU repo or on the compute-runtime GitHub releases page, we can immediately rebuild with 2026.0 and test SYCL graph support.
g.I am on kubuntu 26.04, I am on the latest 26.18.38308.1 driver. and oneapi version 2026.0. I use the ppa:kobuk-team/intel-graphics repository to get the latest intel driver for Ubuntu. Are you using this ppa?
I compile llama.cpp from source.. I have not had any issues until now on compile, except for I have to switch to icpx and icx compilers in the configuration to get the correct behavior even for vulkan build.
It would be great if mtp performance gets fixed for sycl, for right now vulkan on linux is 40% slower on linux compared to windows. so any speed gain barely faster than sycl without mtp on linux.
@agsharathnaik thanks for the info! We're also using the kobuk-team/testing PPA with the 26.18.38308.1 driver. But when we try oneAPI 2026.0, we get:
icpx: error: linker command failed symbol lookup error: libsycl.so.9: undefined symbol: urDeviceWaitExpCould you share your cmake command and any env setup you use to get 2026.0 compiling and running? Specifically:
- Do you use
source setvars.sh --forceor manually set PATH/LD_LIBRARY_PATH? - What cmake flags do you pass for SYCL?
- Does the compiled binary actually run without the urDeviceWaitExp error, or do you also see that at runtime?
We'd love to replicate your working 2026.0 setup and test SYCL graph support.
- Do you use
@agsharathnaik
Yes, I use ppa:kobuk-team/intel-graphics repository to install the GPU driver too.
To verify some issues, I will install the new or old driver from the github release: https://github.com/intel/compute-runtime/releases.for error:
icpx: error: linker command failed symbol lookup error: libsycl.so.9: undefined symbol: urDeviceWaitExprm -r build, then build again.Yes, we will check the MTP performance issue.
@agsharathnaik thanks for the info! We're also using the kobuk-team/testing PPA with the 26.18.38308.1 driver. But when we try oneAPI 2026.0, we get:
icpx: error: linker command failed symbol lookup error: libsycl.so.9: undefined symbol: urDeviceWaitExpCould you share your cmake command and any env setup you use to get 2026.0 compiling and running? Specifically:
- Do you use
source setvars.sh --forceor manually set PATH/LD_LIBRARY_PATH? - What cmake flags do you pass for SYCL?
- Does the compiled binary actually run without the urDeviceWaitExp error, or do you also see that at runtime?
We'd love to replicate your working 2026.0 setup and test SYCL graph support.
I use the below directly in my .profile file. Since I pretty much use this compiler for everything.
# Intel one api environment setup source /opt/intel/oneapi/setvars.shBut I also always compile using static build(no shared library) to avoid library visibility issues. Since I use VScode directly to build this I just alter the CMakePresets.json file to use the preset. The reason for fully switching to intel compiler is I use --cpu-moe for larger models.. this has a problem with vulkan using the gcc compiled runtime never accepting the thread count correctly.
{ "name": "sycl-base", "hidden": true, "generator": "Ninja", "binaryDir": "${sourceDir}/build-${presetName}", "cacheVariables": { "CMAKE_EXPORT_COMPILE_COMMANDS": "ON", "CMAKE_CXX_COMPILER": "icpx", "CMAKE_C_COMPILER": "icx", "GGML_SYCL": "ON", "CMAKE_INSTALL_RPATH": "$ORIGIN;$ORIGIN/.." } { "name": "base", "hidden": true, "generator": "Ninja", "binaryDir": "${sourceDir}/build-${presetName}", "cacheVariables": { "CMAKE_EXPORT_COMPILE_COMMANDS": "ON", "CMAKE_C_COMPILER": "icx", "CMAKE_CXX_COMPILER": "icpx", "CMAKE_INSTALL_RPATH": "$ORIGIN;$ORIGIN/.." } { "name": "x64-linux-gcc", "hidden": true, "cacheVariables": { "CMAKE_C_COMPILER": "icx", "CMAKE_CXX_COMPILER": "icpx" } { "name": "no_sharedlibs", "hidden": true, "cacheVariables": { "BUILD_SHARED_LIBS": "OFF" } }, { "name": "x64-linux-sycl-static-release", "inherits": [ "sycl-base", "release", "sycl_f16", "no_sharedlibs"] }, { "name": "x64-linux-vulkan-static-release", "inherits": [ "base", "vulkan", "release", "no_sharedlibs"] },icx --version
Intel(R) oneAPI DPC++/C++ Compiler 2026.0.0 (2026.0.0.20260331)intel-opencl-icd
26.18.38308.1-126.04ppa1Hope this helps
- Do you use
@agsharathnaik
Sorry, we aren't familiar with Vulkan backend building issue.Could you report your case in another issue for Vulkan?
Thank you!
@arthw —
rm -r buildfixed the compilation, thank you! The binary links successfully now.But there's a second issue: the compiled binary won't RUN. It crashes immediately with:
libsycl.so.9: undefined symbol: urDeviceWaitExp, version LIBUR_LOADER_0.12This is because 2026.0's
libsycl.so.9needsurDeviceWaitExpfrom the UR loader, but the 26.18.38308.1 GPU driver doesn't provide it yet. We also need to add 2025.3's lib path to LD_LIBRARY_PATH because DNNL 2025.3 still links againstlibsycl.so.8.So the status is:
- Compile: fixed with
rm -r build✅ - Run: blocked by missing
urDeviceWaitExpin 26.18 driver ❌
When Intel ships a GPU driver update that supports the new UR loader function, we can rebuild and test SYCL graph support immediately.
Reacted by Neo Zhang Jianyu, Stoney49th, Matt Kucia, Thicc-fil-a, encodatamHirmer and unmotivatedgene- Compile: fixed with
@arthw —
rm -r buildfixed the compilation, thank you! The binary links successfully now.But there's a second issue: the compiled binary won't RUN. It crashes immediately with:
libsycl.so.9: undefined symbol: urDeviceWaitExp, version LIBUR_LOADER_0.12This is because 2026.0's
libsycl.so.9needsurDeviceWaitExpfrom the UR loader, but the 26.18.38308.1 GPU driver doesn't provide it yet. We also need to add 2025.3's lib path to LD_LIBRARY_PATH because DNNL 2025.3 still links againstlibsycl.so.8.So the status is:
- Compile: fixed with
rm -r build✅ - Run: blocked by missing
urDeviceWaitExpin 26.18 driver ❌
When Intel ships a GPU driver update that supports the new UR loader function, we can rebuild and test SYCL graph support immediately.
Try clearing ccache before build.
I have Kubuntu 26.04 ; intel 26.18.38308.1 ; and oneapi 2026.0. I do not see any build failure or runtime failure. I have a lunarlake 256v basically arc2 igpu. No changes, just straight up upgraded from 2025.3 to 2026.0. So issue is not the default driver or oneapi.
- Compile: fixed with
@agsharathnaik
Yes, there is no above issue in my test too.Thank you for your sharing!
Just 21845
@R-SITES @agsharathnaik Fix it by PR: #24070
PR: #21845 merge to main branch fixed the sycl MTP performance issue.
@agsharathnaik — confirming that #21845 did improve SYCL MTP, though the effect is model-architecture-dependent.
Tested on Intel Arc Pro B70 (Battlemage, PCI 8086:e223, 32GB), mainline b9519 (7fe2ae4) with GGML_SYCL_F16=1, oneAPI 2025.3:
Qwen3.6 27B dense Q4_K_S:
Config tok/s vs no-MTP no-MTP 25.23 baseline MTP n_max=2 34.52 +37% Qwen3.6 35B-A3B MoE Q4_K_M:
Config tok/s vs no-MTP no-MTP 82.17 baseline MTP n_max=2 46.93 -43% MTP n_max=3 51.23 -38% The MMVQ improvement from #21845 helps the dense MMVQ path significantly — dense MTP is now faster than no-MTP on SYCL for the first time. MoE models use
mul_mat_idwhich was not touched by this PR, so MTP overhead still dominates there.masonmilby noted MoE paths are planned as a follow-up PR once this foundation settles. That's the piece still needed to close the gap on 35B-A3B and other MoE architectures.
#21845 is merged.
We could use the latest version to try it.Thank you!
14 remaining items
@arthw @grukx -- confirmed working on single Arc B70 (oneAPI 2026.0, SYCL F16, build b9604 with PR #24578 applied).
Crash is gone. Qwen3.6-35B-A3B Q4_K_M + MTP n_max=2 loads clean, no assert.
Full results on single B70:
Model No MTP MTP n=2 Delta 27B Dense Q4_K_S 25.0 tok/s 36.8 tok/s +47% 35B MoE Q4_K_M 82.2 tok/s 63.1 tok/s -23% Dense models are a clear win on single GPU. MoE still trails on single GPU due to MUL_MAT_ID expert dispatch overhead, but grukx dual B70 results (+19% at 4k, +78% at 32k, +143% at 61k) confirm MoE MTP becomes a win with cross-device parallelism.
96-98 graph reuses on dense, 90% draft acceptance. Build is clean.
Thanks for the quick PR and the dual-GPU testing.
Reacted by Neo ZhangReacted by SleepinDevilThank you for your feedback! :)
Reacted by SleepinDevilgithub-actions commented
on Jul 29, 2026 on Jul 29, 2026 – with GitHub ActionsContributorMore actionsThis issue was closed because it has been inactive for 14 days since being marked as stale.
Same GPU as the original report (Arc Pro B70, 32656 MiB), but on a newer
build, and I get a different — worse — result: near-zero draft acceptance and
some prompts returning zero tokens. Posting in case it helps localise whether
this is a model-file issue or a regression after b9292.Environment
- GPU: Intel Arc Pro B70, BMG-G31, device
0x8086:0xe223, 32656 MiB - Backend: SYCL, built from source with
icpx, oneAPI 2026.0 - Build: commit
360e1349f0009c5ad99d21e3c4546b707addc68a(master,
ggml 0.18.1) — i.e. later than the b9292 in the original report, so it
should contain SYCL gated_delta_net K>1 #23174, [SYCL] improve MoE prefill throughput (+70% with Qwen3.6-35B) #23142, Move to backend sampling for MTP draft path #23287 and server : free draft/MTP resources on sleep to fix VRAM leak #23461 - CMake:
-DGGML_SYCL=ON -DGGML_SYCL_F16=ON -DGGML_SYCL_GRAPH=ON -DGGML_SYCL_DNN=ON -DGGML_DNNL=ON -DGGML_SYCL_TARGET=INTEL - Env:
ONEAPI_DEVICE_SELECTOR=level_zero:0,ZES_ENABLE_SYSMAN=1 - Host: Ubuntu 26.04, Ryzen 9 9950X,
libze-intel-gpu126.05.37020.3 - Model:
unsloth/Qwen3.6-35B-A3B-MTP-GGUFUD-Q4_K_XL (MTP fused, 23.3 GB) - Flags:
--spec-type draft-mtp --spec-draft-n-max 2 -np 1 -ngl 99 -c 32768
Baseline agreement
Without MTP I measure 73.71 t/s (greedy,
temperature=0,top_k=1,
max_tokens=256, median of 5, distinct prompts), which is within 0.6% of the
73.3 t/s in the original report. So the two setups agree closely on baseline.With
--spec-type draft-mtpDraft acceptance is ~2.7%, not ~100%:
draft acceptance = 0.02703 ( 13 accepted / 481 generated) statistics draft-mtp: #calls(b,g,a) = 3 246 246, #gen drafts = 246, #acc drafts = 12, #gen tokens = 491, #acc tokens = 13Throughput is erratic, 14-33 t/s. Some requests return an empty completion
— one token generated,content: '', noreasoning_content. Others generate
the full 481 tokens. Output is not garbled multilingual text; it is either
empty or degenerate.Initialisation looks healthy:
srv load_model: creating MTP draft context against the target model common_speculative_impl_draft_mtp: adding speculative implementation 'draft-mtp' common_speculative_impl_draft_mtp: - n_max=2, n_min=0, p_min=0.00, n_embd=2048 common_speculative_impl_draft_mtp: - gpu_layers=-1, cache_k=f16, cache_v=f16, ctx_tgt=yes, ctx_dft=yes, devices=[default] srv load_model: speculative decoding context initializedEarlier build, for comparison
On commit
ac76808(b9237, before #23174) the same model and flags produced
garbled multilingual output that failed chat-template parsing with HTTP 500,
acceptance 0.00198, 18.24 t/s. Control on that build: same GGUF with the MTP
flags omitted gave clean output at 71.54 t/s.So #23174 clearly changed the failure mode here (garbled -> empty/degenerate),
but acceptance is still near zero on this setup.One difference that may matter
The original report used
Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-Q4_K_M.gguf. The
"Native-MTP-Preserved" naming suggests the MTP heads were deliberately
protected during conversion. I am using unsloth'sUD-Q4_K_XL, a dynamic
quant where the MTP heads presumably went through the same quantization as
the rest of the model.If prediction heads are quantization-sensitive, that would explain 100%
acceptance in one case and 2.7% in the other on the same hardware and
essentially the same code. There is precedent: DFlash drafters for this family
are documented as dropping from ~43% to ~28% acceptance when quantized to Q4
because of their sliding-window-attention layers.Happy to test a specific GGUF or build if that would help distinguish
"quantized MTP heads" from "regression after b9292" — this is a stable setup
and the non-MTP path performs normally.- GPU: Intel Arc Pro B70, BMG-G31, device
Could you check this issue on CPU by
--ngl 0?Ran the
-ngl 0test as suggested. Result isolates it cleanly to the SYCL backend — the MTP heads in the GGUF are fine.Same everything except
-ngl: same file (unsloth/Qwen3.6-35B-A3B-MTP-GGUFUD-Q4_K_XL), same build (360e1349f0009c5ad99d21e3c4546b707addc68a, ggml 0.18.1), same flags (--spec-type draft-mtp --spec-draft-n-max 2 -np 1), same host.-ngl backend draft acceptance output 99 SYCL (Arc Pro B70) 0.027 (13/481) empty or degenerate 0 CPU (Ryzen 9 9950X) 0.861 (62/72), 0.845 (250/296) clean, coherent Two separate CPU requests, 100 and 400 max tokens,
temperature=0,top_k=1. Both produced normal English. CPU throughput was 16.2-16.4 t/s, which is unremarkable for a 23 GB MoE with ~3B active params — the point here is correctness, not speed.This retracts the hypothesis in my previous comment. I suggested the MTP heads might have been damaged by unsloth's dynamic quantization, since the original report used a
Native-MTP-Preservedconversion and I did not. That looks wrong: the same file reaches ~85% acceptance on CPU. Whatever is happening is in the SYCL MTP path, not the weights.Also worth noting the failure mode differs between builds on SYCL:
ac76808(b9237, pre-#23174): garbled multilingual text, chat-template parse failure, HTTP 500, acceptance 0.00198360e134(post-SYCL gated_delta_net K>1 #23174): no garbling — instead some requests return a single token with empty content, others generate normally but with ~2.7% acceptance
So #23174 changed the failure mode without resolving it on this setup.
Environment recap: Intel Arc Pro B70 (BMG-G31,
0x8086:0xe223, 32656 MiB), oneAPI 2026.0,libze-intel-gpu126.05.37020.3, Ubuntu 26.04,ONEAPI_DEVICE_SELECTOR=level_zero:0, built with-DGGML_SYCL=ON -DGGML_SYCL_F16=ON -DGGML_SYCL_GRAPH=ON -DGGML_SYCL_DNN=ON -DGGML_DNNL=ON. Non-MTP baseline on the same image is 73.71 t/s.Happy to run further tests — e.g. with
Ran the `-ngl 0` test as suggested. Result isolates it cleanly to the SYCL backend — the MTP heads in the GGUF are fine.GGML_SYCL_GRAPH=OFF,GGML_SYCL_F16=0, a differentn_max, or a specific patch — if any of those would help narrow where in the verification path it diverges.Same everything except
-ngl: same file
(unsloth/Qwen3.6-35B-A3B-MTP-GGUFUD-Q4_K_XL), same build
(360e1349f0009c5ad99d21e3c4546b707addc68a, ggml 0.18.1), same flags
(--spec-type draft-mtp --spec-draft-n-max 2 -np 1), same host.-nglbackend draft acceptance output 99 SYCL (Arc Pro B70) 0.027 (13/481) empty or degenerate 0 CPU (Ryzen 9 9950X) 0.861 (62/72), 0.845 (250/296) clean, coherent Two separate CPU requests, 100 and 400 max tokens,
temperature=0,
top_k=1. Both produced normal English. CPU throughput was 16.2-16.4 t/s,
which is unremarkable for a 23 GB MoE with ~3B active params — the point here
is correctness, not speed.This retracts the hypothesis in my previous comment. I suggested the MTP
heads might have been damaged by unsloth's dynamic quantization, since the
original report used aNative-MTP-Preservedconversion and I did not. That
looks wrong: the same file reaches ~85% acceptance on CPU. Whatever is
happening is in the SYCL MTP path, not the weights.Also worth noting the failure mode differs between builds on SYCL:
ac76808(b9237, pre-[#23174](SYCL gated_delta_net K>1 #23174)):
garbled multilingual text, chat-template parse failure, HTTP 500,
acceptance 0.00198360e134(post-SYCL gated_delta_net K>1 #23174): no garbling — instead some requests return a single
token with empty content, others generate normally but with ~2.7% acceptance
So #23174 changed the failure mode without resolving it on this setup.
Environment recap: Intel Arc Pro B70 (BMG-G31,
0x8086:0xe223, 32656 MiB),
oneAPI 2026.0,libze-intel-gpu126.05.37020.3, Ubuntu 26.04,
ONEAPI_DEVICE_SELECTOR=level_zero:0, built with-DGGML_SYCL=ON -DGGML_SYCL_F16=ON -DGGML_SYCL_GRAPH=ON -DGGML_SYCL_DNN=ON -DGGML_DNNL=ON.
Non-MTP baseline on the same image is 73.71 t/s.Happy to run further tests — e.g. with
GGML_SYCL_GRAPH=OFF,GGML_SYCL_F16=0,
a differentn_max, or a specific patch — if any of those would help narrow
where in the verification path it diverges.Follow-up to the
-ngl 0result: I swept the SYCL runtime knobs to see whether any of them recover MTP on Arc. None do. Posting the negative result so nobody else spends time on these.All runs: same GGUF (
unsloth/Qwen3.6-35B-A3B-MTP-GGUFUD-Q4_K_XL), same build (360e134, ggml 0.18.1),-ngl 99,--spec-type draft-mtp,n_parallel=1,temperature=0,top_k=1, prompt"Say hello in one sentence.", 80 max tokens, container restarted between each.config result baseline (n_max=2) garbled → HTTP 500 on chat-template parse GGML_SYCL_USE_ASYNC_MEM_OP=0 same GGML_SYCL_ENABLE_GRAPH=0 + GGML_SYCL_GRAPH=0 same GGML_SYCL_ENABLE_FUSION=0 same GGML_SYCL_ENABLE_DNN=0 + GGML_SYCL_DNNL=0 same GGML_SYCL_ENABLE_OPT=0 same GGML_SYCL_F16=0 same all of the above together same --spec-draft-n-max 1 garbled, but returns 200 instead of 500 Caveat: several of those
GGML_SYCL_ENABLE_*names came fromstringsonlibggml-sycl.so, so some may be build-time defines rather than runtimegetenvlookups. If any of them are runtime-readable and I have the spelling wrong, that row is meaningless — happy to redo with correct names.n_max=1output, for reference — same token soup, just less dense, which is why it survives template parsing:'мести 😀 ** personn Nues " 11 「FSIZE : arch 籲headhead banjلانheadheadhead 女郎Kepheadhead女郎女郎headhead头放松,head女郎,女郎女郎, 很累升级到的心灵 ...'So
n_maxis not a factor — 1 and 2 fail the same way.n_parallelchanges the presentation but not the failure. Withn_parallelleft on auto (resolved to 4,kv_unified=true) some requests returned a single token with empty content while others generated normally at ~2.7% acceptance. Forced ton_parallel=1it is consistently garbled at 0.000 acceptance. Worth knowing if others report differing symptoms — it may just be slot count.Summary of what is now ruled out on this hardware:
- the GGUF / MTP head quantization (CPU reaches 0.845-0.861 acceptance on the same file)
n_paralleln_max- every SYCL runtime env var I could identify
What is left is the SYCL MTP kernel path itself. Environment variables do not reach it.
Setup recap: Arc Pro B70 (BMG-G31,
0x8086:0xe223, 32656 MiB), oneAPI 2026.0,libze-intel-gpu126.05.37020.3, Ubuntu 26.04, built with-DGGML_SYCL=ON -DGGML_SYCL_F16=ON -DGGML_SYCL_GRAPH=ON -DGGML_SYCL_DNN=ON -DGGML_DNNL=ON -DGGML_SYCL_TARGET=INTEL. Non-MTP baseline 73.71 t/s, clean.Still happy to run a patched build or a debug branch — this is a stable, scripted setup and I can turn a test around quickly.
@raphaelkobi-sys
I want to reproduce this issue. But the LLM unsloth/Qwen3.6-35B-A3B-MTP-GGUF UD-Q4_K_XL is too big.
Is it possible to reproduce it with a smaller LLM, like <20GB?Could you share the whole test cmd with parameters?
Thank you!
Happy to help — here is the full invocation, and a smaller model suggestion
below.Exact command
I run this in a container, but the only things that matter are the
llama-serverarguments and the two SYCL env vars:ONEAPI_DEVICE_SELECTOR=level_zero:0 \ ZES_ENABLE_SYSMAN=1 \ llama-server \ -m /models/Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf \ -ngl 99 \ --ctx-size 8192 \ -fa auto \ --spec-type draft-mtp \ --spec-draft-n-max 2 \ --parallel 1 \ --host 0.0.0.0 --port 8080
--parallel 1matters for reproducibility. Left on auto it resolves to 4 with
kv_unified=trueon my box, and the symptom changes from "garbled text" to
"some requests return a single token with empty content". Same underlying
failure, different presentation — worth pinning so we are comparing like with
like.Request:
curl -s http://localhost:8080/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{"messages":[{"role":"user","content":"Say hello in one sentence."}], "temperature":0,"top_k":1,"max_tokens":80}'
temperature=0, top_k=1(greedy) is deliberate: speculative decoding is
lossless under greedy sampling, so any deviation from the non-speculative
output is a bug rather than sampling noise.Note the model emits a thinking block, so the text lands in
reasoning_content, notcontent. Check both — an emptycontentalone is
not the bug.Control run (should be clean, ~71-73 t/s): same command with
--spec-type
and--spec-draft-n-maxremoved.Build:
cmake -B build -G Ninja \ -DCMAKE_BUILD_TYPE=Release \ -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx \ -DGGML_SYCL=ON -DGGML_SYCL_TARGET=INTEL \ -DGGML_SYCL_F16=ON -DGGML_SYCL_GRAPH=ON -DGGML_SYCL_DNN=ON -DGGML_DNNL=ON \ -DGGML_SYCL_SUPPORT_LEVEL_ZERO=ON -DGGML_SYCL_HOST_MEM_FALLBACK=ON \ -DGGML_NATIVE=ON -DLLAMA_OPENSSL=ON cmake --build build --target llama-server -j$(nproc)oneAPI 2026.0,
libze-intel-gpu126.05.37020.3, Ubuntu 26.04, commit
360e1349f0009c5ad99d21e3c4546b707addc68a.Smaller model
ggml-org/Qwen3.6-27B-GGUFis the better target for you — the base model is
~19 GB at Q4_K_M, and MTP ships as a separate ~3 GB sidecar
(mtp-Qwen3.6-27B-Q8_0.gguformtp-Q4_0) rather than fused into the file.
It is also the model #23149 was originally filed against.I am testing that combination on my B70 now and will post the result here.
Two things I want to confirm before you spend the download: whether it
reproduces at all, and whether the sidecar loading path behaves the same as
the fused-heads path I have been testing. Worth waiting for that.If it does not reproduce on the 27B, that itself would be informative — it
would mean the fused-MTP conversion is implicated rather than the SYCL kernels
alone, which would partly walk back my earlier conclusion.@raphaelkobi-sys
I still didn't reproduce it on B60.
The output is correct.Here are the configure:
Code: tag: b10423 Driver: lspci -nnk | grep -i vga -A3 03:00.0 VGA compatible controller [0300]: Intel Corporation Device [8086:e211] Subsystem: Shenzhen Gunnir Technology Development Co., Ltd Device [1ef7:2542] Kernel driver in use: xe Kernel modules: xe dpkg -l | grep libze-intel-gpu1 ii libze-intel-gpu1 26.27.39122.11-0 amd64 Intel(R) Graphics Compute Runtime for oneAPI Level Zero. icpx --version Intel(R) oneAPI DPC++/C++ Compiler 2026.1.1 (2026.1.1.20260724)Here are my cmds:
rm -rf build ./examples/sycl/build.sh fp16 ZES_ENABLE_SYSMAN=1 ./build/bin/llama-server -m ../models/Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf -ngl 99 --ctx-size 8192 -fa auto --spec-type draft-mtp --spec-draft-n-max 2 --parallel 1 --host 0.0.0.0 --port 8080 -lv 4Could you refer to them?
Thanks for the B60 details — very useful. I matched your driver and compiler and still reproduce on B70, so the difference now looks like the silicon rather than the software stack.
What I changed to match you
yours (B60, works) mine before mine now (B70, still fails) libze-intel-gpu1 26.27.39122.11 26.05.37020.3 26.27.39122.14 icpx 2026.1.1 (20260724) 2026.0.0 2026.1.0 (20260617) build flags build.sh fp16 GRAPH/DNN/DNNL/NATIVE on stripped to match GPU B60 8086:e211 B70 8086:e223 B70 8086:e223 I rebuilt three times: (1) minimal cmake flags — only
GGML_SYCL=ON,GGML_SYCL_TARGET=INTEL,GGML_SYCL_F16=ON; (2) newer oneAPI base bringing icpx 2026.1.0; (3) the 26.27 compute runtime from the kobuk-team PPA.All three produce byte-identical corrupted output, including the same draft acceptance (0.08333, 6/72) and literally the same token sequence:
'мести 😀 ** personn Nues " 11 「FSIZE : arch 籲headhead banjلانheadheadhead 女郎Kepheadhead女郎女郎headhead头放松,head女郎,女郎女郎, 很累升级到的心灵 ...'Deterministic under greedy sampling, unchanged across compiler, driver and build-flag changes.
Control on the same upgraded image
Same container, same model,
--spec-typeremoved:'Here's a thinking process:\n\n1. **Analyze User Input:**\n - **Request:** "Say hello in one sentence."\n - **Constraints:** Must be exactly one sentence...'Clean. So the 26.27 / icpx 2026.1 stack is healthy on this card — only the MTP path is broken.
Everything ruled out so far
- GGUF / MTP head quantization — same file reaches 0.845-0.861 acceptance
with
-ngl 0(CPU) n_parallel— fails at both 1 and auto(4); only the symptom differs (garbled + HTTP 500 at 1, emptycontentat 4)n_max— 1 and 2 fail identically- SYCL runtime env vars — 9 configs, no change
(
GGML_SYCL_USE_ASYNC_MEM_OP=0,ENABLE_GRAPH=0,ENABLE_FUSION=0,ENABLE_DNN=0,ENABLE_OPT=0,GGML_SYCL_F16=0, and all combined) - build flags — minimal build identical to full build
- driver + compiler — now matching yours, identical failure
What is left is B60 (
e211) vs B70 (e223).One difference I have not closed
You are on tag b10423; I am on
360e1349f0009c5ad99d21e3c4546b707addc68a. Given the output is bit-identical across three of my builds I doubt the commit explains it, but I have not proven that. Happy to rebuild at b10423 if you think it is worth ruling out — just say so.Offer
If it is genuinely BMG-G31-specific, I am glad to be the test rig: this is a scripted, reproducible setup and I can turn around a patched build or a debug branch quickly. If there is a kernel you would like instrumented, or a dump of draft-vs-verify logits at the divergence point, tell me what to capture.
Full config: Arc Pro B70, 32656 MiB, Ubuntu 26.04, kernel driver
xe,ONEAPI_DEVICE_SELECTOR=level_zero:0,ZES_ENABLE_SYSMAN=1. Non-MTP baseline 73.71 t/s.- GGUF / MTP head quantization — same file reaches 0.845-0.861 acceptance
with
Correction to my previous comment. I concluded the remaining difference was B60 vs B70 silicon. That was premature — vLLM runs MTP correctly and fast on this same B70, so the hardware is not the blocker.
vLLM on the same card
intel/llm-scaler-vllm(v0.21.1.dev0+gad7125a43), Qwen3.8-27B INT4 (GPTQ),--speculative-config {"method":"mtp","num_speculative_tokens":3}, XPU backend, same Arc Pro B70.Paired measurement, identical weights and prompts, three distinct prompts, warmup discarded,
temperature=0, 256 max tokens:config t/s no speculation 28.23 / 28.38 / 28.10 MTP, num_speculative_tokens=3 47.58 / 46.76 1.67x from speculation alone, on the same quantization.
vLLM's own acceptance metrics:
SpecDecoding metrics: Mean acceptance length: 2.76, Accepted: 317, Drafted: 540, Per-position acceptance rate: 0.817, 0.550, 0.394, Avg Draft acceptance rate: 58.7% SpecDecoding metrics: Mean acceptance length: 2.73, Accepted: 171, Drafted: 297, Per-position acceptance rate: 0.778, 0.556, 0.394, Avg Draft acceptance rate: 57.6%Compare llama.cpp on the same GPU, same architecture (
qwen35), same speculation method: draft acceptance 0.00000, garbled or empty output.vLLM resolves the architecture as
Qwen3_5MTPand shares the target model's embedding and lm_head weights with the drafter.So the silicon is fine
BMG-G31 executes MTP correctly. The fault is in llama.cpp's MTP path on this hardware, not in the GPU.
Caveat on scope: vLLM's XPU backend does not use ggml-sycl — it goes through its own
gdn_attention_core_xpuand FMHA sycl-tla kernels via torch/IPEX. So this shows a runtime driving MTP correctly on B70, not that the SYCL programming model as such is fine. But it does remove "the hardware cannot do this" as an explanation.One detail that may be relevant: vLLM logs
FMHA sycl-tla kernels cannot be captured with XPU graphs, falling back to PIECEWISE graph mode on XPU platform.so it is not using full graph capture either. That weakens graph capture as a candidate explanation for the llama.cpp failure (I had swept
GGML_SYCL_ENABLE_GRAPH=0with no change, which points the same way).Recap of what is ruled out for llama.cpp on B70
- GGUF / MTP head quantization — CPU (
-ngl 0) reaches 0.845-0.861 acceptance on the same file n_parallel(1 and 4),n_max(1 and 2)- every SYCL runtime env var I could identify (9 configs)
- cmake build flags (minimal build byte-identical to full build)
- driver 26.05 -> 26.27, icpx 2026.0 -> 2026.1
- the hardware (this comment)
- also reproduces on Qwen3.8-27B GGUF (native fused MTP, arch
qwen35): empty content, acceptance 0.00000 — so it is not specific to one model or one quant maker
Still happy to run instrumented builds. Setup is scripted and I can turn a test around quickly.
- GGUF / MTP head quantization — CPU (
I will verify this issue on a B70.
I don't think it's the difference of B60 vs B70 silicon.
It should be the LLM and driver/running time/software issue.
We has handled many user issues, all hardware special issues are related to the driver/running time issue, instead of silicon.- vLLM vs llama.cpp SYCL
They have different targets:
- vLLM: server/cloud, best performance, commercial support.
- SYCL backend: client/server, flexible to support new or more LLMs with different quantization format, community support.
To get the best performance on Intel GPU, vLLM apply the new performance lib/tech to reach the performance goal.
llama.cpp use common/popular lib/software to support more hardware with more LLMs.Compare vLLM to llama.cpp, like compare vLLM to PyTorch.
User should choose the right solution for their cases/targets.- vLLM vs llama.cpp SYCL
SYCL MTP on Intel Arc: correct output but no speed gain over baseline
Build: b9292 (master, oneAPI 2025.3.3)
GPU: Intel Arc Pro B70 (32656 MiB)
Model: Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-Q4_K_M.gguf (MoE, 256 experts, 8 active)
Summary
After the recent merges (#23174 GDN K>1, #23142 MoE prefill, #23287 backend sampling), SYCL MTP produces 100% correct output with perfect draft acceptance. This is a major milestone — earlier builds produced garbled text.
However, MTP consistently underperforms the non-MTP baseline:
MTP is ~21% slower than generating without speculation, despite 100% draft accuracy. This is the opposite of the expected result (typically 1.5-2x speedup on CUDA).
What was fixed (since b9159)
What remains
The MTP head forward pass adds significant launch overhead on SYCL. Each MTP step requires:
On SYCL/Intel Arc, the per-kernel dispatch cost is 100-500µs vs ~5µs on CUDA. The 100% draft acceptance means the draft tokens are perfect — the bottleneck is purely kernel dispatch overhead, not compute.
Comparison with Vulkan
The same b9187 code built with
-DGGML_VULKAN=ONachieves 53 tok/s with MTP — similar throughput but Vulkan has lower dispatch overhead per kernel.Launch command for reproduction
Prior related issues
Note
This is not a regression — the output quality and correctness fixes are invaluable. This report is to track that MTP generation speed on SYCL+MoE architectures still needs Intel kernel-level optimization (likely kernel fusion or SYCL graph support for the MTP head pipeline).