Skip to content

SYCL MTP on Intel Arc: correct output but no speed gain over baseline #23533

Description

@R-SITES

SYCL MTP on Intel Arc: correct output but no speed gain over baseline

Build: b9292 (master, oneAPI 2025.3.3)
GPU: Intel Arc Pro B70 (32656 MiB)
Model: Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-Q4_K_M.gguf (MoE, 256 experts, 8 active)

Summary

After the recent merges (#23174 GDN K>1, #23142 MoE prefill, #23287 backend sampling), SYCL MTP produces 100% correct output with perfect draft acceptance. This is a major milestone — earlier builds produced garbled text.

However, MTP consistently underperforms the non-MTP baseline:

Config Tok/s Draft Accept Output
No MTP 73.3 — ✅ Clean
MTP (draft-mtp, n_max=2) 57.8 26/26 (100%) ✅ Clean
MTP (draft-mtp, n_max=3) ~48 ~80% ✅ Clean

MTP is ~21% slower than generating without speculation, despite 100% draft accuracy. This is the opposite of the expected result (typically 1.5-2x speedup on CUDA).

What was fixed (since b9159)

What remains

The MTP head forward pass adds significant launch overhead on SYCL. Each MTP step requires:

  • 1 main model decode (40 layers, ~8 experts each)
  • 1 MTP head decode (1 layer, same expert count)
  • 1 verification pass (batched main model decode)

On SYCL/Intel Arc, the per-kernel dispatch cost is 100-500µs vs ~5µs on CUDA. The 100% draft acceptance means the draft tokens are perfect — the bottleneck is purely kernel dispatch overhead, not compute.

Comparison with Vulkan

The same b9187 code built with -DGGML_VULKAN=ON achieves 53 tok/s with MTP — similar throughput but Vulkan has lower dispatch overhead per kernel.

Launch command for reproduction

source /opt/intel/oneapi/setvars.sh --force && export GGML_SYCL_F16=1 && \
  ~/llama.cpp/build-sycl-mtp-b9292/bin/llama-server \
  --jinja --chat-template-file ~/models/chat_template-v18.jinja \
  -m ~/models/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-Q4_K_M.gguf \
  --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 1 \
  -t 24 -ngl 99 -ub 512 -b 4096

Prior related issues

Note

This is not a regression — the output quality and correctness fixes are invaluable. This report is to track that MTP generation speed on SYCL+MoE architectures still needs Intel kernel-level optimization (likely kernel fusion or SYCL graph support for the MTP head pipeline).

Activity

  1. arthw commented on May 23, 2026

    @arthw
    Contributor

    @R-SITES
    Yes, it's known issue.
    It will take a little more time to check.

    We will check if SYCL graph can help it since oneAPI 2026.0 has support it officially.
    Welcome more suggestion!

    Thank you!

  2. R-SITES commented on May 23, 2026

    @R-SITES
    Author

    @arthw — thanks for confirming this is on your radar. Here is additional data from thorough testing on Intel Arc Pro B70 (Battlemage, PCI 8086:e223) that may help narrow the investigation:

    Test matrix (b9292, oneAPI 2025.3.3, Qwen3.6-35B-A3B MoE Q4_K_M):

    Config Tok/s Draft Accept Output
    No MTP 73.3 — ✅
    MTP n_max=2 57.8 26/26 (100%) ✅
    MTP n_max=3 ~48 ~80% ✅
    Vulkan MTP n_max=2 53 ~81% ✅

    What we tried for SYCL MTP speed (none moved the needle):

    • GGML_SYCL_MMV_Y=4 — register spill, dropped to 25 tok/s. MMV_Y=1 is optimal for Arc EU register file
    • GGML_SYCL_DISABLE_GRAPH=0 re-enabled graphs — MUL_MAT_ID blocking stream->wait() prevents full pipeline recording. The fused MMVQ path dispatches 8 experts in one kernel, but the MoE routing copy calls stream->wait() for CPU-side expert selection, which cannot be captured
    • Manual fusion attempt (sigmoid+mul→single kernel) — compiled correctly but did not change throughput; saving 1 kernel launch out of ~960 per token is negligible
    • oneAPI 2026.0 DPC++ compiler — ggml-sycl backend does not compile with the 2026.0 compiler. Additionally, the 2026.0 runtime requires urDeviceWaitExp from Level Zero v2 API not yet in the public driver (latest: 26.18.38308.1)
    • MoE prefill PR [SYCL] improve MoE prefill throughput (+70% with Qwen3.6-35B) #23142 (applied) — improved prefill, +12% on MTP generation
    • Multi-column MMVQ PR sycl : port multi-column MMVQ from CUDA backend (~45% speculative decoding speedup on Intel Arc) #21845 — broke output on Qwen3.6 MoE

    Root cause analysis:

    Each token requires ~960 expert matmul dispatches. With MTP, verification adds another batch. The kernel launch overhead (100-500µs vs ~5µs CUDA) dominates. 100% draft acceptance confirms this is dispatch overhead, not compute.

    What would help:

    1. Record the full MoE layer graph as a single SYCL graph replayable with zero dispatch overhead
    2. The MUL_MAT_ID blocking wait() needs to become event-based
    3. Alternatively, pre-compute expert mappings on GPU and remove CPU-in-the-loop routing

    Additional note on oneAPI 2026.0 / Ubuntu 26.04 compatibility:

    The oneAPI 2026.0 compiler and runtime are installed but cannot currently function on Ubuntu 26.04 because Intel GPU driver repo (https://apt.repos.intel.com/gpu/ubuntu) does not yet list the "resolute" codename (returns 404). Once the matching Level Zero v2 driver lands in either the Intel GPU repo or on the compute-runtime GitHub releases page, we can immediately rebuild with 2026.0 and test SYCL graph support.

    Happy to test any experimental branch — hardware is available for immediate testing.

  3. arthw commented on May 23, 2026

    @arthw
    Contributor

    @R-SITES
    Thank you for your check and sharing!

    Looks like SYCL Graph is still out of work in SYCL backend.

    I will check other issue.

    Ubuntu 26.04 is not recommended to normal user:

    • it can't be upgrade from old Ubuntu, need to install from scatch.
    • It's still unstable for unknown reason.

    Thank you!

  4. sharathanaik commented on May 28, 2026

    @sharathanaik

    @arthw — t
    Additional note on oneAPI 2026.0 / Ubuntu 26.04 compatibility:

    oneAPI 2026.0 DPC++ compiler — ggml-sycl backend does not compile with the 2026.0 compiler. Additionally, the 2026.0
    runtime requires urDeviceWaitExp from Level Zero v2 API not yet in the public driver (latest: 26.18.38308.1)

    The oneAPI 2026.0 compiler and runtime are installed but cannot currently function on Ubuntu 26.04 because Intel GPU driver repo (https://apt.repos.intel.com/gpu/ubuntu) does not yet list the "resolute" codename (returns 404). Once the matching Level Zero v2 driver lands in either the Intel GPU repo or on the compute-runtime GitHub releases page, we can immediately rebuild with 2026.0 and test SYCL graph support.
    g.

    I am on kubuntu 26.04, I am on the latest 26.18.38308.1 driver. and oneapi version 2026.0. I use the ppa:kobuk-team/intel-graphics repository to get the latest intel driver for Ubuntu. Are you using this ppa?

    I compile llama.cpp from source.. I have not had any issues until now on compile, except for I have to switch to icpx and icx compilers in the configuration to get the correct behavior even for vulkan build.

    It would be great if mtp performance gets fixed for sycl, for right now vulkan on linux is 40% slower on linux compared to windows. so any speed gain barely faster than sycl without mtp on linux.

  5. R-SITES commented on May 28, 2026

    @R-SITES
    Author

    @agsharathnaik thanks for the info! We're also using the kobuk-team/testing PPA with the 26.18.38308.1 driver. But when we try oneAPI 2026.0, we get:

    icpx: error: linker command failed
    symbol lookup error: libsycl.so.9: undefined symbol: urDeviceWaitExp
    

    Could you share your cmake command and any env setup you use to get 2026.0 compiling and running? Specifically:

    • Do you use source setvars.sh --force or manually set PATH/LD_LIBRARY_PATH?
    • What cmake flags do you pass for SYCL?
    • Does the compiled binary actually run without the urDeviceWaitExp error, or do you also see that at runtime?

    We'd love to replicate your working 2026.0 setup and test SYCL graph support.

  6. arthw commented on May 28, 2026

    @arthw
    Contributor

    @agsharathnaik
    Yes, I use ppa:kobuk-team/intel-graphics repository to install the GPU driver too.
    To verify some issues, I will install the new or old driver from the github release: https://github.com/intel/compute-runtime/releases.

    for error:

    icpx: error: linker command failed
    symbol lookup error: libsycl.so.9: undefined symbol: urDeviceWaitExp
    

    rm -r build, then build again.

    Yes, we will check the MTP performance issue.

  7. sharathanaik commented on May 28, 2026

    @sharathanaik

    @agsharathnaik thanks for the info! We're also using the kobuk-team/testing PPA with the 26.18.38308.1 driver. But when we try oneAPI 2026.0, we get:

    icpx: error: linker command failed
    symbol lookup error: libsycl.so.9: undefined symbol: urDeviceWaitExp
    

    Could you share your cmake command and any env setup you use to get 2026.0 compiling and running? Specifically:

    • Do you use source setvars.sh --force or manually set PATH/LD_LIBRARY_PATH?
    • What cmake flags do you pass for SYCL?
    • Does the compiled binary actually run without the urDeviceWaitExp error, or do you also see that at runtime?

    We'd love to replicate your working 2026.0 setup and test SYCL graph support.

    I use the below directly in my .profile file. Since I pretty much use this compiler for everything.

    # Intel one api environment setup
    source /opt/intel/oneapi/setvars.sh
    

    But I also always compile using static build(no shared library) to avoid library visibility issues. Since I use VScode directly to build this I just alter the CMakePresets.json file to use the preset. The reason for fully switching to intel compiler is I use --cpu-moe for larger models.. this has a problem with vulkan using the gcc compiled runtime never accepting the thread count correctly.

        {
            "name": "sycl-base",
            "hidden": true,
            "generator": "Ninja",
            "binaryDir": "${sourceDir}/build-${presetName}",
            "cacheVariables": {
                "CMAKE_EXPORT_COMPILE_COMMANDS": "ON",
                "CMAKE_CXX_COMPILER": "icpx",
                "CMAKE_C_COMPILER": "icx",
                "GGML_SYCL": "ON",
                "CMAKE_INSTALL_RPATH": "$ORIGIN;$ORIGIN/.."
            }
        {
            "name":  "base",
            "hidden": true,
            "generator":   "Ninja",
            "binaryDir":   "${sourceDir}/build-${presetName}",
            "cacheVariables": {
                "CMAKE_EXPORT_COMPILE_COMMANDS": "ON",
                "CMAKE_C_COMPILER": "icx",
                "CMAKE_CXX_COMPILER": "icpx",
                "CMAKE_INSTALL_RPATH": "$ORIGIN;$ORIGIN/.."
            }
    
       {
            "name": "x64-linux-gcc", "hidden": true,
            "cacheVariables": {
                "CMAKE_C_COMPILER": "icx",
                "CMAKE_CXX_COMPILER": "icpx"
            }
    { "name": "no_sharedlibs",   "hidden": true, "cacheVariables": {  "BUILD_SHARED_LIBS": "OFF" } },
    
     { "name": "x64-linux-sycl-static-release", "inherits": [ "sycl-base", "release",  "sycl_f16",  "no_sharedlibs"] },
     { "name": "x64-linux-vulkan-static-release", "inherits": [ "base", "vulkan", "release", "no_sharedlibs"] },
    
    

    icx --version
    Intel(R) oneAPI DPC++/C++ Compiler 2026.0.0 (2026.0.0.20260331)

    intel-opencl-icd
    26.18.38308.1-126.04ppa1

    Hope this helps

  8. NeoZhangJianyu commented on May 29, 2026

    @NeoZhangJianyu
    Contributor

    @agsharathnaik
    Sorry, we aren't familiar with Vulkan backend building issue.

    Could you report your case in another issue for Vulkan?

    Thank you!

  9. R-SITES commented on May 29, 2026

    @R-SITES
    Author

    @arthw — rm -r build fixed the compilation, thank you! The binary links successfully now.

    But there's a second issue: the compiled binary won't RUN. It crashes immediately with:

    libsycl.so.9: undefined symbol: urDeviceWaitExp, version LIBUR_LOADER_0.12
    

    This is because 2026.0's libsycl.so.9 needs urDeviceWaitExp from the UR loader, but the 26.18.38308.1 GPU driver doesn't provide it yet. We also need to add 2025.3's lib path to LD_LIBRARY_PATH because DNNL 2025.3 still links against libsycl.so.8.

    So the status is:

    • Compile: fixed with rm -r build ✅
    • Run: blocked by missing urDeviceWaitExp in 26.18 driver ❌

    When Intel ships a GPU driver update that supports the new UR loader function, we can rebuild and test SYCL graph support immediately.

  10. sharathanaik commented on Jun 3, 2026

    @sharathanaik

    @arthw — rm -r build fixed the compilation, thank you! The binary links successfully now.

    But there's a second issue: the compiled binary won't RUN. It crashes immediately with:

    libsycl.so.9: undefined symbol: urDeviceWaitExp, version LIBUR_LOADER_0.12
    

    This is because 2026.0's libsycl.so.9 needs urDeviceWaitExp from the UR loader, but the 26.18.38308.1 GPU driver doesn't provide it yet. We also need to add 2025.3's lib path to LD_LIBRARY_PATH because DNNL 2025.3 still links against libsycl.so.8.

    So the status is:

    • Compile: fixed with rm -r build ✅
    • Run: blocked by missing urDeviceWaitExp in 26.18 driver ❌

    When Intel ships a GPU driver update that supports the new UR loader function, we can rebuild and test SYCL graph support immediately.

    Try clearing ccache before build.

    I have Kubuntu 26.04 ; intel 26.18.38308.1 ; and oneapi 2026.0. I do not see any build failure or runtime failure. I have a lunarlake 256v basically arc2 igpu. No changes, just straight up upgraded from 2025.3 to 2026.0. So issue is not the default driver or oneapi.

  11. NeoZhangJianyu commented on Jun 3, 2026

    @NeoZhangJianyu
    Contributor

    @agsharathnaik
    Yes, there is no above issue in my test too.

    Thank you for your sharing!

  12. NeoZhangJianyu commented on Jun 3, 2026

    @NeoZhangJianyu
    Contributor

    @R-SITES @agsharathnaik
    Fix it by PR: #24070

  13. sharathanaik commented on Jun 5, 2026

    @sharathanaik

    Just 21845

    @R-SITES @agsharathnaik Fix it by PR: #24070

    PR: #21845 merge to main branch fixed the sycl MTP performance issue.

  14. R-SITES commented on Jun 5, 2026

    @R-SITES
    Author

    @agsharathnaik — confirming that #21845 did improve SYCL MTP, though the effect is model-architecture-dependent.

    Tested on Intel Arc Pro B70 (Battlemage, PCI 8086:e223, 32GB), mainline b9519 (7fe2ae4) with GGML_SYCL_F16=1, oneAPI 2025.3:

    Qwen3.6 27B dense Q4_K_S:

    Config tok/s vs no-MTP
    no-MTP 25.23 baseline
    MTP n_max=2 34.52 +37%

    Qwen3.6 35B-A3B MoE Q4_K_M:

    Config tok/s vs no-MTP
    no-MTP 82.17 baseline
    MTP n_max=2 46.93 -43%
    MTP n_max=3 51.23 -38%

    The MMVQ improvement from #21845 helps the dense MMVQ path significantly — dense MTP is now faster than no-MTP on SYCL for the first time. MoE models use mul_mat_id which was not touched by this PR, so MTP overhead still dominates there.

    masonmilby noted MoE paths are planned as a follow-up PR once this foundation settles. That's the piece still needed to close the gap on 35B-A3B and other MoE architectures.

  15. NeoZhangJianyu commented on Jun 5, 2026

    @NeoZhangJianyu
    Contributor

    #21845 is merged.
    We could use the latest version to try it.

    Thank you!

  16. 14 remaining items

  17. R-SITES commented on Jun 13, 2026

    @R-SITES
    Author

    @arthw @grukx -- confirmed working on single Arc B70 (oneAPI 2026.0, SYCL F16, build b9604 with PR #24578 applied).

    Crash is gone. Qwen3.6-35B-A3B Q4_K_M + MTP n_max=2 loads clean, no assert.

    Full results on single B70:

    Model No MTP MTP n=2 Delta
    27B Dense Q4_K_S 25.0 tok/s 36.8 tok/s +47%
    35B MoE Q4_K_M 82.2 tok/s 63.1 tok/s -23%

    Dense models are a clear win on single GPU. MoE still trails on single GPU due to MUL_MAT_ID expert dispatch overhead, but grukx dual B70 results (+19% at 4k, +78% at 32k, +143% at 61k) confirm MoE MTP becomes a win with cross-device parallelism.

    96-98 graph reuses on dense, 90% draft acceptance. Build is clean.

    Thanks for the quick PR and the dual-GPU testing.

  18. arthw commented on Jun 14, 2026

    @arthw
    Contributor

    Thank you for your feedback! :)

  19. github-actions commented on Jul 29, 2026

    @github-actions
    Contributor

    This issue was closed because it has been inactive for 14 days since being marked as stale.

  20. raphaelkobi-sys commented on Aug 13, 2026

    @raphaelkobi-sys

    Same GPU as the original report (Arc Pro B70, 32656 MiB), but on a newer
    build, and I get a different — worse — result: near-zero draft acceptance and
    some prompts returning zero tokens. Posting in case it helps localise whether
    this is a model-file issue or a regression after b9292.

    Environment

    Baseline agreement

    Without MTP I measure 73.71 t/s (greedy, temperature=0, top_k=1,
    max_tokens=256, median of 5, distinct prompts), which is within 0.6% of the
    73.3 t/s in the original report. So the two setups agree closely on baseline.

    With --spec-type draft-mtp

    Draft acceptance is ~2.7%, not ~100%:

    draft acceptance = 0.02703 (   13 accepted /   481 generated)
    statistics draft-mtp: #calls(b,g,a) = 3 246 246, #gen drafts = 246,
      #acc drafts = 12, #gen tokens = 491, #acc tokens = 13
    

    Throughput is erratic, 14-33 t/s. Some requests return an empty completion
    — one token generated, content: '', no reasoning_content. Others generate
    the full 481 tokens. Output is not garbled multilingual text; it is either
    empty or degenerate.

    Initialisation looks healthy:

    srv load_model: creating MTP draft context against the target model
    common_speculative_impl_draft_mtp: adding speculative implementation 'draft-mtp'
    common_speculative_impl_draft_mtp: - n_max=2, n_min=0, p_min=0.00, n_embd=2048
    common_speculative_impl_draft_mtp: - gpu_layers=-1, cache_k=f16, cache_v=f16,
      ctx_tgt=yes, ctx_dft=yes, devices=[default]
    srv load_model: speculative decoding context initialized
    

    Earlier build, for comparison

    On commit ac76808 (b9237, before #23174) the same model and flags produced
    garbled multilingual output that failed chat-template parsing with HTTP 500,
    acceptance 0.00198, 18.24 t/s. Control on that build: same GGUF with the MTP
    flags omitted gave clean output at 71.54 t/s.

    So #23174 clearly changed the failure mode here (garbled -> empty/degenerate),
    but acceptance is still near zero on this setup.

    One difference that may matter

    The original report used
    Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-Q4_K_M.gguf. The
    "Native-MTP-Preserved" naming suggests the MTP heads were deliberately
    protected during conversion. I am using unsloth's UD-Q4_K_XL, a dynamic
    quant where the MTP heads presumably went through the same quantization as
    the rest of the model.

    If prediction heads are quantization-sensitive, that would explain 100%
    acceptance in one case and 2.7% in the other on the same hardware and
    essentially the same code. There is precedent: DFlash drafters for this family
    are documented as dropping from ~43% to ~28% acceptance when quantized to Q4
    because of their sliding-window-attention layers.

    Happy to test a specific GGUF or build if that would help distinguish
    "quantized MTP heads" from "regression after b9292" — this is a stable setup
    and the non-MTP path performs normally.

  21. NeoZhangJianyu commented on Aug 14, 2026

    @NeoZhangJianyu
    Contributor

    @raphaelkobi-sys

    Could you check this issue on CPU by --ngl 0?

  22. raphaelkobi-sys commented on Aug 14, 2026

    @raphaelkobi-sys

    Ran the -ngl 0 test as suggested. Result isolates it cleanly to the SYCL backend — the MTP heads in the GGUF are fine.

    Same everything except -ngl: same file (unsloth/Qwen3.6-35B-A3B-MTP-GGUF UD-Q4_K_XL), same build (360e1349f0009c5ad99d21e3c4546b707addc68a, ggml 0.18.1), same flags (--spec-type draft-mtp --spec-draft-n-max 2 -np 1), same host.

    -ngl backend draft acceptance output
    99 SYCL (Arc Pro B70) 0.027 (13/481) empty or degenerate
    0 CPU (Ryzen 9 9950X) 0.861 (62/72), 0.845 (250/296) clean, coherent

    Two separate CPU requests, 100 and 400 max tokens, temperature=0, top_k=1. Both produced normal English. CPU throughput was 16.2-16.4 t/s, which is unremarkable for a 23 GB MoE with ~3B active params — the point here is correctness, not speed.

    This retracts the hypothesis in my previous comment. I suggested the MTP heads might have been damaged by unsloth's dynamic quantization, since the original report used a Native-MTP-Preserved conversion and I did not. That looks wrong: the same file reaches ~85% acceptance on CPU. Whatever is happening is in the SYCL MTP path, not the weights.

    Also worth noting the failure mode differs between builds on SYCL:

    • ac76808 (b9237, pre-#23174): garbled multilingual text, chat-template parse failure, HTTP 500, acceptance 0.00198
    • 360e134 (post-SYCL gated_delta_net K>1 #23174): no garbling — instead some requests return a single token with empty content, others generate normally but with ~2.7% acceptance

    So #23174 changed the failure mode without resolving it on this setup.

    Environment recap: Intel Arc Pro B70 (BMG-G31, 0x8086:0xe223, 32656 MiB), oneAPI 2026.0, libze-intel-gpu1 26.05.37020.3, Ubuntu 26.04, ONEAPI_DEVICE_SELECTOR=level_zero:0, built with -DGGML_SYCL=ON -DGGML_SYCL_F16=ON -DGGML_SYCL_GRAPH=ON -DGGML_SYCL_DNN=ON -DGGML_DNNL=ON. Non-MTP baseline on the same image is 73.71 t/s.

    Happy to run further tests — e.g. with GGML_SYCL_GRAPH=OFF, GGML_SYCL_F16=0, a different n_max, or a specific patch — if any of those would help narrow where in the verification path it diverges.

    Ran the `-ngl 0` test as suggested. Result isolates it cleanly to the SYCL backend — the MTP heads in the GGUF are fine.

    Same everything except -ngl: same file
    (unsloth/Qwen3.6-35B-A3B-MTP-GGUF UD-Q4_K_XL), same build
    (360e1349f0009c5ad99d21e3c4546b707addc68a, ggml 0.18.1), same flags
    (--spec-type draft-mtp --spec-draft-n-max 2 -np 1), same host.

    -ngl backend draft acceptance output
    99 SYCL (Arc Pro B70) 0.027 (13/481) empty or degenerate
    0 CPU (Ryzen 9 9950X) 0.861 (62/72), 0.845 (250/296) clean, coherent

    Two separate CPU requests, 100 and 400 max tokens, temperature=0,
    top_k=1. Both produced normal English. CPU throughput was 16.2-16.4 t/s,
    which is unremarkable for a 23 GB MoE with ~3B active params — the point here
    is correctness, not speed.

    This retracts the hypothesis in my previous comment. I suggested the MTP
    heads might have been damaged by unsloth's dynamic quantization, since the
    original report used a Native-MTP-Preserved conversion and I did not. That
    looks wrong: the same file reaches ~85% acceptance on CPU. Whatever is
    happening is in the SYCL MTP path, not the weights.

    Also worth noting the failure mode differs between builds on SYCL:

    So #23174 changed the failure mode without resolving it on this setup.

    Environment recap: Intel Arc Pro B70 (BMG-G31, 0x8086:0xe223, 32656 MiB),
    oneAPI 2026.0, libze-intel-gpu1 26.05.37020.3, Ubuntu 26.04,
    ONEAPI_DEVICE_SELECTOR=level_zero:0, built with -DGGML_SYCL=ON -DGGML_SYCL_F16=ON -DGGML_SYCL_GRAPH=ON -DGGML_SYCL_DNN=ON -DGGML_DNNL=ON.
    Non-MTP baseline on the same image is 73.71 t/s.

    Happy to run further tests — e.g. with GGML_SYCL_GRAPH=OFF, GGML_SYCL_F16=0,
    a different n_max, or a specific patch — if any of those would help narrow
    where in the verification path it diverges.

  23. raphaelkobi-sys commented on Aug 14, 2026

    @raphaelkobi-sys

    Follow-up to the -ngl 0 result: I swept the SYCL runtime knobs to see whether any of them recover MTP on Arc. None do. Posting the negative result so nobody else spends time on these.

    All runs: same GGUF (unsloth/Qwen3.6-35B-A3B-MTP-GGUF UD-Q4_K_XL), same build (360e134, ggml 0.18.1), -ngl 99, --spec-type draft-mtp, n_parallel=1, temperature=0, top_k=1, prompt "Say hello in one sentence.", 80 max tokens, container restarted between each.

    config result
    baseline (n_max=2) garbled → HTTP 500 on chat-template parse
    GGML_SYCL_USE_ASYNC_MEM_OP=0 same
    GGML_SYCL_ENABLE_GRAPH=0 + GGML_SYCL_GRAPH=0 same
    GGML_SYCL_ENABLE_FUSION=0 same
    GGML_SYCL_ENABLE_DNN=0 + GGML_SYCL_DNNL=0 same
    GGML_SYCL_ENABLE_OPT=0 same
    GGML_SYCL_F16=0 same
    all of the above together same
    --spec-draft-n-max 1 garbled, but returns 200 instead of 500

    Caveat: several of those GGML_SYCL_ENABLE_* names came from strings on libggml-sycl.so, so some may be build-time defines rather than runtime getenv lookups. If any of them are runtime-readable and I have the spelling wrong, that row is meaningless — happy to redo with correct names.

    n_max=1 output, for reference — same token soup, just less dense, which is why it survives template parsing:

    'мести 😀 ** personn Nues "   11 「FSIZE : arch 籲headhead banjلانheadheadhead
    女郎Kepheadhead女郎女郎headhead头放松,head女郎,女郎女郎,  很累升级到的心灵 ...'
    

    So n_max is not a factor — 1 and 2 fail the same way.

    n_parallel changes the presentation but not the failure. With n_parallel left on auto (resolved to 4, kv_unified=true) some requests returned a single token with empty content while others generated normally at ~2.7% acceptance. Forced to n_parallel=1 it is consistently garbled at 0.000 acceptance. Worth knowing if others report differing symptoms — it may just be slot count.

    Summary of what is now ruled out on this hardware:

    • the GGUF / MTP head quantization (CPU reaches 0.845-0.861 acceptance on the same file)
    • n_parallel
    • n_max
    • every SYCL runtime env var I could identify

    What is left is the SYCL MTP kernel path itself. Environment variables do not reach it.

    Setup recap: Arc Pro B70 (BMG-G31, 0x8086:0xe223, 32656 MiB), oneAPI 2026.0, libze-intel-gpu1 26.05.37020.3, Ubuntu 26.04, built with -DGGML_SYCL=ON -DGGML_SYCL_F16=ON -DGGML_SYCL_GRAPH=ON -DGGML_SYCL_DNN=ON -DGGML_DNNL=ON -DGGML_SYCL_TARGET=INTEL. Non-MTP baseline 73.71 t/s, clean.

    Still happy to run a patched build or a debug branch — this is a stable, scripted setup and I can turn a test around quickly.

  24. arthw commented on Aug 17, 2026

    @arthw
    Contributor

    @raphaelkobi-sys
    I want to reproduce this issue. But the LLM unsloth/Qwen3.6-35B-A3B-MTP-GGUF UD-Q4_K_XL is too big.
    Is it possible to reproduce it with a smaller LLM, like <20GB?

    Could you share the whole test cmd with parameters?

    Thank you!

  25. raphaelkobi-sys commented on Aug 17, 2026

    @raphaelkobi-sys

    Happy to help — here is the full invocation, and a smaller model suggestion
    below.

    Exact command

    I run this in a container, but the only things that matter are the
    llama-server arguments and the two SYCL env vars:

    ONEAPI_DEVICE_SELECTOR=level_zero:0 \
    ZES_ENABLE_SYSMAN=1 \
    llama-server \
      -m /models/Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf \
      -ngl 99 \
      --ctx-size 8192 \
      -fa auto \
      --spec-type draft-mtp \
      --spec-draft-n-max 2 \
      --parallel 1 \
      --host 0.0.0.0 --port 8080

    --parallel 1 matters for reproducibility. Left on auto it resolves to 4 with
    kv_unified=true on my box, and the symptom changes from "garbled text" to
    "some requests return a single token with empty content". Same underlying
    failure, different presentation — worth pinning so we are comparing like with
    like.

    Request:

    curl -s http://localhost:8080/v1/chat/completions \
      -H 'Content-Type: application/json' \
      -d '{"messages":[{"role":"user","content":"Say hello in one sentence."}],
           "temperature":0,"top_k":1,"max_tokens":80}'

    temperature=0, top_k=1 (greedy) is deliberate: speculative decoding is
    lossless under greedy sampling, so any deviation from the non-speculative
    output is a bug rather than sampling noise.

    Note the model emits a thinking block, so the text lands in
    reasoning_content, not content. Check both — an empty content alone is
    not the bug.

    Control run (should be clean, ~71-73 t/s): same command with --spec-type
    and --spec-draft-n-max removed.

    Build:

    cmake -B build -G Ninja \
      -DCMAKE_BUILD_TYPE=Release \
      -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx \
      -DGGML_SYCL=ON -DGGML_SYCL_TARGET=INTEL \
      -DGGML_SYCL_F16=ON -DGGML_SYCL_GRAPH=ON -DGGML_SYCL_DNN=ON -DGGML_DNNL=ON \
      -DGGML_SYCL_SUPPORT_LEVEL_ZERO=ON -DGGML_SYCL_HOST_MEM_FALLBACK=ON \
      -DGGML_NATIVE=ON -DLLAMA_OPENSSL=ON
    cmake --build build --target llama-server -j$(nproc)

    oneAPI 2026.0, libze-intel-gpu1 26.05.37020.3, Ubuntu 26.04, commit
    360e1349f0009c5ad99d21e3c4546b707addc68a.

    Smaller model

    ggml-org/Qwen3.6-27B-GGUF is the better target for you — the base model is
    ~19 GB at Q4_K_M, and MTP ships as a separate ~3 GB sidecar
    (mtp-Qwen3.6-27B-Q8_0.gguf or mtp-Q4_0) rather than fused into the file.
    It is also the model #23149 was originally filed against.

    I am testing that combination on my B70 now and will post the result here.
    Two things I want to confirm before you spend the download: whether it
    reproduces at all, and whether the sidecar loading path behaves the same as
    the fused-heads path I have been testing. Worth waiting for that.

    If it does not reproduce on the 27B, that itself would be informative — it
    would mean the fused-MTP conversion is implicated rather than the SYCL kernels
    alone, which would partly walk back my earlier conclusion.

  26. arthw commented on Aug 19, 2026

    @arthw
    Contributor

    @raphaelkobi-sys
    I still didn't reproduce it on B60.
    The output is correct.

    Here are the configure:

    Code: tag: b10423
    Driver:
     lspci -nnk | grep -i vga -A3
    03:00.0 VGA compatible controller [0300]: Intel Corporation Device [8086:e211]
    	Subsystem: Shenzhen Gunnir Technology Development Co., Ltd Device [1ef7:2542]
    	Kernel driver in use: xe
    	Kernel modules: xe
    
    dpkg -l | grep libze-intel-gpu1
    ii  libze-intel-gpu1                                 26.27.39122.11-0                          amd64        Intel(R) Graphics Compute Runtime for oneAPI Level Zero.
    
    icpx --version
    Intel(R) oneAPI DPC++/C++ Compiler 2026.1.1 (2026.1.1.20260724)
    
    

    Here are my cmds:

    rm -rf build
    ./examples/sycl/build.sh fp16
    ZES_ENABLE_SYSMAN=1 ./build/bin/llama-server   -m ../models/Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf   -ngl 99   --ctx-size 8192   -fa auto   --spec-type draft-mtp   --spec-draft-n-max 2   --parallel 1   --host 0.0.0.0 --port 8080 -lv 4
    

    Could you refer to them?

  27. raphaelkobi-sys commented on Aug 19, 2026

    @raphaelkobi-sys

    Thanks for the B60 details — very useful. I matched your driver and compiler and still reproduce on B70, so the difference now looks like the silicon rather than the software stack.

    What I changed to match you

      yours (B60, works) mine before mine now (B70, still fails)
    libze-intel-gpu1 26.27.39122.11 26.05.37020.3 26.27.39122.14
    icpx 2026.1.1 (20260724) 2026.0.0 2026.1.0 (20260617)
    build flags build.sh fp16 GRAPH/DNN/DNNL/NATIVE on stripped to match
    GPU B60 8086:e211 B70 8086:e223 B70 8086:e223

    I rebuilt three times: (1) minimal cmake flags — only GGML_SYCL=ON, GGML_SYCL_TARGET=INTEL, GGML_SYCL_F16=ON; (2) newer oneAPI base bringing icpx 2026.1.0; (3) the 26.27 compute runtime from the kobuk-team PPA.

    All three produce byte-identical corrupted output, including the same draft acceptance (0.08333, 6/72) and literally the same token sequence:

    'мести 😀 ** personn Nues "   11 「FSIZE : arch 籲headhead banjلانheadheadhead
    女郎Kepheadhead女郎女郎headhead头放松,head女郎,女郎女郎,  很累升级到的心灵 ...'
    

    Deterministic under greedy sampling, unchanged across compiler, driver and build-flag changes.

    Control on the same upgraded image

    Same container, same model, --spec-type removed:

    'Here's a thinking process:\n\n1.  **Analyze User Input:**\n   - **Request:**
    "Say hello in one sentence."\n   - **Constraints:** Must be exactly one
    sentence...'
    

    Clean. So the 26.27 / icpx 2026.1 stack is healthy on this card — only the MTP path is broken.

    Everything ruled out so far

    • GGUF / MTP head quantization — same file reaches 0.845-0.861 acceptance with -ngl 0 (CPU)
    • n_parallel — fails at both 1 and auto(4); only the symptom differs (garbled + HTTP 500 at 1, empty content at 4)
    • n_max — 1 and 2 fail identically
    • SYCL runtime env vars — 9 configs, no change (GGML_SYCL_USE_ASYNC_MEM_OP=0, ENABLE_GRAPH=0, ENABLE_FUSION=0, ENABLE_DNN=0, ENABLE_OPT=0, GGML_SYCL_F16=0, and all combined)
    • build flags — minimal build identical to full build
    • driver + compiler — now matching yours, identical failure

    What is left is B60 (e211) vs B70 (e223).

    One difference I have not closed

    You are on tag b10423; I am on 360e1349f0009c5ad99d21e3c4546b707addc68a. Given the output is bit-identical across three of my builds I doubt the commit explains it, but I have not proven that. Happy to rebuild at b10423 if you think it is worth ruling out — just say so.

    Offer

    If it is genuinely BMG-G31-specific, I am glad to be the test rig: this is a scripted, reproducible setup and I can turn around a patched build or a debug branch quickly. If there is a kernel you would like instrumented, or a dump of draft-vs-verify logits at the divergence point, tell me what to capture.

    Full config: Arc Pro B70, 32656 MiB, Ubuntu 26.04, kernel driver xe, ONEAPI_DEVICE_SELECTOR=level_zero:0, ZES_ENABLE_SYSMAN=1. Non-MTP baseline 73.71 t/s.

  28. raphaelkobi-sys commented on Aug 19, 2026

    @raphaelkobi-sys

    Correction to my previous comment. I concluded the remaining difference was B60 vs B70 silicon. That was premature — vLLM runs MTP correctly and fast on this same B70, so the hardware is not the blocker.

    vLLM on the same card

    intel/llm-scaler-vllm (v0.21.1.dev0+gad7125a43), Qwen3.8-27B INT4 (GPTQ), --speculative-config {"method":"mtp","num_speculative_tokens":3}, XPU backend, same Arc Pro B70.

    Paired measurement, identical weights and prompts, three distinct prompts, warmup discarded, temperature=0, 256 max tokens:

    config t/s
    no speculation 28.23 / 28.38 / 28.10
    MTP, num_speculative_tokens=3 47.58 / 46.76

    1.67x from speculation alone, on the same quantization.

    vLLM's own acceptance metrics:

    SpecDecoding metrics: Mean acceptance length: 2.76, Accepted: 317, Drafted: 540,
      Per-position acceptance rate: 0.817, 0.550, 0.394, Avg Draft acceptance rate: 58.7%
    SpecDecoding metrics: Mean acceptance length: 2.73, Accepted: 171, Drafted: 297,
      Per-position acceptance rate: 0.778, 0.556, 0.394, Avg Draft acceptance rate: 57.6%
    

    Compare llama.cpp on the same GPU, same architecture (qwen35), same speculation method: draft acceptance 0.00000, garbled or empty output.

    vLLM resolves the architecture as Qwen3_5MTP and shares the target model's embedding and lm_head weights with the drafter.

    So the silicon is fine

    BMG-G31 executes MTP correctly. The fault is in llama.cpp's MTP path on this hardware, not in the GPU.

    Caveat on scope: vLLM's XPU backend does not use ggml-sycl — it goes through its own gdn_attention_core_xpu and FMHA sycl-tla kernels via torch/IPEX. So this shows a runtime driving MTP correctly on B70, not that the SYCL programming model as such is fine. But it does remove "the hardware cannot do this" as an explanation.

    One detail that may be relevant: vLLM logs

    FMHA sycl-tla kernels cannot be captured with XPU graphs,
    falling back to PIECEWISE graph mode on XPU platform.
    

    so it is not using full graph capture either. That weakens graph capture as a candidate explanation for the llama.cpp failure (I had swept GGML_SYCL_ENABLE_GRAPH=0 with no change, which points the same way).

    Recap of what is ruled out for llama.cpp on B70

    • GGUF / MTP head quantization — CPU (-ngl 0) reaches 0.845-0.861 acceptance on the same file
    • n_parallel (1 and 4), n_max (1 and 2)
    • every SYCL runtime env var I could identify (9 configs)
    • cmake build flags (minimal build byte-identical to full build)
    • driver 26.05 -> 26.27, icpx 2026.0 -> 2026.1
    • the hardware (this comment)
    • also reproduces on Qwen3.8-27B GGUF (native fused MTP, arch qwen35): empty content, acceptance 0.00000 — so it is not specific to one model or one quant maker

    Still happy to run instrumented builds. Setup is scripted and I can turn a test around quickly.

  29. arthw commented on Aug 20, 2026

    @arthw
    Contributor

    @raphaelkobi-sys

    I will verify this issue on a B70.
    I don't think it's the difference of B60 vs B70 silicon.
    It should be the LLM and driver/running time/software issue.
    We has handled many user issues, all hardware special issues are related to the driver/running time issue, instead of silicon.

    1. vLLM vs llama.cpp SYCL
      They have different targets:
    • vLLM: server/cloud, best performance, commercial support.
    • SYCL backend: client/server, flexible to support new or more LLMs with different quantization format, community support.

    To get the best performance on Intel GPU, vLLM apply the new performance lib/tech to reach the performance goal.
    llama.cpp use common/popular lib/software to support more hardware with more LLMs.

    Compare vLLM to llama.cpp, like compare vLLM to PyTorch.
    User should choose the right solution for their cases/targets.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions