Repository navigation
vulkan: add Intel Xe flash attention optimization kernels (2/3, Xe-LPG Plus/Xe2/Xe3) - #24406
Conversation
|
Hi @fish-jiang, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
This is way too much complexity and code for one vendor. |
|
Hi @0cc4m I understand your point, it is large (2300+ lines of code). We will circle back and re-look at the code, but any high level feedback you could please offer us how we could improve it? Intel perf w/ the current flash attention shaders is very unoptimized, especially in long context scenarios. We tried to write custom FA shaders that significantly help our platforms and users, across lots of models, and gives improvement in both prefill and decode stages (MegaPR has FA=ON perf graphs). Any feedback you could offer that could help this PR go in would be very beneficial! |
|
A few questions from my perspective:
|
|
Yes, basically what Jeff said. Vulkan is meant to be generic, code should be shared between vendors where possible. That way everyone benefits from optimizations and other improvements. Of course, that is not always possible, and device-specific tuning is still necessary. But 2000 lines of additional code specifically for one vendor is excessive and would make maintenance of the backend much harder. I know Intel's architecture is special and differs in significant ways from AMD and Nvidia, but I also know that many of the issues we are dealing with are down to unoptimized/immature drivers (Linux ANV especially), and that is the main thing that needs to improve, in my opinion. The other problem is that it has been pretty hard for me to figure out how to write shaders in a way that works well on Intel, especially in regards to subgroups and subgroup operators. I've made great progress, but it's still not up there. I do appreciate that you want to help out with that. I hope we can find a less intrusive way to improve the Vulkan performance for Intel GPUs. |
@jeffbolznv |
|
Does this fundamentally come down to Intel having fewer registers per shader core available? In my experience that has been a common problem with trying to choose FA tile size. Do I understand correctly that the split prefill is really not flash attention, it's just softmax(Q*K) spilled to memory and then the second multiply in a separate dispatch, and this whole thing is chunked to put a bound on the temporary buffer size? One of the bigger problems having a totally separate path like this is that it won't be easily testable on other devices, which makes it very difficult to maintain. |
Exactly, smaller GRF bytes per warp is one of the key diffs, and there are more others including smaller subgroup size and larger warp count per workgroup, SLM reshape for message reduce, efficient WG dispatch/memory access pattern.
Agree with the point, split prefill does not strictly comply with FA definition, but softmax operation is still fused as epilog/preprocessing in phase1/phase2. It’s a tradeoff in large head dim case to avoid quad reduce for Q^K or GRF spill with all-in-one FA kernel. For maintenance, other vendors with smaller subgroup/GRF size may also use it or the idea of it, there may be dedicate effort in case of major changes like Vulkan FA op definition upgrade, elsewise would be covered in test-backend-ops and E2E model test. |
|
But the most obvious first attempt to make flash attention run better on Intel Battlemage would be to start using coopmat. We currently require 16x16x16 coopmat shape support for Flash Attention, which IIRC Intel GPUs don't offer. They have 8x8x16? Have you tried adapting the existing coopmat FA shader to support this? |
Yes, we initially tried to add CM1 with existing FA kernel flow, but due to fundamental GPU arch diffs mentioned in above discussion, we were unable to get consistent prefill win over FA off path unless making significant changes in kernel flow. To avoid impacting existing path, those changes were put in separate path. |
|
We can investigate that direction if it's really the only way for Intel, but different shaders for decode, prefill and even specific head sizes would make maintenance too hard. Would it be possible to push it into 2 shaders, one for each stage? |
Thanks for the suggestion, tried experiments of prefill/decode unified 2-phases attention solution, the prefill performance seems to be OK, but decode performance dropped, and need an dedicated decode path to close the gap. |
|
I mean I want to keep the number of shaders to maintain minimal. Currently we have 3 for Flash Attention and that works reasonably well across most devices. The fewer the better. |
|
Hi @0cc4m We've experimented with a consolidated 2 shader solution, however we are seeing significant regressions. We wanted to share the data with you. We looked at the perf in two aspects on an Arrow Lake (ARL-H) Xe2 Intel platform, covering the same models as our mega-pr:
PrefillAspect 1 - We see good perf increases in FA=ON case for 8K input tokens, no FA=ON regressions.
Testing aspect 2, we are able to maintain most prefill gains from our optimizations. Some models gain perf and some lose compared to our mega pr, but average is 1.03x
DecodeAspect 1 - Significant perf loss across every model we tested, so 2 shader solution causes our decode to be much worse than current master.
Testing aspect 2, the decode perf on average is 0.38x compared to our mega pr optimizations in flash attention. Here are the ratios:
We have a hunch that if we increase to 3 shader solution, we may be able to solve the decode regression (decode will be split-phase), but we need time to experiment and measure once more. @0cc4m could we please get your early thoughts on increasing the shader count? We still want to respect the ask to consolidate this PR a lot, but we are having difficulties porting these optimizations on 2 shaders. |
37bc472 to
0ef6e55
Compare
|
Hi @0cc4m just checking on the above! Thank you |
|
In general, the fewer shaders the better. If you need 3 then you need 3, but someone will have to maintain them. Are they coopmat-only? My second concern is that I still have no access to any Battlemage hardware, so I cannot test anything that doesn't run on my A770. |
|
|
If you can help out with hardware that would be good, you can send me a mail with details. |
|
Just dropping by to leave an update, tried out the latest Vulkan and Sycl updates on the Intel B70, and this custom build that included the Intel dev's original 3x PRs is still so much faster PP that it is just not worth running anything else Sycl or Vulkan besides that custom build on my Intel B70 🤯. I am just so happy that the build runs Qwen3.8-27b perfectly without a hitch so far! 🤞 I still hope the performance available on the table with the Intel XMX cores via Vulkan can be merged into the project somehow after you guys figure out the most feasible way to do it for long-term maintainability too. Keep up the great work! |
|
@SleepinDevil and @0cc4m We have identified two active contributors to help maintaining Intel changes long term: - @rillomas and @virajwad. They should be able to support maintaining Llama.cpp/Vulkan changes for Intel. |
|
That's great! What's the status of the shader work, have you found a minimal number of required shaders yet? |
|
It was just your feature branch merged into the master branch. |
|
Will this work on iGPU found on Tiger Lake? I am currently using Vulkan backend on this oldish laptop and it does work better than pure CPU inference, especially at PP. |
|
Hi @arbv unfortunately not, this PR has a dependency on cooperative matrix extension which tiger lake doesn't support (no XMX hardware) |
|
@virajwad, oh, I see. Will not it break the existing support at least? |
|
@arbv Yes it should maintain existing support, this PR is an additional code path for devices with supported XMX hardware |
|
I did actually measure a perf improvement on A770, so we should look into extending support for Xe1 after this lands. |
|
@virajwad Cool. I don't think I am alone who could benefit from running an autonomous, small clanker slightly faster on portable hardware (if "faster" can be used in this sentence given the hardware limitations in compute an memory bandwidth). |
…hader to resolve test op failre on A770 Linux with 26.2.3 mesa driver
Issue root caused. With all experiments so far, most likely the latest Linux MESA driver have issues when dispatching WGs with asymmetric coopMatMulAdd() count into the same XeCore, it was good on 26.0.8 version that I tested yesterday, and that's why the issue was not able to be reproduced. After upgrading to 26.2.3 version the issue comes out instantly. Changed the kernel to use symmetric coopMatMulAdd() count to resolve the potential issue on A770, op tests now all pass with 26.2.3 MESA driver on A770, please let me know if the fix is working. |
|
Thank you, I can confirm it's resolved. I also still measure some tg improvements on A770 if enabled. |
|
Fix the editorconfig issue please. |
Done, thanks for the reminder |
|
Tested 2 models on A770 Linux, observed below TPS perf gain with enforcing the 2 phase FA shader path:
Before change:
Qwen2-14B
Before change:
|
|
Amazing work by all of you that contributed to this! I'm glad to see this merge go through :) |
Couldn't help but come back to this today 😅. @virajwad very eager to see what the Intel team puts together for the consolidated prefill shader to bring back some of that original performance on the prompt processing too (9591 build results in tables below) :) I grabbed the current master Vulkan build and compared the Qwen3.6-27B-Q6_K (Unsloth) model against the current master Vulkan build (11152), the original Mega PR build of Intel's (9591, #24408 (comment)), and the original official Vulkan build (9851) from back when the Mega PR was set up. The results below confirm that the TG collapse at long context has been significantly improved on the Intel B70. I appreciate this is comparing apples to oranges a bit here because there's been so many updates to Vulkan in the few months, but I think most of the TG improvement we see here is because of your team's hard work. I've been following the other Vulkan PRs too and nothing really addressed the long context TG dropoff like this PR did. Very impressed with what you guys have managed with the shaders, great work guys! Looking forward to the prefill shader! Summary table of the below results
Vulkan - Latest build from source 24/09/2026 - 11152
build: 013b31c (11152) Vulkan - Original Official Build from few months ago - 9851
build: 0eca4d4 (9851) - Official Vulkan Release for Windows Vulkan - Original PR 24406 build from when Mega PR was posted few months ago - 9591
build: 37bc472e6 (9591) - Vulkan Built with PR#24406 only |
|
Thanks @SleepinDevil for your kind words. We opened the prefill shader as draft PR, we're testing it more fully but will give update soon! #29357 |
…G Plus/Xe2/Xe3) (ggml-org#24406) * vulkan : Intel FA kernel optimization for split k path * vulkan : Host code update for Intel split k FA kernel path selection, fix A770 Linux op test failures * vulkan : use symmetric coopMatMulAdd() in flash_attn_decode_phase_1 shader to resolve test op failre on A770 Linux with 26.2.3 mesa driver * vulkan : fix editorconfig issue in flash_attn_decode_phase_2.comp --------- Co-authored-by: Liu, Russell <russell.liu@intel.com>
…G Plus/Xe2/Xe3) (ggml-org#24406) * vulkan : Intel FA kernel optimization for split k path * vulkan : Host code update for Intel split k FA kernel path selection, fix A770 Linux op test failures * vulkan : use symmetric coopMatMulAdd() in flash_attn_decode_phase_1 shader to resolve test op failre on A770 Linux with 26.2.3 mesa driver * vulkan : fix editorconfig issue in flash_attn_decode_phase_2.comp --------- Co-authored-by: Liu, Russell <russell.liu@intel.com>
…G Plus/Xe2/Xe3) (ggml-org#24406) * vulkan : Intel FA kernel optimization for split k path * vulkan : Host code update for Intel split k FA kernel path selection, fix A770 Linux op test failures * vulkan : use symmetric coopMatMulAdd() in flash_attn_decode_phase_1 shader to resolve test op failre on A770 Linux with 26.2.3 mesa driver * vulkan : fix editorconfig issue in flash_attn_decode_phase_2.comp --------- Co-authored-by: Liu, Russell <russell.liu@intel.com> (cherry picked from commit 4ceb171)




Overview
Co-authors: @jxia4intel, @sliu39
PR 2/3 of the Intel Xe optimization series — see #24408 (mega PR, draft) for the full feature set.
Target platforms: Xe-LPG Plus, Xe2, Xe3
This PR adds Intel Xe-specific flash attention optimization kernels for both ARLH iGPU (Xe1, UMA, coopmat1) and Xe2/Xe3. Dependency: builds on top of #24404 (Xe-LPG Plus coopmat1 enable). Independent of #24407 (GEMM+CW).
Flash Attention (Intel Xe)
flash_attn_hdim64/96/128) and two-phase split prefill/decode variants(head_dim, gqa_ratio)for runtime dispatch across various GQA ratios without combinatorial pipeline proliferationqk_groups)fa_copy_qstate) between prefill phasesPerformance (Panther Lake B390 + Windows OS)
Requirements