Repository navigation
Conversation
|
Hi @sliu39, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
@0cc4m Please help to review this PR for adding Intel customized FA shader for LLM prefill performance optimization, thanks! |
|
This works incredibly well. Ran the same benchmark I've run on these PRs before: #24406 (comment) & #24408 (comment) The long context PP is amazing, wow, nice work bringing this back! The PP numbers are only ever so slightly lower than the original Mega PR numbers (#24408 (comment)). If this PR is merged into the project, then this is a great achievement for llama.cpp on Intel GPUs. The slightly reduction of PP in the long context is worth this being merged into the project and the slight increase in TG we are seeing from the other updates that have happened too. Very exciting stuff! Environment: |
|
I read through the new shader and it doesnt look fundamentally very different from the existing coopmat1 shader. Can you describe what it's doing differently that affects performance? |
|
Please rebase. |
Overall the attention flow are similar, there are some difference in design and details:
|
|
Is this PR for Xe hardware that had XMX/coopmat support? Or more specifically, is it going to be of any use for platforms with Xe-LP (e.g. Tiger Lake)? |
Hi @arbv this PR requires coopmat support so only with XMX hardware |
e3c41cb to
4b1a27f
Compare
Rebase done, please continue to review |
Thank you for the summary. My concerns are primarily that we will have a shader that only runs on a (subset of) Intel hardware, and that this will make it difficult to make updates as the ops evolve because the code will not be easily testable. One way to address this is to try to unify with the existing flash_attn_cm1 shader, another would be to try to make this new shader support other devices (even if it is less performant than the existing shader). I think some of the issues you list can be handled with variants of the same shader, e.g. specializing matrix sizes and shared memory layout. Changing the work distribution and swapping A/B matrices is more difficult to keep in the same shader source, but we don't actually know that the current choices in the cm1 shader are actually better than what you've done here. I don't want to block the change, but I do want to be able to maintain the shader variant going forward. If it goes in as a separate shader that is Intel-specific, I will probably try to have codex make it runnable on other devices. |
|
Hi @jeffbolznv I understand your concern, and we appreciate your support in helping maintain. If we extended the support of the shader to accommodate m16n16k16 + SIMD32 for NV / AMD functional support, would that help address it? But like you mentioned, maybe unlikely the perf will be better than current shader in the master, although no objections to enabling it for other vendors if it helps. We could keep it default disabled for other vendors, but still testable via some method (like forced env variable or GPU detection change). What do you think? For Intel side, we may need to keep it enabled for subset of devices that support cooperative matrix extension with XMX hardware. Probably easier for devices without XMX to use a scalar path rather than follow a driver emulation. |
Yes, that would be great. |
Overview
Collaborated and co-developed with @virajwad
This PR is targeting to improve Intel platform performance of FA prefill, with below changes:
Additional information
Prefill Perf data with FA on:
Command line:
llama-bench.exe -p 8192 -n 0 -r 3 -fa 1 --delay 10 -ngl 99 -m qwen3-8b-q4_k_m.gguf,Qwen3.8-27B-UD-Q4_K_M.gguf,Qwen3.6-35B-A3B-UD-Q4_K_M.gguf,gpt-oss-20b-Q4_K_M.gguf
ARL-H Windows:
Baseline performance
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc(TM) 140T GPU (32GB) (Intel Corporation) | uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
PR performance
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc(TM) 140T GPU (32GB) (Intel Corporation) | uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
PTL-H Windows:
Baseline performance
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc(TM) B390 GPU (Intel Corporation) | uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
PR performance
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc(TM) B390 GPU (Intel Corporation) | uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
B70 Pro Windows:
Baseline performance
ggml_vulkan: 0 = Intel(R) Arc(TM) Pro B70 Graphics (Intel Corporation) | uma: 0 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
PR performance
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc(TM) Pro B70 Graphics (Intel Corporation) | uma: 0 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
B70 Pro Linux:
Baseline performance
ggml_vulkan: 0 = Intel(R) Graphics (BMG G31) (Intel open-source Mesa driver) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
PR performance
ggml_vulkan: 0 = Intel(R) Graphics (BMG G31) (Intel open-source Mesa driver) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
Requirements