Skip to content

vulkan : Intel FA kernel for prefill. - #29357

Open
sliu39 wants to merge 2 commits into
ggml-org:masterfrom
sliu39:intel/fa-prefill
Open

sliu39 wants to merge 2 commits into
ggml-org:masterfrom
sliu39:intel/fa-prefill

Conversation

@sliu39

@sliu39 sliu39 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Collaborated and co-developed with @virajwad

This PR is targeting to improve Intel platform performance of FA prefill, with below changes:

Additional information

  • FA Customized shader is currently enabled on XE1 ARL-H, XE2/XE3 platforms
  • For A770 (XE1 dGPU), the existing performance test data with coopmat enforce on showed inconsistent result, we will still keep it disabled as the shader require coopmat.

Prefill Perf data with FA on:

Command line:

llama-bench.exe -p 8192 -n 0 -r 3 -fa 1 --delay 10 -ngl 99 -m qwen3-8b-q4_k_m.gguf,Qwen3.8-27B-UD-Q4_K_M.gguf,Qwen3.6-35B-A3B-UD-Q4_K_M.gguf,gpt-oss-20b-Q4_K_M.gguf

ARL-H Windows:

Baseline performance

ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc(TM) 140T GPU (32GB) (Intel Corporation) | uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat

model size params backend ngl fa test t/s
qwen3 8B Q4_K - Medium 4.68 GiB 8.19 B Vulkan 99 1 pp8192 126.58 ± 0.48
qwen35 27B Q4_K - Medium 15.32 GiB 27.32 B Vulkan 99 1 pp8192 83.38 ± 0.22
qwen35moe 35B.A3B Q4_K - Medium 21.10 GiB 35.51 B Vulkan 99 1 pp8192 247.82 ± 0.65
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 1 pp8192 478.14 ± 0.17

PR performance

ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc(TM) 140T GPU (32GB) (Intel Corporation) | uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat

model size params backend ngl fa test t/s
qwen3 8B Q4_K - Medium 4.68 GiB 8.19 B Vulkan 99 1 pp8192 415.31 ± 1.04
qwen35 27B Q4_K - Medium 15.32 GiB 27.32 B Vulkan 99 1 pp8192 119.95 ± 0.07
qwen35moe 35B.A3B Q4_K - Medium 21.10 GiB 35.51 B Vulkan 99 1 pp8192 394.44 ± 0.89
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 1 pp8192 639.02 ± 0.70

PTL-H Windows:

Baseline performance

ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc(TM) B390 GPU (Intel Corporation) | uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat

model size params backend ngl fa test t/s
qwen3 8B Q4_K - Medium 4.68 GiB 8.19 B Vulkan 99 1 pp8192 234.92 ± 0.39
qwen35 27B Q4_K - Medium 15.32 GiB 27.32 B Vulkan 99 1 pp8192 194.93 ± 0.59
qwen35moe 35B.A3B Q4_K - Medium 21.10 GiB 35.51 B Vulkan 99 1 pp8192 632.12 ± 17.56
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 1 pp8192 804.70 ± 84.56

PR performance

ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc(TM) B390 GPU (Intel Corporation) | uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat

model size params backend ngl fa test t/s
qwen3 8B Q4_K - Medium 4.68 GiB 8.19 B Vulkan 99 1 pp8192 911.60 ± 14.56
qwen35 27B Q4_K - Medium 15.32 GiB 27.32 B Vulkan 99 1 pp8192 264.39 ± 0.38
qwen35moe 35B.A3B Q4_K - Medium 21.10 GiB 35.51 B Vulkan 99 1 pp8192 914.80 ± 67.39
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 1 pp8192 1357.34 ± 125.30

B70 Pro Windows:

Baseline performance
ggml_vulkan: 0 = Intel(R) Arc(TM) Pro B70 Graphics (Intel Corporation) | uma: 0 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat

model size params backend ngl fa test t/s
qwen3 8B Q4_K - Medium 4.68 GiB 8.19 B Vulkan 99 1 pp8192 1064.67 ± 4.37
qwen35 27B Q4_K - Medium 15.32 GiB 27.32 B Vulkan 99 1 pp8192 659.10 ± 0.11
qwen35moe 35B.A3B Q4_K - Medium 20.60 GiB 34.66 B Vulkan 99 1 pp8192 2088.50 ± 3.32
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 1 pp8192 2311.37 ± 2.56

PR performance
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc(TM) Pro B70 Graphics (Intel Corporation) | uma: 0 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat

model size params backend ngl fa test t/s
qwen3 8B Q4_K - Medium 4.68 GiB 8.19 B Vulkan 99 1 pp8192 2955.83 ± 1.87
qwen35 27B Q4_K - Medium 15.32 GiB 27.32 B Vulkan 99 1 pp8192 932.13 ± 0.63
qwen35moe 35B.A3B Q4_K - Medium 20.60 GiB 34.66 B Vulkan 99 1 pp8192 3005.11 ± 4.24
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 1 pp8192 3967.29 ± 6.77

B70 Pro Linux:

Baseline performance
ggml_vulkan: 0 = Intel(R) Graphics (BMG G31) (Intel open-source Mesa driver) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat

model size params backend ngl fa test t/s
qwen3 8B Q4_K - Medium 4.68 GiB 8.19 B Vulkan 99 1 pp8192 783.67 ± 0.20
qwen35 27B Q4_K - Medium 15.32 GiB 27.32 B Vulkan 99 1 pp8192 497.22 ± 1.17
qwen35moe 35B.A3B Q4_K - Medium 20.60 GiB 34.66 B Vulkan 99 1 pp8192 1060.55 ± 5.09
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 1 pp8192 1237.51 ± 2.16

PR performance
ggml_vulkan: 0 = Intel(R) Graphics (BMG G31) (Intel open-source Mesa driver) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat

model size params backend ngl fa test t/s
qwen3 8B Q4_K - Medium 4.68 GiB 8.19 B Vulkan 99 1 pp8192 1518.19 ± 1.05
qwen35 27B Q4_K - Medium 15.32 GiB 27.32 B Vulkan 99 1 pp8192 647.78 ± 2.44
qwen35moe 35B.A3B Q4_K - Medium 20.60 GiB 34.66 B Vulkan 99 1 pp8192 1311.62 ± 7.55
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 1 pp8192 1841.35 ± 3.78

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, used claude code, then lots of manual review/tweaking.

@github-actions github-actions Bot added Vulkan Issues specific to the Vulkan backend ggml changes relating to the ggml tensor library for machine learning labels Sep 24, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 24, 2026

Copy link
Copy Markdown

Hi @sliu39, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@sliu39

sliu39 commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

@0cc4m Please help to review this PR for adding Intel customized FA shader for LLM prefill performance optimization, thanks!

@SleepinDevil

Copy link
Copy Markdown

This works incredibly well.

Ran the same benchmark I've run on these PRs before: #24406 (comment) & #24408 (comment)

The long context PP is amazing, wow, nice work bringing this back! The PP numbers are only ever so slightly lower than the original Mega PR numbers (#24408 (comment)).

If this PR is merged into the project, then this is a great achievement for llama.cpp on Intel GPUs. The slightly reduction of PP in the long context is worth this being merged into the project and the slight increase in TG we are seeing from the other updates that have happened too. Very exciting stuff!

Environment:
GPU Tested On: Intel Arc Pro B70
OS: Windows 11

llama-bench  -p 8192  -n 128  -d 0,8192,65536  -r 3  -fa 1  --delay 10  -ngl 99  --device Vulkan1  -m ..\LLM-models\Qwen3.6-27B-MTP-Q6_K.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce RTX 3080 Ti (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2
ggml_vulkan: 1 = Intel(R) Arc(TM) Pro B70 Graphics (Intel Corporation) | uma: 0 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
| model                          |       size |     params | backend    | ngl |  fa | dev          |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ------------ | --------------: | -------------------: |
| qwen35 27B Q6_K                |  21.30 GiB |    27.32 B | Vulkan     |  99 |   1 | Vulkan1      |          pp8192 |        958.80 ± 1.78 |
| qwen35 27B Q6_K                |  21.30 GiB |    27.32 B | Vulkan     |  99 |   1 | Vulkan1      |           tg128 |         21.12 ± 0.04 |
| qwen35 27B Q6_K                |  21.30 GiB |    27.32 B | Vulkan     |  99 |   1 | Vulkan1      |  pp8192 @ d8192 |        858.46 ± 0.41 |
| qwen35 27B Q6_K                |  21.30 GiB |    27.32 B | Vulkan     |  99 |   1 | Vulkan1      |   tg128 @ d8192 |         20.49 ± 0.02 |
| qwen35 27B Q6_K                |  21.30 GiB |    27.32 B | Vulkan     |  99 |   1 | Vulkan1      | pp8192 @ d65536 |        514.35 ± 0.18 |
| qwen35 27B Q6_K                |  21.30 GiB |    27.32 B | Vulkan     |  99 |   1 | Vulkan1      |  tg128 @ d65536 |         17.09 ± 0.01 |

build: 3c03a35bc (11184)

Comment thread ggml/src/ggml-vulkan/vulkan-shaders/flash_attn_prefill.comp Outdated
@jeffbolznv

Copy link
Copy Markdown
Contributor

I read through the new shader and it doesnt look fundamentally very different from the existing coopmat1 shader. Can you describe what it's doing differently that affects performance?

@0cc4m

0cc4m commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Please rebase.

@sliu39

sliu39 commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor Author

I read through the new shader and it doesnt look fundamentally very different from the existing coopmat1 shader. Can you describe what it's doing differently that affects performance?

Overall the attention flow are similar, there are some difference in design and details:

  • CM1 FA kernel require m16n16k16 coopMat shape that not supported in driver yet
  • For native SIMD8 architecture like ARL-H, the matrix core is m8n8k16 native. In this case even if driver add support for m16n16k16 and handle it underflow, it is likely to have efficiency issues for extra mov instructions and GRF spill issues
  • The customized shader designed for adaptive m8n8k16+SIMD8 and m16n8k16+SIMD16 support
  • The customized shader have different load K to SLM design for efficient matrixA load, it's dedicated for Intel architecture SLM message saving
  • The customized shader have different load V to SLM design for efficient matrixA load, also dedicated for Intel architecture.
  • The customized shader was using asymmetric subgroup workload balance for MATP calculation and SLM V Load.
  • To collaborate with the matV load difference, the customized shader swapped the P/V order to use V as matrixA, with matP shape difference accordingly, and is using different GEMM PV and acc flow: ACC = coopMatMulAdd(V, P, ACC); out= fma(out, compO, ACC); combined the output rescaling (mul) and accumulation (add) with single fma().

@arbv

arbv commented Sep 25, 2026 •

Copy link
Copy Markdown

Is this PR for Xe hardware that had XMX/coopmat support? Or more specifically, is it going to be of any use for platforms with Xe-LP (e.g. Tiger Lake)?

@virajwad

Copy link
Copy Markdown
Contributor

Is this PR for Xe hardware that had XMX/coopmat support? Or more specifically, is it going yo be of any use for platforms with Xe-LP (e.g. Tiger Lake)?

Hi @arbv this PR requires coopmat support so only with XMX hardware

@sliu39 sliu39 reopened this Sep 26, 2026
@sliu39

sliu39 commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

Please rebase.

Rebase done, please continue to review

@jeffbolznv

Copy link
Copy Markdown
Contributor

Overall the attention flow are similar, there are some difference in design and details:

Thank you for the summary. My concerns are primarily that we will have a shader that only runs on a (subset of) Intel hardware, and that this will make it difficult to make updates as the ops evolve because the code will not be easily testable. One way to address this is to try to unify with the existing flash_attn_cm1 shader, another would be to try to make this new shader support other devices (even if it is less performant than the existing shader).

I think some of the issues you list can be handled with variants of the same shader, e.g. specializing matrix sizes and shared memory layout. Changing the work distribution and swapping A/B matrices is more difficult to keep in the same shader source, but we don't actually know that the current choices in the cm1 shader are actually better than what you've done here.

I don't want to block the change, but I do want to be able to maintain the shader variant going forward. If it goes in as a separate shader that is Intel-specific, I will probably try to have codex make it runnable on other devices.

@virajwad

virajwad commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Hi @jeffbolznv I understand your concern, and we appreciate your support in helping maintain.

If we extended the support of the shader to accommodate m16n16k16 + SIMD32 for NV / AMD functional support, would that help address it? But like you mentioned, maybe unlikely the perf will be better than current shader in the master, although no objections to enabling it for other vendors if it helps. We could keep it default disabled for other vendors, but still testable via some method (like forced env variable or GPU detection change). What do you think?

For Intel side, we may need to keep it enabled for subset of devices that support cooperative matrix extension with XMX hardware. Probably easier for devices without XMX to use a scalar path rather than follow a driver emulation.

@jeffbolznv

Copy link
Copy Markdown
Contributor

If we extended the support of the shader to accommodate m16n16k16 + SIMD32 for NV / AMD functional support, would that help address it?

Yes, that would be great.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants