Skip to content

ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 - #29423

Merged
ggerganov merged 3 commits into
ggml-org:masterfrom
SongXiaoXi:cpu_fa_tile
Sep 28, 2026
Merged

ggerganov merged 3 commits into
ggml-org:masterfrom
SongXiaoXi:cpu_fa_tile

Conversation

@SongXiaoXi

@SongXiaoXi SongXiaoXi commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Overview

The CPU tiled flash attention kernel is gated on the value head size being a multiple of the SIMD vector width. Head sizes like 72 fall back to the per-query kernel, which is an order of magnitude slower and also accumulates P*V in F16.

This PR removes the restriction on 64-bit x86 and handles the leftover columns of the P*V tile GEMM with masked AVX-512 loads/stores/FMA instead of scalar per-column dot products. The behavior is unchanged outside 64-bit x86.

Additional information

FA operator, 16 heads, F16 K/V, Ryzen 9 9950X, 16 threads:

D Q=KV master, ms this PR, ms speedup
40 784 34.48 1.61 21.4x
40 3136 548.90 19.66 27.9x
72 784 25.13 2.03 12.4x
72 3136 399.84 24.85 16.1x
72 5776 1393.14 84.70 16.4x

Qwen3-VL BF16 mmproj on CPU:

image master, ms this PR, ms speedup
448x448 967.9 338.7 2.86x
896x896 13147.9 1822.8 7.21x
1120x1120 28878.4 3487.1 8.28x

tested on x86_64 only.

Correctness: 225/225 CPU FA backend tests pass, including a new DK=65/DV=67 case with mask, sinks and ALiBi.

Requirements

@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning labels Sep 25, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 25, 2026

Copy link
Copy Markdown

Hi @SongXiaoXi, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

Comment thread ggml/src/ggml-cpu/simd-gemm.h
@am17an

am17an commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

/bot review

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

❌ Code review failed.

Error: failed to fetch diff (503) from https://github.com/ggml-org/llama.cpp/pull/29423.diff

@SongXiaoXi

Copy link
Copy Markdown
Contributor Author

/bot review

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 26, 2026

Copy link
Copy Markdown

Sorry @SongXiaoXi, only members of the ggml-org organization can trigger /bot commands.

@SongXiaoXi

Copy link
Copy Markdown
Contributor Author

Hi @am17an, is there anything I can do to get the bot review working or help move this PR forward?

@am17an

am17an commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

/bot review

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

❌ Code review failed.

Error: failed to fetch diff (503) from https://github.com/ggml-org/llama.cpp/pull/29423.diff

@SongXiaoXi

Copy link
Copy Markdown
Contributor Author

Thanks for retrying! It failed with the same failed to fetch diff (503) error again. Is this likely a GitHub issue or a bot issue? I can rebase the branch or try anything else you think might help.

@am17an

am17an commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Idk. Time to retire the bot. Generated from one that works. Fix this one and it should be good with me.


[P2] Soft-capping gives padded KV entries nonzero attention weight — ops.cpp:9323
Removing the dimension guard exposes an existing tiled-kernel bug to previously correct inputs: padding is set to -INFINITY before soft-capping, which converts it to finite -logit_softcap. Without a mask, those nonexistent keys contribute to softmax.
Reproduced with DK=65, DV=67, Q=75, KV=17, four heads, zero Q/K, unit V, no mask, and logit_softcap=1:

  • Before PR: 1.000000
  • PR, AVX2 and AVX-512: 0.495767
    Apply the padding exclusion after soft-capping and add an unmasked regression case. The new test uses a mask and disables soft-capping, so it misses this.
    Otherwise, validation passed: 1,011 selected attention tests on each ISA, another 224 head-size-40 tests on AVX-512, and 6,750 GEMM shapes per SSE3/AVX/AVX2/AVX-512 configuration with protected-page boundary checks.

@am17an am17an added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 28, 2026
@ggerganov
ggerganov merged commit 6f767fe into ggml-org:master Sep 28, 2026
27 of 31 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
…ims on x86 (ggml-org#29423)

* ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86

* add AVX2 support for masked loading and storing in simd_gemm_ukernel_tail

* ggml-cpu: fix FA softcap handling for padded KV tiles
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…ims on x86 (ggml-org#29423)

* ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86

* add AVX2 support for masked loading and storing in simd_gemm_ukernel_tail

* ggml-cpu: fix FA softcap handling for padded KV tiles
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants