Skip to content

kleidiai : add SME2 f32 kernel - #24414

Merged
ggerganov merged 2 commits into
ggml-org:masterfrom
chaxu01:feature/sme2-f32-kernel
Jul 14, 2026
Merged

ggerganov merged 2 commits into
ggml-org:masterfrom
chaxu01:feature/sme2-f32-kernel

Conversation

@chaxu01

@chaxu01 chaxu01 commented Jun 10, 2026

Copy link
Copy Markdown
Collaborator

Overview

This patch integrates the KleidiAI SME2 F32 GEMM/GEMV kernels into the kleidiai backend and enable runtime dispatch to the optimized SME2 implementations when supported by the target system.

Additional information

Benchmark results on an Apple M4 Pro (2 SME units) using Llama-3.2-1B F32 show substantial improvements over the native F32 implementation:

Threads Native F32 (t/s) SME2 F32 (t/s) Speedup
1 19.70 457.57 23.2×
2 37.78 892.35 23.6×
4 77.28 561.52 7.3×
6 104.57 452.55 4.3×
8 122.40 609.17 5.0×

Requirements

@chaxu01
chaxu01 requested a review from ggerganov as a code owner June 10, 2026 13:05
@github-actions github-actions Bot added the ggml changes relating to the ggml tensor library for machine learning label Jun 10, 2026
@chaxu01

chaxu01 commented Jun 23, 2026

Copy link
Copy Markdown
Collaborator Author

@ggerganov — just wanted to highlight this PR when you have a chance. The patch integrates the KleidiAI SME2 F32 GEMM/GEMV kernels into the KleidiAI backend and adds runtime dispatch to the optimized SME2 implementations when supported by the target system.

Happy to provide any additional details if helpful. Thanks!

@chaxu01
chaxu01 force-pushed the feature/sme2-f32-kernel branch from c8cfe47 to 287abd6 Compare June 30, 2026 06:11
@chaxu01

chaxu01 commented Jun 30, 2026

Copy link
Copy Markdown
Collaborator Author

@ggerganov, I've just updated this PR with an improvement to the SME2 F32 path by enabling dynamic scheduling.

The original implementation used static column partitioning, which works well up to the number of available SME units but doesn't scale well beyond that. With dynamic scheduling enabled, the F32 kernels maintain much better utilization when the thread count exceeds the available SME units.

Threads Static SME2 F32 Dynamic Scheduling
1 457.57 450.83
2 892.35 1067.71
4 561.52 938.98
6 452.55 1179.90
8 609.17 1114.89

Whenever you have a chance, I'd appreciate it if you could take a look. Happy to address any feedback or make further changes if needed. Thanks!

@chaxu01

chaxu01 commented Jul 14, 2026

Copy link
Copy Markdown
Collaborator Author

@ggerganov — just wanted to follow up on this PR when you have a chance.

The PR has been updated with dynamic scheduling for the SME2 F32 path, which significantly improves scaling once the thread count exceeds the available SME units.

When you have a chance to take a look, I'd really appreciate any feedback or suggestions you may have. I'm happy to address any comments or make further improvements if needed. Thanks!

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, I missed the notifications for this.

@ggerganov ggerganov added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Jul 14, 2026
@ggerganov
ggerganov merged commit 47c7869 into ggml-org:master Jul 14, 2026
27 checks passed
RehanQasim-dev pushed a commit to aifoundry-org/llama.cpp that referenced this pull request Jul 23, 2026
* kleidiai : add SME2 f32 kernel

* enable dynamic scheduling for SME2 f32 kernel
RehanQasim-dev pushed a commit to aifoundry-org/llama.cpp that referenced this pull request Jul 23, 2026
* kleidiai : add SME2 f32 kernel

* enable dynamic scheduling for SME2 f32 kernel
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
* kleidiai : add SME2 f32 kernel

* enable dynamic scheduling for SME2 f32 kernel
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* kleidiai : add SME2 f32 kernel

* enable dynamic scheduling for SME2 f32 kernel
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* kleidiai : add SME2 f32 kernel

* enable dynamic scheduling for SME2 f32 kernel
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* kleidiai : add SME2 f32 kernel

* enable dynamic scheduling for SME2 f32 kernel
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants