Skip to content

metal : add GATED_LINEAR_ATTN op - #22176

Closed
TheTom wants to merge 1 commit into
ggml-org:masterfrom
TheTom:metal/gla-op-v2
Closed

TheTom wants to merge 1 commit into
ggml-org:masterfrom
TheTom:metal/gla-op-v2

Conversation

@TheTom

@TheTom TheTom commented Apr 20, 2026

Copy link
Copy Markdown

Add Metal backend support for GGML_OP_GATED_LINEAR_ATTN (GLA). Supports head_size 64 and 128, f32 only.

Overview

Metal is currently the only major GPU backend without GLA support. Without this, RWKV6Qwen2 models fall back to CPU for this op on every layer, causing GPU/CPU sync overhead that hurts performance significantly on lower-end Apple Silicon.

The kernel follows the existing RWKV WKV6 Metal structure: one thread per head element, threadgroup shared memory for k/q/gate, and a float4-vectorized inner loop.

supports_op restricts execution to f32 with head sizes 64 or 128.

Correctness

test-backend-ops (7/7 passed on both devices):

Before (master a6cc43c):

GATED_LINEAR_ATTN: not supported [MTL0] (0/0 tests)

After:

GATED_LINEAR_ATTN(head_size=64,  n_seq_tokens=1,   n_seqs=1): OK
GATED_LINEAR_ATTN(head_size=64,  n_seq_tokens=32,  n_seqs=1): OK
GATED_LINEAR_ATTN(head_size=64,  n_seq_tokens=32,  n_seqs=4): OK
GATED_LINEAR_ATTN(head_size=64,  n_seq_tokens=128, n_seqs=4): OK
GATED_LINEAR_ATTN(head_size=128, n_seq_tokens=1,   n_seqs=1): OK
GATED_LINEAR_ATTN(head_size=128, n_seq_tokens=32,  n_seqs=1): OK
GATED_LINEAR_ATTN(head_size=128, n_seq_tokens=32,  n_seqs=4): OK
7/7 tests passed

Performance

Benchmarked with QRWKV6-7B-Instruct Q4_K_M, llama-bench, 3 runs each.

M2 Mac Mini 32GB (Apple8)

Test Baseline (CPU fallback) Metal GLA Speedup
pp1 4.73 t/s 25.19 t/s +432%
pp512 194.88 t/s 214.69 t/s +10.2%
pp2048 200.33 t/s 214.95 t/s +7.3%
pp8192 201.46 t/s 215.38 t/s +6.9%
tg128 20.35 t/s 24.93 t/s +22.5%

M5 Max 128GB (Apple10)

Test Baseline (CPU fallback) Metal GLA Delta
pp1 72.49 t/s 71.14 t/s within noise
pp512 1260.14 t/s 1262.71 t/s within noise
pp2048 1261.07 t/s 1243.36 t/s within noise
pp8192 1195.14 t/s 1190.42 t/s within noise
tg128 70.64 t/s 73.19 t/s +3.6%

The M5 Max has enough GPU bandwidth that GLA is not the bottleneck at this model size. The M2 Mini shows the real impact of CPU fallback: 5.3x on single-token prefill and 22.5% on decode.

Perplexity

wikitext-2, 5 chunks, ctx=512:

Device Baseline Metal GLA
M5 Max 7.7383 7.7383
M2 Mini 7.7386 7.7382

Zero quality impact.

Notes

docs/ops.md and docs/ops/Metal.csv regenerated using test-backend-ops support --output csv and ./scripts/create_ops_docs.py.

Mentions #14909

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. AI used for research and guidance, all code manually reviewed and understood.

Add Metal backend support for GGML_OP_GATED_LINEAR_ATTN (GLA).
Supports head_size 64 and 128, f32 only.

Tested with test-backend-ops (7/7 passed):
- M5 Max 128GB (Apple10)
- M2 Mac Mini (Apple8)
@ggml-gh-bot

ggml-gh-bot Bot commented Apr 20, 2026

Copy link
Copy Markdown

Hi @TheTom, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 4 open PRs.

  • Large PR: Large changes require prior discussion (e.g. an issue or RFC) and maintainers may not be able to review this PR as-is. Consider splitting it into smaller, focused PRs.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@TheTom

TheTom commented Apr 20, 2026

Copy link
Copy Markdown
Author

Duplicate of #21452. Closing in favor of the original PR which has been updated with the clean commit and fresh benchmarks.

@TheTom TheTom closed this Apr 20, 2026
@github-actions github-actions Bot added documentation Improvements or additions to documentation testing Everything test related ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) labels Apr 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant