Skip to content

metal : add GATED_LINEAR_ATTN op - #21452

Open
TheTom wants to merge 2 commits into
ggml-org:masterfrom
TheTom:metal/gla-op
Open

TheTom wants to merge 2 commits into
ggml-org:masterfrom
TheTom:metal/gla-op

Conversation

@TheTom

@TheTom TheTom commented Apr 5, 2026 •

Copy link
Copy Markdown

Add Metal backend support for GGML_OP_GATED_LINEAR_ATTN (GLA). Supports head_size 64 and 128, f32 only.

Overview

Metal is currently the only major GPU backend without GLA support. Without this, RWKV6Qwen2 models fall back to CPU for this op on every layer, causing GPU/CPU sync overhead that hurts performance significantly on lower-end Apple Silicon.

The kernel follows the existing RWKV WKV6 Metal structure: one thread per head element, threadgroup shared memory for k/q/gate, and a float4-vectorized inner loop.

supports_op restricts execution to f32 with head sizes 64 or 128.

Correctness

test-backend-ops (7/7 passed on both devices):

Before (master a6cc43c):

GATED_LINEAR_ATTN: not supported [MTL0] (0/0 tests)

After:

GATED_LINEAR_ATTN(head_size=64,  n_seq_tokens=1,   n_seqs=1): OK
GATED_LINEAR_ATTN(head_size=64,  n_seq_tokens=32,  n_seqs=1): OK
GATED_LINEAR_ATTN(head_size=64,  n_seq_tokens=32,  n_seqs=4): OK
GATED_LINEAR_ATTN(head_size=64,  n_seq_tokens=128, n_seqs=4): OK
GATED_LINEAR_ATTN(head_size=128, n_seq_tokens=1,   n_seqs=1): OK
GATED_LINEAR_ATTN(head_size=128, n_seq_tokens=32,  n_seqs=1): OK
GATED_LINEAR_ATTN(head_size=128, n_seq_tokens=32,  n_seqs=4): OK
7/7 tests passed

Performance

Benchmarked with QRWKV6-7B-Instruct Q4_K_M, llama-bench, 3 runs each.

M2 Mac Mini 32GB (Apple8)

Test Baseline (CPU fallback) Metal GLA Speedup
pp1 4.73 t/s 25.19 t/s +432%
pp512 194.88 t/s 214.69 t/s +10.2%
pp2048 200.33 t/s 214.95 t/s +7.3%
pp8192 201.46 t/s 215.38 t/s +6.9%
tg128 20.35 t/s 24.93 t/s +22.5%

M5 Max 128GB (Apple10)

Test Baseline (CPU fallback) Metal GLA Delta
pp1 72.49 t/s 71.14 t/s within noise
pp512 1260.14 t/s 1262.71 t/s within noise
pp2048 1261.07 t/s 1243.36 t/s within noise
pp8192 1195.14 t/s 1190.42 t/s within noise
tg128 70.64 t/s 73.19 t/s +3.6%

The M5 Max has enough GPU bandwidth that GLA is not the bottleneck at this model size. The M2 Mini shows the real impact of CPU fallback: 5.3x on single-token prefill and 22.5% on decode.

Perplexity

wikitext-2, 5 chunks, ctx=512:

Device Baseline Metal GLA
M5 Max 7.7383 7.7383
M2 Mini 7.7386 7.7382

Zero quality impact.

Notes

docs/ops.md and docs/ops/Metal.csv regenerated using test-backend-ops support --output csv and ./scripts/create_ops_docs.py.

Mentions #14909

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. AI used for research and guidance, all code manually reviewed and understood.

@TheTom
TheTom requested review from a team and ggerganov as code owners April 5, 2026 00:58
@github-actions github-actions Bot added documentation Improvements or additions to documentation testing Everything test related ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) labels Apr 5, 2026
@ggml-gh-bot

This comment was marked as resolved.

@TheTom

TheTom commented Apr 5, 2026

Copy link
Copy Markdown
Author

Regarding the automated flags:

  • Multiple open PRs: I currently have another open PR (metal: add opt-in V skip for negligible attention weights #21119). I have converted it to draft it to comply with the new contributor limit of one open PR.

  • Large PR: The functional change is ~170 lines across the Metal backend plus 3 test lines. The remaining diff is from regenerating docs/ops/Metal.csv and docs/ops.md via the required workflow, which picked up unrelated updates from recent changes in main.

@TheTom
TheTom force-pushed the metal/gla-op branch 2 times, most recently from 67ec85e to ed69c7b Compare April 18, 2026 18:00
@TheTom
TheTom requested a review from a team as a code owner April 18, 2026 18:00
@github-actions github-actions Bot added the Nvidia GPU Issues specific to Nvidia GPUs label Apr 18, 2026
@CISC
CISC removed the request for review from a team April 18, 2026 19:09
@CISC CISC removed the Nvidia GPU Issues specific to Nvidia GPUs label Apr 18, 2026
@CISC

CISC commented Apr 18, 2026

Copy link
Copy Markdown
Member

You seem to have included your other PR in this one.

@TheTom

TheTom commented Apr 18, 2026

Copy link
Copy Markdown
Author

Thank you @CISC . When fixing the conflict with ./scripts/create_ops_docs.py I accidentally included other fixes for another branch. Will address.

@github-actions github-actions Bot added the Nvidia GPU Issues specific to Nvidia GPUs label Apr 18, 2026
@CISC CISC removed the Nvidia GPU Issues specific to Nvidia GPUs label Apr 18, 2026
@TheTom

TheTom commented Apr 20, 2026

Copy link
Copy Markdown
Author

Requesting review from @ggml-org/ggml-metal

@aminya

aminya commented May 5, 2026 •

Copy link
Copy Markdown

@CISC Would you mind taking another look? The issue is addressed.
Note that, this PR is a precursor to adding TurboQuant support
TheTom#113

@CISC

CISC commented May 5, 2026

Copy link
Copy Markdown
Member

@CISC Would you mind taking another look? The issue is addressed.

I cannot review Metal PRs (I don't even have the necessary hardware), wait for @ggml-org/ggml-metal

@TheTom

TheTom commented May 17, 2026

Copy link
Copy Markdown
Author

@ggerganov any chance I can get a review?

Add Metal backend support for GGML_OP_GATED_LINEAR_ATTN (GLA).
Supports head_size 64 and 128, f32 only.

Tested with test-backend-ops (7/7 passed):
- M5 Max 128GB (Apple10)
- M2 Mac Mini (Apple8)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants