Skip to content

perf(builtin): enable flash attention on Metal, CoreML and CUDA - #536

Open
buxuku wants to merge 1 commit into
mainfrom
perf/enable-flash-attn
Open

buxuku wants to merge 1 commit into
mainfrom
perf/enable-flash-attn

Conversation

@buxuku

@buxuku buxuku commented Oct 10, 2026

Copy link
Copy Markdown
Owner

Summary

The built-in whisper.cpp engine passed flash_attn: false on every call, while upstream has enabled flash attention (FA) by default since v1.8.0 (ggml-org/whisper.cpp#3441). This turns it on for the backends where there is evidence that it helps, and leaves every other backend exactly as it was.

What changes

  • Add shouldUseFlashAttn(backend) to main/helpers/engines/transcribeShared.ts: true for metal, coreml, cuda; false for cpu, vulkan, custom.
  • Use it at both addon call sites: builtinEngine.ts (main transcription pipeline) and voiceClone/referenceTranscriber.ts (reference-clip transcription).
  • Add a table-driven test in scripts/test-engine-units.ts. The expectation table is typed Record<WhisperBackend, boolean>, so adding a backend later fails to compile until someone decides its FA setting.

The addon already reads flash_attn from JS (its init log prints flash attn = 0/1), so no addon rebuild is needed. scripts/longgap/run*.ts still hard-code flash_attn: false; they are experiment harnesses and I left them alone so their numbers stay comparable.

Why not simply true everywhere

Opening these up later is a one-line change in shouldUseFlashAttn (return true; for blanket-on), ideally after an A/B with the matching addon.

Testing

Why enable it, from upstream's scripts/bench-all-gg.txt (same device and commit, FA off vs on, small to large models): Apple Silicon encode -15 to -33 %, decode -9 to -22 %; NVIDIA encode -53 to -70 %, decode -6 to -24 %, but only on Blackwell (RTX 5090, DGX Spark). It is not a free win everywhere: one row (M1 Pro, tiny) shows decode 7.6 % slower with FA, and encoder/decoder times in that file do not convert directly to end-to-end time.

Local A/B on an Apple M1 Pro, using the addon.node / addon.coreml.node the app ships, called the way the app calls them (process.dlopen + promisify) with the app's whisperParams. FA off and on were run in alternating processes; the first call in each process was discarded as Metal pipeline warm-up and the rest are medians. The machine was shared and under load, so treat timings as indicative. English audio, language set to en.

model, audio encode decode whole call word edit distance, off vs on
base fp16, 4 min speech -29.1 % -8.6 % -10.7 % 6 (of 694 words)
base q8_0, 4 min speech -28.4 % -20.7 % -18.7 % 3 (of 696 words)
base fp16, 30 s -29.0 % -13.7 % -8.8 % 0 (last segment ends 30.00 s -> 29.96 s)
base fp16, 11 s -29.8 % -23.6 % -17.1 % 0 (one segment instead of two)

Output is not bit-identical with FA (float accumulation order changes), so a few decoding forks and segment boundaries move. On the 4-minute clip, tokens that line up between the two transcripts shift by a median of 50 ms (p90 400 ms, max 950 ms). To put that in context, running the same clip on CPU instead of Metal, both with FA off, gives the same spread (edit distance 6, median 50 ms, p90 410 ms, max 950 ms), and Metal with FA on matched the CPU transcript word for word. Repeating the same setting in separate processes gave identical output in every repeat I made. This is one clip and one model family, so it shows FA stays within normal backend noise there, not that it is lossless in general.

CoreML: addon.coreml.node with ggml-tiny and its encoder .mlmodelc, 11 s clip. It loads and runs with flash attn = 1, output identical to FA off, encode time unchanged (encoder is on the ANE). Decode went 44-55 ms -> 38-40 ms, which is too small a model to mean anything.

  • npm run test:engines: 982 passed, 0 failed (976 before, +6 new); the sherpa and Qwen scaffold scripts pass.
  • tsc --noEmit clean for renderer, main and automation; prettier clean.
  • Mutation check: making the helper always return true fails exactly the cpu, vulkan and custom cases.
  • To confirm on a device, the task log's whisperParams: line now shows flash_attn: true on Metal, CoreML and CUDA.

Not run:

  • CUDA, Vulkan and CPU, and anything on Windows or Linux: no such hardware here. CUDA is enabled on the strength of upstream's Blackwell-only numbers; older NVIDIA architectures and the CUDA 11.8 / 12.2 addon builds were not exercised.
  • Models larger than base (none installed locally), so the bigger encoder-bound gains upstream reports for small to large models were not reproduced.
  • CoreML beyond the tiny model and a short clip; a full app run through the UI.

The built-in whisper.cpp engine hard-coded flash_attn=false, while upstream
has enabled it by default since v1.8.0. Add shouldUseFlashAttn(backend) and
use it at both addon call sites (the main transcription pipeline and the
voice-clone reference-clip transcriber).

Enabled for metal, coreml and cuda, where upstream benchmarks and a local
Metal A/B show lower encode/decode time. Left off, i.e. unchanged behaviour,
for cpu (no data), vulkan (the shipped ggml predates the flash-attention
shader out-of-bounds fix, llama.cpp #29988) and custom addons (unknown build).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant