Skip to content

vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain - #29520

Merged
0cc4m merged 1 commit into
ggml-org:masterfrom
fxgsell:vk-hc-post-gate-fusion
Sep 28, 2026
Merged

0cc4m merged 1 commit into
ggml-org:masterfrom
fxgsell:vk-hc-post-gate-fusion

Conversation

@fxgsell

@fxgsell fxgsell commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

The PR is a Vulkan-only fusion that runs qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain as one kernel and stops the reorder pass from splitting it; it gives exact output and +3.8% decode, passes all 18,993 tests.

Overview

qwen4exp (Qwen3.8-Flash-Next) builds each hyper-connection's scatter weights as w = 2*sigmoid(inject/hc): three tiny ops on a 4-element tensor, right before dsv4_hc_post. That happens twice per layer, 96 times per token. On Vulkan, decode is limited by the number of kernel dispatches, not by arithmetic, so those 288 dispatches per token cost real time.

The change, all inside the Vulkan backend:

Part What it does
New fusion HC_POST_GATE Detects SCALE -> SIGMOID -> SCALE -> DSV4_HC_POST and runs it as one dispatch; hc_post computes the gate while loading post[] into shared memory
Match check Fuses only when the op is sigmoid, neither scale has a bias, the input is f32, and the shapes match; anything else runs unfused
Reorder registration Without it, Vulkan's reorder pass split all 96 chains per token (0 left adjacent), so the fusion could never fire
Allocation guard add_pattern_alloc_deps keeps the input alive through the fused output, so the overlap check doesn't silently cancel the fusion
Test A gated variable on the existing test_dsv4_hc_post; no new test file

Checks:

Check Result
KL divergence vs master (40 x 512 tokens) 0.000000 (max 5.8e-5), same top token 100%
Decode, Flash-Next UD-IQ1_M 49.84 → 51.72 t/s (+3.8%)
Decode at 16k context 42.87 → 44.27 t/s (+3.3%)
hc_post op tests, R9700 and 7900 XT 9/9 each
Full test-backend-ops on Vulkan0 18,993/18,993 (master 18,990/18,990)
Prefill, 32k and 64k context no change

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, the code, tests and benchmarks were AI-written. Design/review by human.

…nel and stops the reorder pass from splitting it
@fxgsell
fxgsell requested review from a team and ggerganov as code owners September 27, 2026 08:33
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown

Hi @fxgsell, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • AI-generated content: While code is allowed to be generated by AI, please write the PR description and commit messages on your own without the help of AI.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@github-actions github-actions Bot added testing Everything test related Vulkan Issues specific to the Vulkan backend ggml changes relating to the ggml tensor library for machine learning labels Sep 27, 2026
@0cc4m 0cc4m changed the title qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain as one kernel vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain Sep 28, 2026
@0cc4m
0cc4m merged commit 03a667a into ggml-org:master Sep 28, 2026
21 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning testing Everything test related Vulkan Issues specific to the Vulkan backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants