Skip to content

opencl: route large narrow-hidden q6_K lm_head to the flat GEMV - #26427

Merged
max-krasnyansky merged 1 commit into
ggml-org:masterfrom
qualcomm:hq/q6k-flat-large-m-size-gate-r0730
Aug 3, 2026
Merged

max-krasnyansky merged 1 commit into
ggml-org:masterfrom
qualcomm:hq/q6k-flat-large-m-size-gate-r0730

Conversation

@wanghqc

@wanghqc wanghqc commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

Overview

The problem:

  • The opencl backend has two q6_K GEMV kernels: the fast flat one and the slower gemv_noshuffle fallback.
  • The rule that picks between them requires the weight to be big in both dimensions (ne[1] >= 32768 && ne[0] >= 2048).
  • gemma-4 / gemma3n E2B ties its token embedding as the lm_head: a Q6_K weight of [1536 × 262144] — 330 MiB. It is huge, but its first dimension is 1536 < 2048, so the rule rejects it and it runs on the slow kernel every generated token:

Solution:

  • Keep the existing rule exactly as it is, and OR-in one extra way to qualify: total weight size ≥ 256 MiB.
  • One exception: the escape is skipped on the A7X (740), whose compiler miscompiles the flat q6_K kernel (wrong answers).
  • Nothing changes for any weight the old rule already accepted, on any device.

Additional information

Performance on the X2-90 for gemma-4 E2B:

  • Per-kernel is now ~23 GB/s => ~56 GB/s
  • End to end (e2e) TG128 is now 27.62 → 38.58 = +39.7%

Requirements

@wanghqc
wanghqc requested a review from a team as a code owner August 2, 2026 02:24
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend labels Aug 2, 2026
* add a direct size condition for `large` weights; the original
  dimension condition is insufficient -- q6_K lm_head for gemma-4 E2B
  has [1536, 262144], which is big enough to slowdown gemv_noshuffle but
  does not satisfy the dimension condition (ne0 >= 2048)
@lhez
lhez force-pushed the hq/q6k-flat-large-m-size-gate-r0730 branch from 0b1c734 to 2183862 Compare August 2, 2026 14:49
@max-krasnyansky
max-krasnyansky merged commit 39eab74 into ggml-org:master Aug 3, 2026
24 of 27 checks passed
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
* add a direct size condition for `large` weights; the original
  dimension condition is insufficient -- q6_K lm_head for gemma-4 E2B
  has [1536, 262144], which is big enough to slowdown gemv_noshuffle but
  does not satisfy the dimension condition (ne0 >= 2048)
brittlewis12 pushed a commit to brittlewis12/llama.cpp that referenced this pull request Aug 17, 2026
* add a direct size condition for `large` weights; the original
  dimension condition is insufficient -- q6_K lm_head for gemma-4 E2B
  has [1536, 262144], which is big enough to slowdown gemv_noshuffle but
  does not satisfy the dimension condition (ne0 >= 2048)
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
* add a direct size condition for `large` weights; the original
  dimension condition is insufficient -- q6_K lm_head for gemma-4 E2B
  has [1536, 262144], which is big enough to slowdown gemv_noshuffle but
  does not satisfy the dimension condition (ne0 >= 2048)
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* add a direct size condition for `large` weights; the original
  dimension condition is insufficient -- q6_K lm_head for gemma-4 E2B
  has [1536, 262144], which is big enough to slowdown gemv_noshuffle but
  does not satisfy the dimension condition (ne0 >= 2048)
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* add a direct size condition for `large` weights; the original
  dimension condition is insufficient -- q6_K lm_head for gemma-4 E2B
  has [1536, 262144], which is big enough to slowdown gemv_noshuffle but
  does not satisfy the dimension condition (ne0 >= 2048)
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* add a direct size condition for `large` weights; the original
  dimension condition is insufficient -- q6_K lm_head for gemma-4 E2B
  has [1536, 262144], which is big enough to slowdown gemv_noshuffle but
  does not satisfy the dimension condition (ne0 >= 2048)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants