Skip to content

metal : add TQ2_0 support - #26980

Merged
ggerganov merged 2 commits into
masterfrom
gg/metal-tq2_0
Aug 13, 2026
Merged

ggerganov merged 2 commits into
masterfrom
gg/metal-tq2_0

Conversation

@ggerganov

@ggerganov ggerganov commented Aug 12, 2026 •

Copy link
Copy Markdown
Member

Overview

Implement missing TQ2_0 ops in the Metal backend.

Hoisting the per-byte coefficients

The TQ2_0 mul_mv kernel packs 4 quantized values per byte v, stored as 2-bit fields

$$f_j = (v \gg 2j) \,\&\, 3, \qquad j = 0..3$$

A byte's dot-product contribution is $\sum_j (f_j - 1)y_j$. Each field is a difference of consecutive floor-divisions by powers of 4:

$$f_j = \left\lfloor v / 4^{j} \right\rfloor - 4 \left\lfloor v / 4^{j+1} \right\rfloor$$

Substituting all four fields and collecting terms, the interior floors cancel (telescoping), leaving only the boundary floors — the byte divided by successive powers of 4:

$$\sum_j (f_j - 1)\,y_j \;=\; y_0 v + (y_1 - 4y_0)\left\lfloor \tfrac{v}{4} \right\rfloor + (y_2 - 4y_1)\left\lfloor \tfrac{v}{16} \right\rfloor + (y_3 - 4y_2)\left\lfloor \tfrac{v}{64} \right\rfloor - (y_0 + y_1 + y_2 + y_3)$$

Hence the inner loop's four f values are exactly $v$, $\lfloor v/4\rfloor$, $\lfloor v/16\rfloor$, $\lfloor v/64\rfloor$ — the survivors of the bit-field decomposition.

The coefficients $y_0$, $y_1-4y_0$, $y_2-4y_1$, $y_3-4y_2$ and the constant $y_0+y_1+y_2+y_3$ depend only on y, which is the same for every row. So the kernel "hoists" them out of the row loop, building them once per block and reusing them for all nr0 rows:

// once per block:  coef[j] = (Y0[j], Y1[j]-4Y0[j], Y2[j]-4Y1[j], Y3[j]-4Y2[j])
// once per block:  sumy    = (Y0[j]+Y1[j]) + (Y2[j]+Y3[j])
// per row:         sum = -sumy + coef[j][0]*v + coef[j][1]*f1 + coef[j][2]*f2 + coef[j][3]*f3

Net effect: the y-dependent setup runs once per block instead of once per (block × row), removing work proportional to the row count — and the per-byte work becomes pure float ops (floor + multiply) instead of integer bit masking/shifting.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. llama.cpp:DeepSeek-v4-Flash-0731

Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in
the Metal backend.

Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
@ggerganov
ggerganov requested a review from a team as a code owner August 12, 2026 18:43
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) testing Everything test related labels Aug 12, 2026
- float ops over integer ops
- precalculate sums
- hoist coef out of the inner loop
- contiguous y loads

llama.cpp:DeepSeek-v4-Flash-0731
@ggerganov
ggerganov merged commit 4a84b0a into master Aug 13, 2026
4 checks passed
@ggerganov
ggerganov deleted the gg/metal-tq2_0 branch August 13, 2026 11:48
brittlewis12 pushed a commit to brittlewis12/llama.cpp that referenced this pull request Aug 17, 2026
* metal: add TQ2_0 support

Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in
the Metal backend.

Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731

* cont : optimize mul_mv kernel

- float ops over integer ops
- precalculate sums
- hoist coef out of the inner loop
- contiguous y loads

llama.cpp:DeepSeek-v4-Flash-0731
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
* metal: add TQ2_0 support

Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in
the Metal backend.

Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731

* cont : optimize mul_mv kernel

- float ops over integer ops
- precalculate sums
- hoist coef out of the inner loop
- contiguous y loads

llama.cpp:DeepSeek-v4-Flash-0731
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* metal: add TQ2_0 support

Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in
the Metal backend.

Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731

* cont : optimize mul_mv kernel

- float ops over integer ops
- precalculate sums
- hoist coef out of the inner loop
- contiguous y loads

llama.cpp:DeepSeek-v4-Flash-0731
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* metal: add TQ2_0 support

Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in
the Metal backend.

Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731

* cont : optimize mul_mv kernel

- float ops over integer ops
- precalculate sums
- hoist coef out of the inner loop
- contiguous y loads

llama.cpp:DeepSeek-v4-Flash-0731
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant