metal : add TQ2_0 support - #26980
Merged
Merged
Conversation
Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in the Metal backend. Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
ggerganov
force-pushed
the
gg/metal-tq2_0
branch
from
August 13, 2026 11:22
2b27b16 to
8740776
Compare
- float ops over integer ops - precalculate sums - hoist coef out of the inner loop - contiguous y loads llama.cpp:DeepSeek-v4-Flash-0731
ggerganov
force-pushed
the
gg/metal-tq2_0
branch
from
August 13, 2026 11:33
8740776 to
beeda4d
Compare
brittlewis12
pushed a commit
to brittlewis12/llama.cpp
that referenced
this pull request
Aug 17, 2026
* metal: add TQ2_0 support Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in the Metal backend. Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 * cont : optimize mul_mv kernel - float ops over integer ops - precalculate sums - hoist coef out of the inner loop - contiguous y loads llama.cpp:DeepSeek-v4-Flash-0731
thecodacus
pushed a commit
to thecodacus/llama.cpp
that referenced
this pull request
Sep 7, 2026
* metal: add TQ2_0 support Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in the Metal backend. Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 * cont : optimize mul_mv kernel - float ops over integer ops - precalculate sums - hoist coef out of the inner loop - contiguous y loads llama.cpp:DeepSeek-v4-Flash-0731
zbrad
pushed a commit
to zbrad/llama.cpp
that referenced
this pull request
Sep 10, 2026
* metal: add TQ2_0 support Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in the Metal backend. Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 * cont : optimize mul_mv kernel - float ops over integer ops - precalculate sums - hoist coef out of the inner loop - contiguous y loads llama.cpp:DeepSeek-v4-Flash-0731
pl752
pushed a commit
to pl752/llama.cpp
that referenced
this pull request
Sep 15, 2026
* metal: add TQ2_0 support Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in the Metal backend. Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 * cont : optimize mul_mv kernel - float ops over integer ops - precalculate sums - hoist coef out of the inner loop - contiguous y loads llama.cpp:DeepSeek-v4-Flash-0731
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Implement missing
TQ2_0ops in the Metal backend.Hoisting the per-byte coefficients
The TQ2_0
mul_mvkernel packs 4 quantized values per bytev, stored as 2-bit fieldsA byte's dot-product contribution is$\sum_j (f_j - 1)y_j$ . Each field is a difference of consecutive floor-divisions by powers of 4:
Substituting all four fields and collecting terms, the interior floors cancel (telescoping), leaving only the boundary floors — the byte divided by successive powers of 4:
Hence the inner loop's four$v$ , $\lfloor v/4\rfloor$ , $\lfloor v/16\rfloor$ , $\lfloor v/64\rfloor$ — the survivors of the bit-field decomposition.
fvalues are exactlyThe coefficients$y_0$ , $y_1-4y_0$ , $y_2-4y_1$ , $y_3-4y_2$ and the constant $y_0+y_1+y_2+y_3$ depend only on
y, which is the same for every row. So the kernel "hoists" them out of the row loop, building them once per block and reusing them for allnr0rows:Net effect: the
y-dependent setup runs once per block instead of once per (block × row), removing work proportional to the row count — and the per-byte work becomes pure float ops (floor+ multiply) instead of integer bit masking/shifting.Requirements