Repository navigation
ggml : add PQ2_0 and PTQ1_0 ternary types; llama : apply Hadamard-folded weights - #29077
QuentinDanblon wants to merge 1 commit into
Conversation
…ded weights PQ2_0 (type 142) and PTQ1_0 (type 143) are group-128 ternary types used by Prism ML's Ternary-Bonsai GGUFs. PQ2_0 reuses the Q2_0 2-bit codec and PTQ1_0 the TQ1_0 trit packing, both with one fp16 scale per group of 128 weights. The ids match the values already written into the published files, so they are kept as-is rather than renumbered. CPU gets reference quantizers/dequantizers plus a vec_dot that dequantizes the block and accumulates in F32: correct but slow. SIMD kernels and the GPU backends are deliberate follow-ups. The second half is what makes those files correct to run at all. The metadata declares that a set of matmul weights is stored folded into a rotated basis (y = W' * (H * (s * P * x))), so a runtime that ignores it loads without a warning and produces incoherent text. The metadata is validated before it is trusted (version, block size, transform, axis, sign mode, non-empty weight list, sign widths summing exactly to the sign values, architectures and tensor kinds verified to route every matmul through build_lora_mm/build_lora_mm_id, and token_embd as the only inverse-after-lookup table). One rotation matrix is built per (block size, buffer type) and one sign vector per input width, as persistent tensors; the per-weight association reaches the graph through llm_graph_params. This reuses llama_mul_mat_hadamard and GGML_HINT_SRC0_IS_HADAMARD. Verified against the reference fork on a 27B pack: perplexity 16.2538 here versus 16.2588 there (wikitext-2-raw, -c 64 --chunks 2), and test-quantize-fns passes for both new types.
|
Hi @QuentinDanblon, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
Please leave this for PrismML to submit themselves. cc/@khosravipasha |
|
@CISC @QuentinDanblon Main reason we did PTQ1_0 is TQ1_0 needs group size 256 but our model is 128 group size. PQ2_0 is same as Q2_0 just group size is 128 instead of 64 so only slightly smaller. Model weights are same for all. |
|
If it was me, I'd just reduce group size of old TQ types, old BitNet model can be simply repackaged for new group size if someone needs them. |
|
One observation from the cheap seats, not a request for changes — PrismML know this machinery far better than I do, and the shape of this PR looks right. In master, That is not a hypothetical — it is what happened on PrismML's own fork, and it was fixed there on 2026-09-21 in #205 ( Raising it only so it does not have to be rediscovered later. Nothing needs to happen on this PR's side — and if carrying the MTP lookup along with the fold is already the plan, please ignore this entirely. |
|
I was following this PR to see when Bonsai 2 is supported in llama.cpp officially. Is there another way of tracking what's left to do implement pls? |
currently: https://github.com/ggml-org/llama.cpp/pulls?q=is%3Apr+author%3Abri-prism |
|
Hi everyone, First, I'd like to say a big thank you to all the maintainers and contributors here. The quality and depth of the discussion in this thread, and the care you all put into every PR, are genuinely one of the reasons llama.cpp has become the reference runtime for so many of us. The way @CISC and the team have handled this series of contributions, guiding it toward a clean, maintainable merge while keeping the project's standards intact, is a pleasure to watch. I'm writing from the community side, as a regular user who is very excited to see Ternary-Bonsai-2 support landing upstream, so please take this as appreciation plus a status update, not as pressure. For people still following this PR (like me): the umbrella PR here has been closed, and the work is now being brought in as a series of smaller PRs, which is exactly the right shape for this. As of today, the open pieces I can see from the series are:
And #29058 remains the umbrella for tracking the remaining pieces. (If the list has moved on since I wrote this, apologies, but I hope it helps as a map for anyone new to the thread.) What is left, per my reading of the discussion: the tensor types themselves (PQ2_0/PTQ1_0) plus the Hadamard-folded weight runtime, the CUDA kernels for those types, and the fast CPU paths. The CUDA kernels are the piece that makes the packs genuinely fast rather than just loadable, so they are the big one for people like me running consumer GPUs. The reason I'm nudge-ing gently: there is a lot of community interest and a lot of real-world demand waiting on this. The published 27B packs run great on the Prism fork, and the more people who can run them on mainline, the bigger the ecosystem gets. I know the maintainer bandwidth is precious, so the smallest ask is simply: whenever it is convenient, could the series be brought to a close so these packs work on a plain Thank you, maintainers, contributors, and community, for the amazing work you are all doing. This project continues to be remarkable, and this thread is a great example of why. Regards, |
Adds support for the ternary GGUFs published as
prism-ml/Ternary-Bonsai-2-27B-gguf
(Apache-2.0, 27B,
qwen35). Two things are missing today, and both are needed: the tensortypes, and the activation-side transform for weights that are stored folded into a rotated basis.
Requested in #29058. Today such a file either fails to parse ("the type ids sit past
GGML_TYPE_COUNT") or, when the type is one llama.cpp already knows (Q2_0), loads without awarning and produces incoherent text, because no mainline runtime applies the rotation declared in
the file's metadata.
What is in here
1. The two tensor types (
ggml)PQ2_0Q2_0PTQ1_0TQ1_0They are group-128 variants of
Q2_0/TQ1_0, so the reference quantizers/dequantizers are ports ofthe existing ones with the group size changed. The ids are deliberately high: the published files
already carry 142/143, and matching them is the only way to read those artifacts.
CPU wiring:
from_float(reference quantizer),to_float,vec_dot_type = Q8_0, and avec_dot.The
vec_dothere is a reference kernel — dequantize the block, accumulate in F32 — which iscorrect but slow (~0.4 tok/s for the 27B on 16 threads). SIMD paths for
ggml-cpuare a deliberatefollow-up, together with the GPU backends (see below).
2. The folded-weight runtime (
llama)prism.hadamard.*metadata describes weights stored asW'with the model defined asy = W' * (H * (s * P * x)), whereHis a normalized Sylvester-Hadamard matrix applied in blocks ofblock_sizealong the input axis,sa per-input-width sign vector, andPan optional featurepermutation for the grouped GDN path. The metadata is self-describing and is validated before it is
trusted:
axis names, sign mode, and non-empty weight list;
value must be
+1/-1;build_lora_mm/build_lora_mm_idareaccepted (
llama,qwen3,qwen3moe,qwen35,qwen35moe,qwen3next), and only tensor kindsverified to be on those paths may be folded. Anything else throws instead of running wrong math;
token_embd.weight, which is the only row-lookup table the graphrestores (
h = s * (H z)after the lookup).One rotation matrix is built per (block size, buffer type) and one sign vector per input width, both
as persistent model tensors; the per-weight association lives in the model and reaches the graph
through
llm_graph_params. Transformations on the same activation are memoized per graph build soweights sharing an activation do not each rebuild one.
This reuses what already exists upstream:
llama_mul_mat_hadamardand theGGML_HINT_SRC0_IS_HADAMARDhint, the FWHT kernels, and the same "one cached rotation matrix per shape" pattern as the attention
KV-rotation feature.
Verification
wikitext-2-raw/wiki.test.raw,-c 64 --chunks 2: this branch givesPPL = 16.2538 and Prism's fork built from the same tag gives 16.2588 — a 0.03% difference,
consistent with this branch dequantizing to F32 while the fork uses Q8 dot kernels. If the transform
were wrong or missing, the number is nonsense rather than 0.03% off.
tests/test-quantize-fnspasses for both new types (the ternary thresholds in that test now includethem; without that entry the types are judged against the default threshold).
llama-cliloads the 27B PQ2_0 file, andllama-benchruns it (pp8/tg1, CPU only).Not in here
PQ2_0/PTQ1_0support in this PR, so the typesfall back to CPU there. Porting the fork's kernels is the next step and is what makes these packs
practical (the same 27B runs at ~135 tok/s on an RTX 5090 with them).
vec_dot).LLAMA_FTYPE_MOSTLY_PQ2_0/PTQ1_0and quantize-tool support: the published files reportgeneral.file_type141/143, which currently maps to "unknown" in the loader's log. Worth pinningdown, but nothing depends on it yet.
Testing notes
Built CPU-only with
cmake -B build -DGGML_CUDA=OFF -DLLAMA_CURL=OFF, gcc in a container; the modelwas loaded from the published pack (6.7 GiB, 402
PQ2_0tensors, 401 folded weights, 2 rotationmatrices and 3 sign vectors at block size 1024).