Skip to content

ggml : add PQ2_0 and PTQ1_0 ternary types; llama : apply Hadamard-folded weights - #29077

Closed
QuentinDanblon wants to merge 1 commit into
ggml-org:masterfrom
QuentinDanblon:prism-ternary-hadamard-fold
Closed

QuentinDanblon wants to merge 1 commit into
ggml-org:masterfrom
QuentinDanblon:prism-ternary-hadamard-fold

Conversation

@QuentinDanblon

Copy link
Copy Markdown

Adds support for the ternary GGUFs published as
prism-ml/Ternary-Bonsai-2-27B-gguf
(Apache-2.0, 27B, qwen35). Two things are missing today, and both are needed: the tensor
types, and the activation-side transform for weights that are stored folded into a rotated basis.

Requested in #29058. Today such a file either fails to parse ("the type ids sit past
GGML_TYPE_COUNT") or, when the type is one llama.cpp already knows (Q2_0), loads without a
warning and produces incoherent text, because no mainline runtime applies the rotation declared in
the file's metadata.

What is in here

1. The two tensor types (ggml)

type id group block layout
PQ2_0 142 128 34 B fp16 scale + each trit in a 2-bit slot, same codec as Q2_0
PTQ1_0 143 128 28 B fp16 scale + dense base-3 trits (5/byte, remainder at 4/byte), as in TQ1_0

They are group-128 variants of Q2_0/TQ1_0, so the reference quantizers/dequantizers are ports of
the existing ones with the group size changed. The ids are deliberately high: the published files
already carry 142/143, and matching them is the only way to read those artifacts.

CPU wiring: from_float (reference quantizer), to_float, vec_dot_type = Q8_0, and a vec_dot.
The vec_dot here is a reference kernel — dequantize the block, accumulate in F32 — which is
correct but slow (~0.4 tok/s for the 27B on 16 threads). SIMD paths for ggml-cpu are a deliberate
follow-up, together with the GPU backends (see below).

2. The folded-weight runtime (llama)

prism.hadamard.* metadata describes weights stored as W' with the model defined as
y = W' * (H * (s * P * x)), where H is a normalized Sylvester-Hadamard matrix applied in blocks of
block_size along the input axis, s a per-input-width sign vector, and P an optional feature
permutation for the grouped GDN path. The metadata is self-describing and is validated before it is
trusted:

  • version, block size (power of two, must divide every named weight's input dimension), transform and
    axis names, sign mode, and non-empty weight list;
  • explicit sign mode must supply widths whose lengths sum exactly to the sign-value count, and every
    value must be +1/-1;
  • only architectures whose matmuls are all routed through build_lora_mm/build_lora_mm_id are
    accepted (llama, qwen3, qwen3moe, qwen35, qwen35moe, qwen3next), and only tensor kinds
    verified to be on those paths may be folded. Anything else throws instead of running wrong math;
  • the inverse path is limited to token_embd.weight, which is the only row-lookup table the graph
    restores (h = s * (H z) after the lookup).

One rotation matrix is built per (block size, buffer type) and one sign vector per input width, both
as persistent model tensors; the per-weight association lives in the model and reaches the graph
through llm_graph_params. Transformations on the same activation are memoized per graph build so
weights sharing an activation do not each rebuild one.

This reuses what already exists upstream: llama_mul_mat_hadamard and the GGML_HINT_SRC0_IS_HADAMARD
hint, the FWHT kernels, and the same "one cached rotation matrix per shape" pattern as the attention
KV-rotation feature.

Verification

  • Numerical equivalence with the reference implementation. Perplexity on
    wikitext-2-raw/wiki.test.raw, -c 64 --chunks 2: this branch gives
    PPL = 16.2538 and Prism's fork built from the same tag gives 16.2588 — a 0.03% difference,
    consistent with this branch dequantizing to F32 while the fork uses Q8 dot kernels. If the transform
    were wrong or missing, the number is nonsense rather than 0.03% off.
  • tests/test-quantize-fns passes for both new types (the ternary thresholds in that test now include
    them; without that entry the types are judged against the default threshold).
  • llama-cli loads the 27B PQ2_0 file, and llama-bench runs it (pp8/tg1, CPU only).

Not in here

  • GPU kernels. CUDA/Metal/Vulkan/SYCL have no PQ2_0/PTQ1_0 support in this PR, so the types
    fall back to CPU there. Porting the fork's kernels is the next step and is what makes these packs
    practical (the same 27B runs at ~135 tok/s on an RTX 5090 with them).
  • Fast CPU kernels (repack, SIMD vec_dot).
  • LLAMA_FTYPE_MOSTLY_PQ2_0/PTQ1_0 and quantize-tool support: the published files report
    general.file_type 141/143, which currently maps to "unknown" in the loader's log. Worth pinning
    down, but nothing depends on it yet.

Testing notes

Built CPU-only with cmake -B build -DGGML_CUDA=OFF -DLLAMA_CURL=OFF, gcc in a container; the model
was loaded from the published pack (6.7 GiB, 402 PQ2_0 tensors, 401 folded weights, 2 rotation
matrices and 3 sign vectors at block size 1024).

…ded weights

PQ2_0 (type 142) and PTQ1_0 (type 143) are group-128 ternary types used by Prism
ML's Ternary-Bonsai GGUFs. PQ2_0 reuses the Q2_0 2-bit codec and PTQ1_0 the TQ1_0
trit packing, both with one fp16 scale per group of 128 weights. The ids match the
values already written into the published files, so they are kept as-is rather than
renumbered.

CPU gets reference quantizers/dequantizers plus a vec_dot that dequantizes the block
and accumulates in F32: correct but slow. SIMD kernels and the GPU backends are
deliberate follow-ups.

The second half is what makes those files correct to run at all. The metadata declares
that a set of matmul weights is stored folded into a rotated basis
(y = W' * (H * (s * P * x))), so a runtime that ignores it loads without a warning and
produces incoherent text. The metadata is validated before it is trusted (version, block
size, transform, axis, sign mode, non-empty weight list, sign widths summing exactly to
the sign values, architectures and tensor kinds verified to route every matmul through
build_lora_mm/build_lora_mm_id, and token_embd as the only inverse-after-lookup table).
One rotation matrix is built per (block size, buffer type) and one sign vector per input
width, as persistent tensors; the per-weight association reaches the graph through
llm_graph_params. This reuses llama_mul_mat_hadamard and GGML_HINT_SRC0_IS_HADAMARD.

Verified against the reference fork on a 27B pack: perplexity 16.2538 here versus
16.2588 there (wikitext-2-raw, -c 64 --chunks 2), and test-quantize-fns passes for both
new types.
@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning labels Sep 18, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 18, 2026

Copy link
Copy Markdown

Hi @QuentinDanblon, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

  • AI-generated content: While code is allowed to be generated by AI, please write the PR description and commit messages on your own without the help of AI.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@QuentinDanblon
QuentinDanblon marked this pull request as ready for review September 18, 2026 11:40
@CISC

CISC commented Sep 18, 2026

Copy link
Copy Markdown
Member

Please leave this for PrismML to submit themselves. cc/@khosravipasha

@khosravipasha

khosravipasha commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

@CISC
Thanks will be more few small PRs soon, one already got merged. #27779
Our goal is to supprot new changes (hadamard + sign flips) with official Q2_0.
Uploaded this Q2_0 gguf in meantime to help with testing/dev work: https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf-dev

@QuentinDanblon
PTQ1_0 and PQ2_0 might stay in fork only since adding new types is increased maintenance work. We are happy to add if needed.

Main reason we did PTQ1_0 is TQ1_0 needs group size 256 but our model is 128 group size.

PQ2_0 is same as Q2_0 just group size is 128 instead of 64 so only slightly smaller. Model weights are same for all.

@Nekotekina

Copy link
Copy Markdown
Contributor

If it was me, I'd just reduce group size of old TQ types, old BitNet model can be simply repackaged for new group size if someone needs them.

@zhaoyilun

Copy link
Copy Markdown

One observation from the cheap seats, not a request for changes — PrismML know this machinery far better than I do, and the shape of this PR looks right.

In master, src/models/qwen35.cpp reads its MTP token embedding with a raw ggml_get_rows(ctx0, tok_embd_w, inp->tokens) and no inverse transform. That is harmless today, because nothing on master folds a Hadamard rotation into token_embd.weight. After this PR it would stop being harmless: a folded export run with --spec-type draft-mtp gets its graph refused by the verifier:

llama_verify_hadamard_graph: latent lookup 'mtp_tok_embd-64' consumed by op=RMS_NORM name='norm-64' src0 hint=0
llama_init_from_model: failed to initialize the context: Hadamard-latent table 'token_embd.weight' is read without the inverse transform

That is not a hypothetical — it is what happened on PrismML's own fork, and it was fixed there on 2026-09-21 in #205 (422590f5): 14 lines in the same function, applying the same rot + signs inverse the trunk already applies in llm_graph_context::build_inp_embd(). That fix landed after this PR was opened, so it is naturally not part of this diff.

Raising it only so it does not have to be rediscovered later. Nothing needs to happen on this PR's side — and if carrying the MTP lookup along with the fold is already the plan, please ignore this entirely.

@CISC CISC closed this Sep 22, 2026
@KodeMunkie

Copy link
Copy Markdown

I was following this PR to see when Bonsai 2 is supported in llama.cpp officially. Is there another way of tracking what's left to do implement pls?

@Green-Sky

Copy link
Copy Markdown
Collaborator

I was following this PR to see when Bonsai 2 is supported in llama.cpp officially. Is there another way of tracking what's left to do implement pls?

currently: https://github.com/ggml-org/llama.cpp/pulls?q=is%3Apr+author%3Abri-prism

@XBold

XBold commented Sep 23, 2026

Copy link
Copy Markdown

Hi everyone,

First, I'd like to say a big thank you to all the maintainers and contributors here. The quality and depth of the discussion in this thread, and the care you all put into every PR, are genuinely one of the reasons llama.cpp has become the reference runtime for so many of us. The way @CISC and the team have handled this series of contributions, guiding it toward a clean, maintainable merge while keeping the project's standards intact, is a pleasure to watch.

I'm writing from the community side, as a regular user who is very excited to see Ternary-Bonsai-2 support landing upstream, so please take this as appreciation plus a status update, not as pressure.

For people still following this PR (like me): the umbrella PR here has been closed, and the work is now being brought in as a series of smaller PRs, which is exactly the right shape for this. As of today, the open pieces I can see from the series are:

And #29058 remains the umbrella for tracking the remaining pieces. (If the list has moved on since I wrote this, apologies, but I hope it helps as a map for anyone new to the thread.)

What is left, per my reading of the discussion: the tensor types themselves (PQ2_0/PTQ1_0) plus the Hadamard-folded weight runtime, the CUDA kernels for those types, and the fast CPU paths. The CUDA kernels are the piece that makes the packs genuinely fast rather than just loadable, so they are the big one for people like me running consumer GPUs.

The reason I'm nudge-ing gently: there is a lot of community interest and a lot of real-world demand waiting on this. The published 27B packs run great on the Prism fork, and the more people who can run them on mainline, the bigger the ecosystem gets. I know the maintainer bandwidth is precious, so the smallest ask is simply: whenever it is convenient, could the series be brought to a close so these packs work on a plain llama-cpp install?

Thank you, maintainers, contributors, and community, for the amazing work you are all doing. This project continues to be remarkable, and this thread is a great example of why.

Regards,
XBold

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants