Skip to content

cuda: add TQ2_0 support to the dequantize path and MMVQ - #28769

Closed
AlexGabbia wants to merge 8 commits into
ggml-org:masterfrom
AlexGabbia:feat/maple-cuda-tq2_0
Closed

AlexGabbia wants to merge 8 commits into
ggml-org:masterfrom
AlexGabbia:feat/maple-cuda-tq2_0

Conversation

@AlexGabbia

@AlexGabbia AlexGabbia commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Adds TQ2_0 support to the CUDA dequantize path and to MMVQ, needed to run ternary models on the GPU. Depends on #27000, which adds the Maple architecture this was developed against; until that lands the diff below also contains those commits and will shrink to the 6 CUDA files once it is merged.

TQ2_0 had no case in ggml_get_to_fp32_cuda / _fp16_ / _bf16_ or their _nc variants, so a GPU run of a ternary model asserted in ggml_cuda_mul_mat_cublas_impl. The MMVQ vector dot had a second, quieter problem: it was written assuming Q2_0's packed layout, where 4 consecutive elements share a byte.

TQ2_0 does not work that way. It stores 4 bit planes of 32 bytes, so element n sits at byte 32*(n/128) + n%32 with shift 2*((n%128)/32). The old code read the wrong bytes and paired weights with the wrong activations, which produced plausible-looking but wrong output rather than a crash. This rewrites the vector dot against that layout; the formula matches dequantize_row_tq2_0 and quantize_row_tq2_0_ref in ggml, the Metal dequantize_tq2_0 kernel, and ggml_vec_dot_tq2_0_q8_K on the CPU side.

Additional information

test-backend-ops, MUL_MAT and MUL_MAT_ID, on an RTX 5070 Ti Laptop (sm_120a, CUDA 12.8): 17 TQ2_0 cases pass, covering n=1..9 (so the ncols_dst variants of MMVQ, not just batch 1), k=256 and k=4096, and MUL_MAT_ID with n=1, n=16 and n=32. The full MUL_MAT/MUL_MAT_ID sweep is 2193/2193, so no regression on the other types.

Perplexity, same model, same file, same machine, only the backend changed:

path GPU CPU delta
cuBLAS, batch 512, full test split 85.1846 +/- 0.87378 85.2080 +/- 0.87385 0.03%
MMVQ, batch 1, 22 chunks 102.8874 +/- 5.42801 102.6038 +/- 5.40022 0.28%

The two sides use independent implementations (cuBLAS dequantizes to F16 and MMVQ uses q8_1, the CPU uses q8_K), so agreement at that level is the useful signal. mmvq.cu also compiles for sm_121a with CUDA 13.0, which is the arch the DGX Spark prefetch path is guarded on.

Caveats:

  • HIP is compiled but not numerically tested. The HIP path uses plain integer arithmetic rather than an AMD perm intrinsic, because it cannot be tested on the NVIDIA machine this was developed on. Both paths were checked to agree exhaustively over all 1024 input combinations, but that is an equivalence check, not a test on AMD hardware. CI compiles HIP (build-cuda-ubuntu.yml has a hip: job) but does not exercise TQ2_0 numerically on AMD, because ci/run.sh runs test-backend-ops with -b CPU only and test-llama-archs uses F16 tensors. If anyone has an AMD GPU, test-backend-ops -o MUL_MAT,MUL_MAT_ID -p TQ2_0 would close it.
  • This is a correctness fix, not a fast path. cuBLAS still dequantizes the whole weight tensor on every forward, so it does not make ternary models fast on GPU the way a normal quantized model is. A real MMQ kernel for TQ2_0 is the next step and is not in this PR.
  • Metal is untouched here. TQ2_0 and ggml_swiglu_clamp are both supported there, but Maple has not been run on Metal.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, code generated with AI assistance under my direction and reviewed by me before submitting

Add MODEL_ARCH.MAPLE, its "maple" name, and the tensor list for the
Maple 20B-A1B ternary MoE architecture: token embeddings, output,
attention with Q/K RMS norms, and per-expert FFN tensors.
Register MapleForCausalLM in the HF architecture map and add the
converter for the Maple 20B-A1B ternary MoE model: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, partial rotary factor 0.5, and
per-expert weight stacking into merged 3D tensors.
Add the Maple 20B-A1B ternary MoE architecture: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, and ternary TQ1_0/TQ2_0
quantization support.

- register LLM_ARCH_MAPLE between MAMBA2 and JAMBA
- implement llama_model_maple: Q/K RMS norms after projection (GEMMA4
  style), rope applied only on SWA layers (nope_on_global_attention),
  ISWA KV cache, and MoE FFN with swiglu gate clamp at +7 (DEEPSEEK4
  style)
- mark MAPLE as unsupported by the model saver (roundtrip skipped)
Maple is always-MoE: the model throws when n_expert == 0, so the
test harness must only run the MoE config for LLM_ARCH_MAPLE.
- load_arch_hparams: use n_ff_exp_arr + n_ff_exp() accessor (upstream
  changed these from a scalar member during the rebase)
- sliding_window_pattern: get_arr, the pattern is mandatory for this arch
- partial_rotary_factor: read only from rope_parameters (base.py mirrors
  the top-level key automatically)
- document why TOKEN_EMBD/OUTPUT are forced to F16 (they are the two
  dense tensors in Maple, and the reference GGUFs ship them as F16)
- add @ModelBase.example("deepgrove/maple-preview")
get_arr for maple.attention.sliding_window_pattern requires an array, but
the harness only emitted a per-layer array for the arches in its list, so
test-llama-archs -a maple failed to load the model.

Assisted-by: DeepSeek Harness
The loader prefilled 7.0 and read the key optionally. The converter now
writes it and the loader reads it as required, because llama-graph.cpp
skips the clamp when the limit is 0 and an optional read would silently
run unclamped. The test harness provides the key for the same reason.

Also drops tensor_force_quant: base.py already forces FFN_GATE_INP to F32
and TOKEN_EMBD/OUTPUT to F16 for ternary file types.

Assisted-by: DeepSeek Harness
TQ2_0 had no case in ggml_get_to_fp32_cuda / _fp16_ / _bf16_ and their
_nc variants, so a GPU run of a ternary model asserted in
ggml_cuda_mul_mat_cublas_impl. Add the dequantize path and wire TQ2_0
into the MMVQ dispatch.

The vector dot cannot reuse Q2_0's code: Q2_0 packs 4 consecutive
elements per byte, while TQ2_0 stores 4 bit planes of 32 bytes, so
element n sits at byte 32*(n/128) + n%32 with shift 2*((n%128)/32).

The HIP path uses plain arithmetic instead of an AMD perm intrinsic,
because it cannot be tested on the NVIDIA machine this was developed on.
Both paths were checked to agree exhaustively.

Assisted-by: DeepSeek Harness
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 11, 2026

Copy link
Copy Markdown

Hi @AlexGabbia, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@AlexGabbia

Copy link
Copy Markdown
Contributor Author

Closing this to respect the 1 open PR limit for new contributors, as flagged. The branch stays on my fork and the work is finished and verified; I will reopen once #27000 is merged, which is also when this diff becomes just the 6 CUDA files.

@AlexGabbia AlexGabbia closed this Sep 11, 2026
@github-actions github-actions Bot added model Model specific testing Everything test related ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend conversion labels Sep 11, 2026
@Green-Sky

Copy link
Copy Markdown
Collaborator

Having draft PRs is fine + its now merged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning model Model specific testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants