Repository navigation
cuda: add TQ2_0 support to the dequantize path and MMVQ - #28769
AlexGabbia wants to merge 8 commits into
Conversation
Add MODEL_ARCH.MAPLE, its "maple" name, and the tensor list for the Maple 20B-A1B ternary MoE architecture: token embeddings, output, attention with Q/K RMS norms, and per-expert FFN tensors.
Register MapleForCausalLM in the HF architecture map and add the converter for the Maple 20B-A1B ternary MoE model: 24 layers, 256 experts with 8 active, sliding-window attention (SWA-512) interleaved with global attention at a 3:1 ratio, partial rotary factor 0.5, and per-expert weight stacking into merged 3D tensors.
Add the Maple 20B-A1B ternary MoE architecture: 24 layers, 256 experts with 8 active, sliding-window attention (SWA-512) interleaved with global attention at a 3:1 ratio, and ternary TQ1_0/TQ2_0 quantization support. - register LLM_ARCH_MAPLE between MAMBA2 and JAMBA - implement llama_model_maple: Q/K RMS norms after projection (GEMMA4 style), rope applied only on SWA layers (nope_on_global_attention), ISWA KV cache, and MoE FFN with swiglu gate clamp at +7 (DEEPSEEK4 style) - mark MAPLE as unsupported by the model saver (roundtrip skipped)
Maple is always-MoE: the model throws when n_expert == 0, so the test harness must only run the MoE config for LLM_ARCH_MAPLE.
- load_arch_hparams: use n_ff_exp_arr + n_ff_exp() accessor (upstream
changed these from a scalar member during the rebase)
- sliding_window_pattern: get_arr, the pattern is mandatory for this arch
- partial_rotary_factor: read only from rope_parameters (base.py mirrors
the top-level key automatically)
- document why TOKEN_EMBD/OUTPUT are forced to F16 (they are the two
dense tensors in Maple, and the reference GGUFs ship them as F16)
- add @ModelBase.example("deepgrove/maple-preview")
get_arr for maple.attention.sliding_window_pattern requires an array, but the harness only emitted a per-layer array for the arches in its list, so test-llama-archs -a maple failed to load the model. Assisted-by: DeepSeek Harness
The loader prefilled 7.0 and read the key optionally. The converter now writes it and the loader reads it as required, because llama-graph.cpp skips the clamp when the limit is 0 and an optional read would silently run unclamped. The test harness provides the key for the same reason. Also drops tensor_force_quant: base.py already forces FFN_GATE_INP to F32 and TOKEN_EMBD/OUTPUT to F16 for ternary file types. Assisted-by: DeepSeek Harness
TQ2_0 had no case in ggml_get_to_fp32_cuda / _fp16_ / _bf16_ and their _nc variants, so a GPU run of a ternary model asserted in ggml_cuda_mul_mat_cublas_impl. Add the dequantize path and wire TQ2_0 into the MMVQ dispatch. The vector dot cannot reuse Q2_0's code: Q2_0 packs 4 consecutive elements per byte, while TQ2_0 stores 4 bit planes of 32 bytes, so element n sits at byte 32*(n/128) + n%32 with shift 2*((n%128)/32). The HIP path uses plain arithmetic instead of an AMD perm intrinsic, because it cannot be tested on the NVIDIA machine this was developed on. Both paths were checked to agree exhaustively. Assisted-by: DeepSeek Harness
|
Hi @AlexGabbia, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
Closing this to respect the 1 open PR limit for new contributors, as flagged. The branch stays on my fork and the work is finished and verified; I will reopen once #27000 is merged, which is also when this diff becomes just the 6 CUDA files. |
|
Having draft PRs is fine + its now merged. |
Overview
Adds TQ2_0 support to the CUDA dequantize path and to MMVQ, needed to run ternary models on the GPU. Depends on #27000, which adds the Maple architecture this was developed against; until that lands the diff below also contains those commits and will shrink to the 6 CUDA files once it is merged.
TQ2_0 had no case in
ggml_get_to_fp32_cuda/_fp16_/_bf16_or their_ncvariants, so a GPU run of a ternary model asserted inggml_cuda_mul_mat_cublas_impl. The MMVQ vector dot had a second, quieter problem: it was written assuming Q2_0's packed layout, where 4 consecutive elements share a byte.TQ2_0 does not work that way. It stores 4 bit planes of 32 bytes, so element
nsits at byte32*(n/128) + n%32with shift2*((n%128)/32). The old code read the wrong bytes and paired weights with the wrong activations, which produced plausible-looking but wrong output rather than a crash. This rewrites the vector dot against that layout; the formula matchesdequantize_row_tq2_0andquantize_row_tq2_0_refin ggml, the Metaldequantize_tq2_0kernel, andggml_vec_dot_tq2_0_q8_Kon the CPU side.Additional information
test-backend-ops, MUL_MAT and MUL_MAT_ID, on an RTX 5070 Ti Laptop (sm_120a, CUDA 12.8): 17 TQ2_0 cases pass, coveringn=1..9(so thencols_dstvariants of MMVQ, not just batch 1),k=256andk=4096, and MUL_MAT_ID withn=1,n=16andn=32. The full MUL_MAT/MUL_MAT_ID sweep is 2193/2193, so no regression on the other types.Perplexity, same model, same file, same machine, only the backend changed:
The two sides use independent implementations (cuBLAS dequantizes to F16 and MMVQ uses q8_1, the CPU uses q8_K), so agreement at that level is the useful signal.
mmvq.cualso compiles forsm_121awith CUDA 13.0, which is the arch the DGX Spark prefetch path is guarded on.Caveats:
build-cuda-ubuntu.ymlhas ahip:job) but does not exercise TQ2_0 numerically on AMD, becauseci/run.shrunstest-backend-opswith-b CPUonly andtest-llama-archsuses F16 tensors. If anyone has an AMD GPU,test-backend-ops -o MUL_MAT,MUL_MAT_ID -p TQ2_0would close it.ggml_swiglu_clampare both supported there, but Maple has not been run on Metal.Requirements