llama: add Maple 20B-A1B ternary MoE architecture (CPU) - #27000
Conversation
This comment was marked as resolved.
This comment was marked as resolved.
|
I don't see
|
|
Thanks for the review. :) This PR supports the official DeepGrove GGUFs https://huggingface.co/deepgrove/maple-preview-GGUF which load and generate correctly on CPU — I tested the TQ2_0 file on an Apple M4: 216 t/s prefill, 88 t/s generation. The stamsam GGUFs are built on a non-mainline fork (PrismML), so they don't help mainline support. I agree a 128-group ternary would be a good improvement and I'd be happy to work on it as a follow-up. |
|
Please provide perplexity numbers, as suggested by the contributor guidelines. |
|
Thanks — working on it. The official DeepGrove GGUFs include both TQ1_0 and TQ2_0 variants (with F16 and Q4_K heads). Since Maple is natively ternary, the F16 GGUFs from DeepGrove are the same ternary model with F16 head/embeddings — not a full-precision baseline. I'll run perplexity comparing TQ2_0 vs TQ1_0 (both official DeepGrove GGUFs) and report the numbers here. |
|
Someone with big internet please check if the conversion from source bf16 works, ideally @AlexGabbia . Some q8_0 ggufs for comparison would be nice too, pretty sure q8_0 or q4_0 would be lossless (for the ternary tensors). |
|
Ran a couple of ppl tests. I used their
The Maple uses swa over 512 tokens in the sparse attending layers, so 2048 should capture that. Overall the ppl is suspiciously high, but that might be because of masked training, not sure. |
|
Update with perplexity numbers — all 4 official DeepGrove GGUFs tested. Setup: wikitext-2, CPU M4 (10-core), 20 chunks, 8 threads,
Conversion BF16 → GGUF F16: works — 291 tensors, 40.5 GB, using the converter in this PR. Observations:
On the Q8_0: your q2_0 head-Q4_K embd-q8_0 numbers already cover that comparison. The conversion from TQ1_0 to q2_0 is lossless for the ternary tensors as you said, and the embedding variants (F16/Q8_0/Q6_K) show negligible PPL difference — so I think that base is covered. |
|
Your numbers are way too different. I ran the full test split. What did you run and how many chunks? edit:
Not really. If you rebase on master, you have access to metal tq2_0 btw. |
|
Thanks for the catch — you're right, my numbers aren't comparable to yours. What I ran: only the first 20 chunks (10,240 tokens) of wikitext-2, A full test-split run takes several hours on this M4 CPU (10-core, 8 threads) — I've started it now ( On Metal: fair point — master has TQ2_0 Metal (#26980) and our branch is 19 commits behind. I'll rebase and re-run the smoke test on Metal as a follow-up. |
|
For fun and because the model is so fast, I ran it on the full train split.
|
Thanks for your help on this PR @Green-Sky @AlexGabbia! Yeah I think any group size that is divisible by row dim will work as the model was trained with row wise scales. The default 256 group scales should work 👍 . |
cbfdb24 to
e935fb0
Compare
Thats great, thanks for noting that. Btw, have you considered adding more sparsity via things like ngram embeddings (longcat) or per-layer embeddings (gemma). Would be easy ways to add 5-10B without hurting inference speed much. Ideally even in ternary or binary. :) |
|
Ran the full wikitext-2 test split this time (not just 20 chunks). All 4 official DeepGrove GGUFs, CUDA on a RTX 5070 Ti, -ngl 99 -t 4. Branch rebased on latest master.
TQ1_0 and TQ2_0 give identical PPL — same weights, different packing, makes sense. Compared to your q2_0 numbers @Green-Sky: tq2_0 scores lower (79.76 vs 86.92 @512, 61.14 vs 64.83 @2048). Probably because TQ2_0 keeps the original ternary weights as-is while q2_0 re-quantizes them. head-Q4_K costs ~1.5 PPL @512, ~0.7 @2048 over head-F16. @2048 is much lower than @512 across the board — Maple's SWA-512 needs the context. |
Hm, I wonder whats wrong here, but the q2_0 weights themself are lossless. I made sure by requantizing them back to tq1_0 and the 2 files are hash identical. So the divergents needs to come from the ops running in different quantizaitons.
This reads wrong, slap your agent writing for you on the wrists. |
|
Fair point on the phrasing — corrected: the lower PPL @2048 vs @512 reflects the global-attention layers (every 4th layer) capturing long-range patterns that SWA-512 layers miss at short context. My agent earned the wrist-slap :D |
|
I reproduced your hash check — q2_0 → TQ1_0 round-trip is bit-identical. So the divergence isn't in the weights. To isolate the variable, I ran the same q2_0 file (head-Q4_K embd-f16, 6060.33 MiB / 2.51 BPW — matches your numbers exactly) across backends on the full wikitext-2 test split, same -c 512 -b 512:
On x86, q2_0 is consistent across backends (CUDA 80.81 vs CPU 80.78 — 0.03 gap). Same for TQ2_0. But your 86.92 is 6 PPL higher than my x86 CPU run on the same file. The only variable left is CPU architecture. Are you on ARM/Apple Silicon? If so, the gap is in the q2_0 ARM kernel, not in the model or the quant type — and TQ2_0 native is the numerically stable path. Happy to help debug the ARM kernel if you can share a repro. The architecture itself is simple enough that a standalone test should pinpoint it. (Fixed the SWA phrasing btw — the @2048 improvement comes from the global-attention layers, not from SWA "needing" context.) |
|
What I did was CUDA (RTX 2070), I see now that I forgot to mention this. Ok lets sync up the launch command, to make sure the issue is not there.
For f16 embeddings. I also checked with Also ran Which has even less difference between cpu and accelerator than your setup. |
|
I rand ppl on the full test split using the provided tq1_0 head-q4_k gguf on my cpu. Which falls in line with my other values.
|
|
Ok, would be great to have this done before the finished model releases. Would be great if someone else also can run ppl numbers and tests, because its kinda stuck right now. Also @AlexGabbia pls rebase. |
- load_arch_hparams: use n_ff_exp_arr + n_ff_exp() accessor (upstream
changed these from a scalar member during the rebase)
- sliding_window_pattern: get_arr, the pattern is mandatory for this arch
- partial_rotary_factor: read only from rope_parameters (base.py mirrors
the top-level key automatically)
- document why TOKEN_EMBD/OUTPUT are forced to F16 (they are the two
dense tensors in Maple, and the reference GGUFs ship them as F16)
- add @ModelBase.example("deepgrove/maple-preview")
get_arr for maple.attention.sliding_window_pattern requires an array, but the harness only emitted a per-layer array for the arches in its list, so test-llama-archs -a maple failed to load the model. Assisted-by: DeepSeek Harness
e935fb0 to
b92b8bf
Compare
|
@CISC @Green-Sky rebased on current master. All six review comments are in. On On One thing worth flagging: @Green-Sky the PPL gap was my wikitext file. Thanks for posting the hash, that was the missing piece. Yours is
That is 5.77 PPL from the file alone. Against your 86.96 on head-Q4_K: the Q4_K head costs me +1.39 over F16 on the same file, so on the upstream file I would expect about 86.58, a 0.38 residual. That is about your own CPU vs CUDA spread, so there is no kernel bug here. My mistake, and I will redo the rest of the table on the upstream file. Also correcting my earlier table: the row I marked CUDA for the TQ2_0 file was not actually running the ternary matmuls on the GPU. It took 20 min against 1 min 18 s for the real CUDA run. CUDA stays out of this PR, the description already says CPU-only. For the record I have the TQ2_0 dequantize path and the MMVQ vec dot fixed locally and will open that separately. |
The loader prefilled 7.0 and read the key optionally. The converter now writes it and the loader reads it as required, because llama-graph.cpp skips the clamp when the limit is 0 and an optional read would silently run unclamped. The test harness provides the key for the same reason. Also drops tensor_force_quant: base.py already forces FFN_GATE_INP to F32 and TOKEN_EMBD/OUTPUT to F16 for ternary file types. Assisted-by: DeepSeek Harness
|
@CISC both applied.
|
|
@Green-Sky here are the numbers on the upstream wikitext file, all four official GGUFs, CPU
Your tq1_0 head-Q4_K on the same file and the same backend is 87.3064 +/- 0.89495, so the like-for-like pair is 87.0225 against 87.3064. That is a 0.28 residual, the same order as the run-to-run spread I see here (the two F16 variants differ by 0.33 even though TQ1_0 and TQ2_0 represent the same ternary weights). For completeness, with the local CUDA TQ2_0 dequantize path, TQ2_0 head-F16 gives 85.1846 +/- 0.87378 on the GPU, 0.023 away from the CPU number on the same file. So the gap was the wikitext file, and it is closed. Thanks for pushing on it. |
|
@AlexGabbia glad we finally figured out what went wrong. :) |
|
@Green-Sky @CISC one small thing that would help, and it is not about this PR's scope. CUDA support for TQ2_0 is ready on a branch, draft PR #28769. The one gap I cannot close myself is AMD: HIP compiles in CI, but TQ2_0 is never exercised numerically there, because That is 17 cases and would settle it. Thanks. |
|
o.O |
Just needed to get CIs running, they were gone again... |
|
@AlexGabbia Check the typing error. |
ty flagged the stack() closure: it takes no argument, while LazyBase is annotated with func: Callable[[Any], Any]. Pass the tensor list through args instead of closing over it, the same way kimi_k3 does, so the callable shape matches. Assisted-by: DeepSeek Harness
|
ty was right: For the record on the other four failures: ubuntu and both gpu-webgpu jobs fail on artifact and model downloads, and their ctest runs report 100% passed. The openvino one is SWIGLU_CLAMP at ERR 3.07e-7 against a 1e-7 tolerance, which is inside the op itself - it is already on master, and this PR only adds LLM_ARCH_MAPLE to its arch list. |
| case LLM_ARCH_LAGUNA: | ||
| case LLM_ARCH_GRANITE_SWA: | ||
| case LLM_ARCH_DOTS3NOTE: // TODO: need to handle SWA pattern and MLA+SWA config | ||
| case LLM_ARCH_MAPLE: |
There was a problem hiding this comment.
Any reason to not implement the model saver for this model?
There was a problem hiding this comment.
Same reason as the others, the SWA pattern.
* gguf-py: add Maple tensor constants
Add MODEL_ARCH.MAPLE, its "maple" name, and the tensor list for the
Maple 20B-A1B ternary MoE architecture: token embeddings, output,
attention with Q/K RMS norms, and per-expert FFN tensors.
* convert: add Maple HF->GGUF converter
Register MapleForCausalLM in the HF architecture map and add the
converter for the Maple 20B-A1B ternary MoE model: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, partial rotary factor 0.5, and
per-expert weight stacking into merged 3D tensors.
* llama: add Maple architecture (20B-A1B ternary MoE)
Add the Maple 20B-A1B ternary MoE architecture: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, and ternary TQ1_0/TQ2_0
quantization support.
- register LLM_ARCH_MAPLE between MAMBA2 and JAMBA
- implement llama_model_maple: Q/K RMS norms after projection (GEMMA4
style), rope applied only on SWA layers (nope_on_global_attention),
ISWA KV cache, and MoE FFN with swiglu gate clamp at +7 (DEEPSEEK4
style)
- mark MAPLE as unsupported by the model saver (roundtrip skipped)
* tests: mark Maple as MoE-mandatory
Maple is always-MoE: the model throws when n_expert == 0, so the
test harness must only run the MoE config for LLM_ARCH_MAPLE.
* maple: apply review feedback (n_ff_exp_arr, get_arr, rope params)
- load_arch_hparams: use n_ff_exp_arr + n_ff_exp() accessor (upstream
changed these from a scalar member during the rebase)
- sliding_window_pattern: get_arr, the pattern is mandatory for this arch
- partial_rotary_factor: read only from rope_parameters (base.py mirrors
the top-level key automatically)
- document why TOKEN_EMBD/OUTPUT are forced to F16 (they are the two
dense tensors in Maple, and the reference GGUFs ship them as F16)
- add @ModelBase.example("deepgrove/maple-preview")
* tests: add Maple to the SWA pattern array list
get_arr for maple.attention.sliding_window_pattern requires an array, but
the harness only emitted a per-layer array for the arches in its list, so
test-llama-archs -a maple failed to load the model.
Assisted-by: DeepSeek Harness
* maple: move swiglu_clamp_exp to the converter
The loader prefilled 7.0 and read the key optionally. The converter now
writes it and the loader reads it as required, because llama-graph.cpp
skips the clamp when the limit is 0 and an optional read would silently
run unclamped. The test harness provides the key for the same reason.
Also drops tensor_force_quant: base.py already forces FFN_GATE_INP to F32
and TOKEN_EMBD/OUTPUT to F16 for ternary file types.
Assisted-by: DeepSeek Harness
* convert: fix the LazyBase func signature in the Maple converter
ty flagged the stack() closure: it takes no argument, while LazyBase is
annotated with func: Callable[[Any], Any]. Pass the tensor list through
args instead of closing over it, the same way kimi_k3 does, so the
callable shape matches.
Assisted-by: DeepSeek Harness
* gguf-py: add Maple tensor constants
Add MODEL_ARCH.MAPLE, its "maple" name, and the tensor list for the
Maple 20B-A1B ternary MoE architecture: token embeddings, output,
attention with Q/K RMS norms, and per-expert FFN tensors.
* convert: add Maple HF->GGUF converter
Register MapleForCausalLM in the HF architecture map and add the
converter for the Maple 20B-A1B ternary MoE model: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, partial rotary factor 0.5, and
per-expert weight stacking into merged 3D tensors.
* llama: add Maple architecture (20B-A1B ternary MoE)
Add the Maple 20B-A1B ternary MoE architecture: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, and ternary TQ1_0/TQ2_0
quantization support.
- register LLM_ARCH_MAPLE between MAMBA2 and JAMBA
- implement llama_model_maple: Q/K RMS norms after projection (GEMMA4
style), rope applied only on SWA layers (nope_on_global_attention),
ISWA KV cache, and MoE FFN with swiglu gate clamp at +7 (DEEPSEEK4
style)
- mark MAPLE as unsupported by the model saver (roundtrip skipped)
* tests: mark Maple as MoE-mandatory
Maple is always-MoE: the model throws when n_expert == 0, so the
test harness must only run the MoE config for LLM_ARCH_MAPLE.
* maple: apply review feedback (n_ff_exp_arr, get_arr, rope params)
- load_arch_hparams: use n_ff_exp_arr + n_ff_exp() accessor (upstream
changed these from a scalar member during the rebase)
- sliding_window_pattern: get_arr, the pattern is mandatory for this arch
- partial_rotary_factor: read only from rope_parameters (base.py mirrors
the top-level key automatically)
- document why TOKEN_EMBD/OUTPUT are forced to F16 (they are the two
dense tensors in Maple, and the reference GGUFs ship them as F16)
- add @ModelBase.example("deepgrove/maple-preview")
* tests: add Maple to the SWA pattern array list
get_arr for maple.attention.sliding_window_pattern requires an array, but
the harness only emitted a per-layer array for the arches in its list, so
test-llama-archs -a maple failed to load the model.
Assisted-by: DeepSeek Harness
* maple: move swiglu_clamp_exp to the converter
The loader prefilled 7.0 and read the key optionally. The converter now
writes it and the loader reads it as required, because llama-graph.cpp
skips the clamp when the limit is 0 and an optional read would silently
run unclamped. The test harness provides the key for the same reason.
Also drops tensor_force_quant: base.py already forces FFN_GATE_INP to F32
and TOKEN_EMBD/OUTPUT to F16 for ternary file types.
Assisted-by: DeepSeek Harness
* convert: fix the LazyBase func signature in the Maple converter
ty flagged the stack() closure: it takes no argument, while LazyBase is
annotated with func: Callable[[Any], Any]. Pass the tensor list through
args instead of closing over it, the same way kimi_k3 does, so the
callable shape matches.
Assisted-by: DeepSeek Harness
* gguf-py: add Maple tensor constants
Add MODEL_ARCH.MAPLE, its "maple" name, and the tensor list for the
Maple 20B-A1B ternary MoE architecture: token embeddings, output,
attention with Q/K RMS norms, and per-expert FFN tensors.
* convert: add Maple HF->GGUF converter
Register MapleForCausalLM in the HF architecture map and add the
converter for the Maple 20B-A1B ternary MoE model: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, partial rotary factor 0.5, and
per-expert weight stacking into merged 3D tensors.
* llama: add Maple architecture (20B-A1B ternary MoE)
Add the Maple 20B-A1B ternary MoE architecture: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, and ternary TQ1_0/TQ2_0
quantization support.
- register LLM_ARCH_MAPLE between MAMBA2 and JAMBA
- implement llama_model_maple: Q/K RMS norms after projection (GEMMA4
style), rope applied only on SWA layers (nope_on_global_attention),
ISWA KV cache, and MoE FFN with swiglu gate clamp at +7 (DEEPSEEK4
style)
- mark MAPLE as unsupported by the model saver (roundtrip skipped)
* tests: mark Maple as MoE-mandatory
Maple is always-MoE: the model throws when n_expert == 0, so the
test harness must only run the MoE config for LLM_ARCH_MAPLE.
* maple: apply review feedback (n_ff_exp_arr, get_arr, rope params)
- load_arch_hparams: use n_ff_exp_arr + n_ff_exp() accessor (upstream
changed these from a scalar member during the rebase)
- sliding_window_pattern: get_arr, the pattern is mandatory for this arch
- partial_rotary_factor: read only from rope_parameters (base.py mirrors
the top-level key automatically)
- document why TOKEN_EMBD/OUTPUT are forced to F16 (they are the two
dense tensors in Maple, and the reference GGUFs ship them as F16)
- add @ModelBase.example("deepgrove/maple-preview")
* tests: add Maple to the SWA pattern array list
get_arr for maple.attention.sliding_window_pattern requires an array, but
the harness only emitted a per-layer array for the arches in its list, so
test-llama-archs -a maple failed to load the model.
Assisted-by: DeepSeek Harness
* maple: move swiglu_clamp_exp to the converter
The loader prefilled 7.0 and read the key optionally. The converter now
writes it and the loader reads it as required, because llama-graph.cpp
skips the clamp when the limit is 0 and an optional read would silently
run unclamped. The test harness provides the key for the same reason.
Also drops tensor_force_quant: base.py already forces FFN_GATE_INP to F32
and TOKEN_EMBD/OUTPUT to F16 for ternary file types.
Assisted-by: DeepSeek Harness
* convert: fix the LazyBase func signature in the Maple converter
ty flagged the stack() closure: it takes no argument, while LazyBase is
annotated with func: Callable[[Any], Any]. Pass the tensor list through
args instead of closing over it, the same way kimi_k3 does, so the
callable shape matches.
Assisted-by: DeepSeek Harness
Overview
This adds the Maple 20B-A1B architecture to llama.cpp. Maple is DeepGrove's ternary MoE — 24 layers, 256 experts (8 active), SWA-512 interleaved with global attention at 3:1, ternary weights via TQ1_0/TQ2_0. Ported from
deepgrove-ai/llama.cpp(commit8ce8ca6c6d) with their go-ahead (see deepgrove-ai/llama.cpp#1).Split into 4 commits: gguf-py constants → converter → architecture → test entry.
CPU-only for now — TQ kernels already exist in ggml. Metal and CUDA can follow.
Additional information
test-llama-archs -a maplepasses (NMSE 8.75e-08)deepgrove/maple-previewTQ2_0 on M4 CPU: ~216 t/s prefill, ~88 t/s genRequirements