Skip to content

llama: add Maple 20B-A1B ternary MoE architecture (CPU) - #27000

Merged
ggerganov merged 8 commits into
ggml-org:masterfrom
AlexGabbia:feat/maple-support
Sep 14, 2026
Merged

ggerganov merged 8 commits into
ggml-org:masterfrom
AlexGabbia:feat/maple-support

Conversation

@AlexGabbia

@AlexGabbia AlexGabbia commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Overview

This adds the Maple 20B-A1B architecture to llama.cpp. Maple is DeepGrove's ternary MoE — 24 layers, 256 experts (8 active), SWA-512 interleaved with global attention at 3:1, ternary weights via TQ1_0/TQ2_0. Ported from deepgrove-ai/llama.cpp (commit 8ce8ca6c6d) with their go-ahead (see deepgrove-ai/llama.cpp#1).

Split into 4 commits: gguf-py constants → converter → architecture → test entry.

CPU-only for now — TQ kernels already exist in ggml. Metal and CUDA can follow.

Additional information

  • test-llama-archs -a maple passes (NMSE 8.75e-08)
  • Smoke tested with deepgrove/maple-preview TQ2_0 on M4 CPU: ~216 t/s prefill, ~88 t/s gen
  • DeepGrove confirmed group size 256 works fine for this model (row-wise scales)
  • Perplexity numbers being collected — will post results below

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES — code was generated by an AI agent under my direction, based on DeepGrove's original implementation. I reviewed and tested everything before submitting.

@github-actions github-actions Bot added model Model specific testing Everything test related conversion labels Aug 13, 2026
@ggml-gh-bot

This comment was marked as resolved.

@ggml-gh-bot ggml-gh-bot Bot added the draft PR will be changed to draft by github-actions bot label Aug 13, 2026
@Green-Sky

Green-Sky commented Aug 13, 2026 •

Copy link
Copy Markdown
Collaborator

I don't see ternary TQ1_0/TQ2_0 being applicable, since they are group size 256, while the model was trained on 128. We still don't have 128 group side ternary, so only q2_0 will work.

https://huggingface.co/stamsam/maple-preview-gguf/blob/main/maple-q2_0.gguf

@github-actions
github-actions Bot marked this pull request as draft August 13, 2026 09:59
@github-actions github-actions Bot removed the draft PR will be changed to draft by github-actions bot label Aug 13, 2026
@AlexGabbia

Copy link
Copy Markdown
Contributor Author

Thanks for the review. :) This PR supports the official DeepGrove GGUFs https://huggingface.co/deepgrove/maple-preview-GGUF which load and generate correctly on CPU — I tested the TQ2_0 file on an Apple M4: 216 t/s prefill, 88 t/s generation. The stamsam GGUFs are built on a non-mainline fork (PrismML), so they don't help mainline support. I agree a 128-group ternary would be a good improvement and I'd be happy to work on it as a follow-up.

@Green-Sky

Copy link
Copy Markdown
Collaborator

Please provide perplexity numbers, as suggested by the contributor guidelines.

@AlexGabbia

Copy link
Copy Markdown
Contributor Author

Thanks — working on it. The official DeepGrove GGUFs include both TQ1_0 and TQ2_0 variants (with F16 and Q4_K heads). Since Maple is natively ternary, the F16 GGUFs from DeepGrove are the same ternary model with F16 head/embeddings — not a full-precision baseline. I'll run perplexity comparing TQ2_0 vs TQ1_0 (both official DeepGrove GGUFs) and report the numbers here.

@Green-Sky

Copy link
Copy Markdown
Collaborator

Someone with big internet please check if the conversion from source bf16 works, ideally @AlexGabbia .

Some q8_0 ggufs for comparison would be nice too, pretty sure q8_0 or q4_0 would be lossless (for the ternary tensors).

@Green-Sky

Green-Sky commented Aug 13, 2026 •

Copy link
Copy Markdown
Collaborator

Ran a couple of ppl tests.

I used their TQ1_0-head-Q4_K as base and quantized it to q2_0 variants. tq1_0 should convert to q2_0 losslessly and q2_0 runs on cuda.

quant model size ppl512 (default) ppl2048
q2_0 head-Q4_K embd-f16 6060.33 MiB (2.51 BPW) 86.9194 +/- 0.89249 64.8277 +/- 0.68759
q2_0 head-q4_K embd-q8_0 5782.12 MiB (2.40 BPW) 86.9098 +/- 0.89197 64.8001 +/- 0.68666
q2_0 head-q4_K embd-q6_K 5710.26 MiB (2.37 BPW) 87.1487 +/- 0.89405 64.4658 +/- 0.68262

The q2_0 head-Q4_K embd-f16 gguf should be lossless compared to base.
Base size was 4747.45 MiB (1.97 BPW).

Maple uses swa over 512 tokens in the sparse attending layers, so 2048 should capture that.

Overall the ppl is suspiciously high, but that might be because of masked training, not sure.

@AlexGabbia

AlexGabbia commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor Author

Update with perplexity numbers — all 4 official DeepGrove GGUFs tested.

Setup: wikitext-2, CPU M4 (10-core), 20 chunks, 8 threads, -dev none (CPU-only as this PR is CPU-scoped).

Model Size PPL @512 PPL @2048
tq2_0 head-F16 embd-f16 6,054 MB 91.76 ± 5.15 56.65 ± 1.59
tq2_0 head-Q4_K embd-f16 5,902 MB 94.15 ± 5.31 57.81 ± 1.63
tq1_0 head-F16 embd-f16 5,431 MB 91.76 ± 5.15 56.65 ± 1.59
tq1_0 head-Q4_K embd-f16 4,984 MB 94.15 ± 5.31 57.81 ± 1.63

Conversion BF16 → GGUF F16: works — 291 tensors, 40.5 GB, using the converter in this PR.

Observations:

  • TQ1_0 and TQ2_0 produce identical PPL when using the same head type — expected, since both are packings of the same {-1,0,+1} ternary weights.
  • The @2048 PPL is significantly lower than @512, consistent with Maple's SWA-512 architecture (as you noted).
  • The head-Q4_K variant adds ~2.4 PPL points over head-F16, a small cost for ~150 MB size reduction.
  • All 4 official GGUFs use embd-f16 — the only difference between the F16 and Q4_K variants is the output/head tensor type.
  • My numbers are in the same range as yours (86.92 @512, 64.83 @2048) — the difference is likely chunk count and sampling.

On the Q8_0: your q2_0 head-Q4_K embd-q8_0 numbers already cover that comparison. The conversion from TQ1_0 to q2_0 is lossless for the ternary tensors as you said, and the embedding variants (F16/Q8_0/Q6_K) show negligible PPL difference — so I think that base is covered.

@Green-Sky

Green-Sky commented Aug 14, 2026 •

Copy link
Copy Markdown
Collaborator

Your numbers are way too different. I ran the full test split. What did you run and how many chunks?

edit:

-dev none (CPU-only as this PR is CPU-scoped).

Not really. If you rebase on master, you have access to metal tq2_0 btw.

@AlexGabbia

Copy link
Copy Markdown
Contributor Author

Thanks for the catch — you're right, my numbers aren't comparable to yours.

What I ran: only the first 20 chunks (10,240 tokens) of wikitext-2, -c 512 -b 512 -t 8, which is why the std is ±5.15 and the estimate is dominated by the noisy early chunks. Your full test-split numbers (86.92 @512) are the meaningful ones; mine just confirmed TQ1_0 ≡ TQ2_0 (same ternary weights, same PPL per-chunk).

A full test-split run takes several hours on this M4 CPU (10-core, 8 threads) — I've started it now (tq2_0 head-F16 and tq1_0 head-Q4_K, same command as yours, full file). I'll post the numbers as soon as it finishes.

On Metal: fair point — master has TQ2_0 Metal (#26980) and our branch is 19 commits behind. I'll rebase and re-run the smoke test on Metal as a follow-up.

@Green-Sky

Copy link
Copy Markdown
Collaborator

For fun and because the model is so fast, I ran it on the full train split.

model ppl 2048
q2_0 head-Q4_K embd-f16 68.4850 +/- 0.25039
q2_0 head-q4_K embd-q6_K 68.4664 +/- 0.25006

@Extirpater

Copy link
Copy Markdown

I don't see ternary TQ1_0/TQ2_0 being applicable, since they are group size 256, while the model was trained on 128. We still don't have 128 group side ternary, so only q2_0 will work.

https://huggingface.co/stamsam/maple-preview-gguf/blob/main/maple-q2_0.gguf

Thanks for your help on this PR @Green-Sky @AlexGabbia! Yeah I think any group size that is divisible by row dim will work as the model was trained with row wise scales. The default 256 group scales should work 👍 .

@Green-Sky

Green-Sky commented Aug 15, 2026 •

Copy link
Copy Markdown
Collaborator

Thanks for your help on this PR @Green-Sky @AlexGabbia! Yeah I think any group size that is divisible by row dim will work as the model was trained with row wise scales. The default 256 group scales should work 👍 .

Thats great, thanks for noting that.

Btw, have you considered adding more sparsity via things like ngram embeddings (longcat) or per-layer embeddings (gemma). Would be easy ways to add 5-10B without hurting inference speed much. Ideally even in ternary or binary. :)

@AlexGabbia

Copy link
Copy Markdown
Contributor Author

Ran the full wikitext-2 test split this time (not just 20 chunks). All 4 official DeepGrove GGUFs, CUDA on a RTX 5070 Ti, -ngl 99 -t 4. Branch rebased on latest master.

model @512 @2048
tq2_0 head-F16 79.76 ± 0.81 61.14 ± 0.65
tq1_0 head-F16 79.76 ± 0.81 61.14 ± 0.65
tq2_0 head-Q4_K 81.22 ± 0.83 61.84 ± 0.66
tq1_0 head-Q4_K 81.22 ± 0.83 61.84 ± 0.66

TQ1_0 and TQ2_0 give identical PPL — same weights, different packing, makes sense.

Compared to your q2_0 numbers @Green-Sky: tq2_0 scores lower (79.76 vs 86.92 @512, 61.14 vs 64.83 @2048). Probably because TQ2_0 keeps the original ternary weights as-is while q2_0 re-quantizes them.

head-Q4_K costs ~1.5 PPL @512, ~0.7 @2048 over head-F16.

@2048 is much lower than @512 across the board — Maple's SWA-512 needs the context.

@Green-Sky

Copy link
Copy Markdown
Collaborator

Compared to your q2_0 numbers @Green-Sky: tq2_0 scores lower (79.76 vs 86.92 @512, 61.14 vs 64.83 @2048). Probably because TQ2_0 keeps the original ternary weights as-is while q2_0 re-quantizes them.

Hm, I wonder whats wrong here, but the q2_0 weights themself are lossless. I made sure by requantizing them back to tq1_0 and the 2 files are hash identical.

54016e4d543bd688829e67103fc85b8396db94b7f8eb3f81fa95884e44393872  maple-preview-TQ1_0-head-Q4_K.gguf
quant to q2_0 ->
3bb9bd3e3755d06810b98f29a2edb0127199dc1b55a40a3c306cd69ac5615e1e  maple-preview-q2_0-head-q4_K.gguf
quant to tq1_0 ->
54016e4d543bd688829e67103fc85b8396db94b7f8eb3f81fa95884e44393872  maple-preview-requant_tq1_0.gguf

So the divergents needs to come from the ops running in different quantizaitons.
But not sure, the gap is rather large. Not sure different quant + different backend account for all that.

Maple's SWA-512 needs the context.

This reads wrong, slap your agent writing for you on the wrists.

@AlexGabbia

Copy link
Copy Markdown
Contributor Author

Fair point on the phrasing — corrected: the lower PPL @2048 vs @512 reflects the global-attention layers (every 4th layer) capturing long-range patterns that SWA-512 layers miss at short context. My agent earned the wrist-slap :D
On the divergence: you're right that the weights are lossless — I reproduced your hash check. So the gap (79.76 TQ2_0 vs 86.92 q2_0 @512) must come from the matmul ops running in different quantizations and/or backends (CUDA vs CPU).
To isolate the variable: I'll run TQ2_0 on CPU (same backend as you) and q2_0 on CUDA (same backend as me). That should tell us whether it's the backend or the quant type causing the gap. Will post results when ready.

@AlexGabbia

Copy link
Copy Markdown
Contributor Author

I reproduced your hash check — q2_0 → TQ1_0 round-trip is bit-identical. So the divergence isn't in the weights. To isolate the variable, I ran the same q2_0 file (head-Q4_K embd-f16, 6060.33 MiB / 2.51 BPW — matches your numbers exactly) across backends on the full wikitext-2 test split, same -c 512 -b 512:

Model Head Backend PPL @512 ±
TQ2_0 F16 CUDA (RTX 5070 Ti) 79.76 0.81
TQ2_0 F16 CPU x86 (AVX2/FMA) 79.88 0.81
q2_0 Q4_K CUDA (RTX 5070 Ti) 80.81 0.83
q2_0 Q4_K CPU x86 (AVX2/FMA) 80.78 0.82
q2_0 Q4_K CPU (Green-Sky) 86.92 0.89

On x86, q2_0 is consistent across backends (CUDA 80.81 vs CPU 80.78 — 0.03 gap). Same for TQ2_0. But your 86.92 is 6 PPL higher than my x86 CPU run on the same file.

The only variable left is CPU architecture. Are you on ARM/Apple Silicon? If so, the gap is in the q2_0 ARM kernel, not in the model or the quant type — and TQ2_0 native is the numerically stable path.

Happy to help debug the ARM kernel if you can share a repro. The architecture itself is simple enough that a standalone test should pinpoint it.


(Fixed the SWA phrasing btw — the @2048 improvement comes from the global-attention layers, not from SWA "needing" context.)

@Green-Sky

Copy link
Copy Markdown
Collaborator

What I did was CUDA (RTX 2070), I see now that I forgot to mention this.

Ok lets sync up the launch command, to make sure the issue is not there.

$ llama-perplexity -m models/maple-preview-q2_0-head-q4_K.gguf -f wikitext-2-raw/wiki.test.raw -c 512

For f16 embeddings.

I also checked with -b 512 and it results in 86.9592 +/- 0.89297.

Also ran $ test-llama-archs -a maple

main: using seed 36176602
|     Model arch.|                                Device|Config|   NMSE vs. CPU|Roundtrip|
|----------------|--------------------------------------|------|---------------|---------|
|           maple|               NVIDIA GeForce RTX 2070|   MoE|  OK (5.14e-13)|     SKIP|
|           maple|AMD Ryzen 9 PRO 3900 12-Core Processor|   MoE|  OK (0.00e+00)|     SKIP|
|           maple|                                  Meta|   MoE|  OK (5.14e-13)|     SKIP|

Which has even less difference between cpu and accelerator than your setup.

@Green-Sky

Copy link
Copy Markdown
Collaborator

I rand ppl on the full test split using the provided tq1_0 head-q4_k gguf on my cpu.
-f wikitext-2-raw/wiki.test.raw -c 512 -b 512
87.3064 +/- 0.89495

Which falls in line with my other values.

173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08 wikitext-2-raw/wiki.test.raw

@Green-Sky

Copy link
Copy Markdown
Collaborator

Ok, would be great to have this done before the finished model releases.

Would be great if someone else also can run ppl numbers and tests, because its kinda stuck right now.

Also @AlexGabbia pls rebase.

- load_arch_hparams: use n_ff_exp_arr + n_ff_exp() accessor (upstream
  changed these from a scalar member during the rebase)
- sliding_window_pattern: get_arr, the pattern is mandatory for this arch
- partial_rotary_factor: read only from rope_parameters (base.py mirrors
  the top-level key automatically)
- document why TOKEN_EMBD/OUTPUT are forced to F16 (they are the two
  dense tensors in Maple, and the reference GGUFs ship them as F16)
- add @ModelBase.example("deepgrove/maple-preview")
get_arr for maple.attention.sliding_window_pattern requires an array, but
the harness only emitted a per-layer array for the arches in its list, so
test-llama-archs -a maple failed to load the model.

Assisted-by: DeepSeek Harness
@AlexGabbia

Copy link
Copy Markdown
Contributor Author

@CISC @Green-Sky rebased on current master.

All six review comments are in. maple.py:14, maple.py:27, maple.cpp:7 and maple.cpp:10 are applied as suggested.

On maple.py:43: in the official TQ2_0 head-F16 GGUF the ternary part is 168 tensors, and TOKEN_EMBD / OUTPUT are the only dense tensors with real weight mass (2048 x 151936 each). Everything else dense is norms and the router, which is already F32. F16 on those two matches what DeepGrove ships. Happy to drop it to super() if you would rather have the default.

On maple.cpp:17: the converter does not write the key because config.json has no clamp value to convert, unlike deepseek or hy_v4 where it comes from config. So the fill is the fallback, and the official GGUFs carry maple.swiglu_clamp_exp = [7.0] * 24 anyway.

One thing worth flagging: maple.cpp:10 as suggested broke the arch test. The harness only emits a per-layer array for the arches in its list, so it failed with array key not found in model: maple.attention.sliding_window_pattern. I added LLM_ARCH_MAPLE to that list, same as dots3note which uses the same call. test-llama-archs -a maple passes now on the rebased branch, 8.99e-08 on GPU and Meta, 0.00e+00 on CPU, and the full suite passes too.

@Green-Sky the PPL gap was my wikitext file. Thanks for posting the hash, that was the missing piece. Yours is 173c87a5..., which is what scripts/get-wikitext-2.sh downloads. Mine has CRLF line endings and 2890 extra line breaks, so it tokenizes differently. Same model, same machine, same backend, only the file changed:

wikitext chunks PPL @512
upstream wiki.test.raw 584 85.1846 +/- 0.87378
mine 583 79.4119 +/- 0.81008

That is 5.77 PPL from the file alone. Against your 86.96 on head-Q4_K: the Q4_K head costs me +1.39 over F16 on the same file, so on the upstream file I would expect about 86.58, a 0.38 residual. That is about your own CPU vs CUDA spread, so there is no kernel bug here. My mistake, and I will redo the rest of the table on the upstream file.

Also correcting my earlier table: the row I marked CUDA for the TQ2_0 file was not actually running the ternary matmuls on the GPU. It took 20 min against 1 min 18 s for the real CUDA run.

CUDA stays out of this PR, the description already says CPU-only. For the record I have the TQ2_0 dequantize path and the MMVQ vec dot fixed locally and will open that separately.

Comment thread conversion/maple.py Outdated
Comment thread src/models/maple.cpp Outdated
The loader prefilled 7.0 and read the key optionally. The converter now
writes it and the loader reads it as required, because llama-graph.cpp
skips the clamp when the limit is 0 and an optional read would silently
run unclamped. The test harness provides the key for the same reason.

Also drops tensor_force_quant: base.py already forces FFN_GATE_INP to F32
and TOKEN_EMBD/OUTPUT to F16 for ternary file types.

Assisted-by: DeepSeek Harness
@AlexGabbia

Copy link
Copy Markdown
Contributor Author

@CISC both applied.

conversion/maple.py:49 - removed. You are right, base.py:1057 already forces TOKEN_EMBD and OUTPUT to F16 when the ftype is TQ1_0 or TQ2_0, and base.py:1023 already forces FFN_GATE_INP to F32, so the whole tensor_force_quant override was dead weight. The four official GGUFs match that: token_embd is F16 in all four, ffn_gate_inp is F32 in all four, and output is F16 or Q4_K depending on the variant.

src/models/maple.cpp:19 - moved to maple.py:set_gguf_parameters(), and the loader now just reads the key. I made the read required rather than optional on purpose: llama-graph.cpp:2227 skips the clamp entirely when the limit is 0, so an optional read would silently run unclamped and drift from the reference. That needs the test harness to supply the key, the same way it already does for hy_v4, deepseek4 and bailingmoe3, so I added a small block there. All four official GGUFs carry maple.swiglu_clamp_exp = [7.0] * 24, so nothing in the wild breaks. Say the word if you would rather have it optional.

test-llama-archs -a maple passes on the new head (8.98e-08 on GPU and Meta, 0.00e+00 on CPU), and an official GGUF still loads.

@AlexGabbia

Copy link
Copy Markdown
Contributor Author

@Green-Sky here are the numbers on the upstream wikitext file, all four official GGUFs, CPU -ngl 0 -t 16, -c 512 -b 512, full test split:

quant head PPL @512
TQ2_0 F16 85.2080 +/- 0.87385
TQ2_0 Q4_K 86.7029 +/- 0.88975
TQ1_0 F16 85.5340 +/- 0.87627
TQ1_0 Q4_K 87.0225 +/- 0.89211

Your tq1_0 head-Q4_K on the same file and the same backend is 87.3064 +/- 0.89495, so the like-for-like pair is 87.0225 against 87.3064. That is a 0.28 residual, the same order as the run-to-run spread I see here (the two F16 variants differ by 0.33 even though TQ1_0 and TQ2_0 represent the same ternary weights).

For completeness, with the local CUDA TQ2_0 dequantize path, TQ2_0 head-F16 gives 85.1846 +/- 0.87378 on the GPU, 0.023 away from the CPU number on the same file.

So the gap was the wikitext file, and it is closed. Thanks for pushing on it.

@AlexGabbia
AlexGabbia marked this pull request as ready for review September 11, 2026 15:49
@Green-Sky

Copy link
Copy Markdown
Collaborator

@AlexGabbia glad we finally figured out what went wrong. :)

@AlexGabbia

Copy link
Copy Markdown
Contributor Author

@Green-Sky @CISC one small thing that would help, and it is not about this PR's scope.

CUDA support for TQ2_0 is ready on a branch, draft PR #28769. The one gap I cannot close myself is AMD: HIP compiles in CI, but TQ2_0 is never exercised numerically there, because ci/run.sh runs test-backend-ops with -b CPU only and test-llama-archs uses F16 tensors. If anyone has an AMD GPU and a few minutes:

test-backend-ops -o MUL_MAT,MUL_MAT_ID -p TQ2_0

That is 17 cases and would settle it. Thanks.

@CISC CISC closed this Sep 11, 2026
@CISC CISC reopened this Sep 11, 2026
@Green-Sky

Copy link
Copy Markdown
Collaborator

o.O

@CISC

CISC commented Sep 11, 2026

Copy link
Copy Markdown
Member

o.O

Just needed to get CIs running, they were gone again...

@CISC

CISC commented Sep 11, 2026

Copy link
Copy Markdown
Member

@AlexGabbia Check the typing error.

ty flagged the stack() closure: it takes no argument, while LazyBase is
annotated with func: Callable[[Any], Any]. Pass the tensor list through
args instead of closing over it, the same way kimi_k3 does, so the
callable shape matches.

Assisted-by: DeepSeek Harness
@AlexGabbia

Copy link
Copy Markdown
Contributor Author

@CISC fixed in 0d0b53c.

ty was right: stack() took no argument while LazyBase is annotated with func: Callable[[Any], Any]. The tensor list now goes through args instead of the closure, the same way kimi_k3.py does it. I verified with ty 0.0.80 before pushing, and reverted the fix once to confirm the checker really does flag the original.

For the record on the other four failures: ubuntu and both gpu-webgpu jobs fail on artifact and model downloads, and their ctest runs report 100% passed. The openvino one is SWIGLU_CLAMP at ERR 3.07e-7 against a 1e-7 tolerance, which is inside the op itself - it is already on master, and this PR only adds LLM_ARCH_MAPLE to its arch list.

@CISC CISC added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 12, 2026
@ggerganov
ggerganov merged commit 3d10bcd into ggml-org:master Sep 14, 2026
27 of 29 checks passed
Comment thread src/llama-model-saver.cpp
case LLM_ARCH_LAGUNA:
case LLM_ARCH_GRANITE_SWA:
case LLM_ARCH_DOTS3NOTE: // TODO: need to handle SWA pattern and MLA+SWA config
case LLM_ARCH_MAPLE:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any reason to not implement the model saver for this model?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same reason as the others, the SWA pattern.

pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* gguf-py: add Maple tensor constants

Add MODEL_ARCH.MAPLE, its "maple" name, and the tensor list for the
Maple 20B-A1B ternary MoE architecture: token embeddings, output,
attention with Q/K RMS norms, and per-expert FFN tensors.

* convert: add Maple HF->GGUF converter

Register MapleForCausalLM in the HF architecture map and add the
converter for the Maple 20B-A1B ternary MoE model: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, partial rotary factor 0.5, and
per-expert weight stacking into merged 3D tensors.

* llama: add Maple architecture (20B-A1B ternary MoE)

Add the Maple 20B-A1B ternary MoE architecture: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, and ternary TQ1_0/TQ2_0
quantization support.

- register LLM_ARCH_MAPLE between MAMBA2 and JAMBA
- implement llama_model_maple: Q/K RMS norms after projection (GEMMA4
  style), rope applied only on SWA layers (nope_on_global_attention),
  ISWA KV cache, and MoE FFN with swiglu gate clamp at +7 (DEEPSEEK4
  style)
- mark MAPLE as unsupported by the model saver (roundtrip skipped)

* tests: mark Maple as MoE-mandatory

Maple is always-MoE: the model throws when n_expert == 0, so the
test harness must only run the MoE config for LLM_ARCH_MAPLE.

* maple: apply review feedback (n_ff_exp_arr, get_arr, rope params)

- load_arch_hparams: use n_ff_exp_arr + n_ff_exp() accessor (upstream
  changed these from a scalar member during the rebase)
- sliding_window_pattern: get_arr, the pattern is mandatory for this arch
- partial_rotary_factor: read only from rope_parameters (base.py mirrors
  the top-level key automatically)
- document why TOKEN_EMBD/OUTPUT are forced to F16 (they are the two
  dense tensors in Maple, and the reference GGUFs ship them as F16)
- add @ModelBase.example("deepgrove/maple-preview")

* tests: add Maple to the SWA pattern array list

get_arr for maple.attention.sliding_window_pattern requires an array, but
the harness only emitted a per-layer array for the arches in its list, so
test-llama-archs -a maple failed to load the model.

Assisted-by: DeepSeek Harness

* maple: move swiglu_clamp_exp to the converter

The loader prefilled 7.0 and read the key optionally. The converter now
writes it and the loader reads it as required, because llama-graph.cpp
skips the clamp when the limit is 0 and an optional read would silently
run unclamped. The test harness provides the key for the same reason.

Also drops tensor_force_quant: base.py already forces FFN_GATE_INP to F32
and TOKEN_EMBD/OUTPUT to F16 for ternary file types.

Assisted-by: DeepSeek Harness

* convert: fix the LazyBase func signature in the Maple converter

ty flagged the stack() closure: it takes no argument, while LazyBase is
annotated with func: Callable[[Any], Any]. Pass the tensor list through
args instead of closing over it, the same way kimi_k3 does, so the
callable shape matches.

Assisted-by: DeepSeek Harness
quimmedes pushed a commit to quimmedes/cafe-llama.cpp that referenced this pull request Sep 16, 2026
* gguf-py: add Maple tensor constants

Add MODEL_ARCH.MAPLE, its "maple" name, and the tensor list for the
Maple 20B-A1B ternary MoE architecture: token embeddings, output,
attention with Q/K RMS norms, and per-expert FFN tensors.

* convert: add Maple HF->GGUF converter

Register MapleForCausalLM in the HF architecture map and add the
converter for the Maple 20B-A1B ternary MoE model: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, partial rotary factor 0.5, and
per-expert weight stacking into merged 3D tensors.

* llama: add Maple architecture (20B-A1B ternary MoE)

Add the Maple 20B-A1B ternary MoE architecture: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, and ternary TQ1_0/TQ2_0
quantization support.

- register LLM_ARCH_MAPLE between MAMBA2 and JAMBA
- implement llama_model_maple: Q/K RMS norms after projection (GEMMA4
  style), rope applied only on SWA layers (nope_on_global_attention),
  ISWA KV cache, and MoE FFN with swiglu gate clamp at +7 (DEEPSEEK4
  style)
- mark MAPLE as unsupported by the model saver (roundtrip skipped)

* tests: mark Maple as MoE-mandatory

Maple is always-MoE: the model throws when n_expert == 0, so the
test harness must only run the MoE config for LLM_ARCH_MAPLE.

* maple: apply review feedback (n_ff_exp_arr, get_arr, rope params)

- load_arch_hparams: use n_ff_exp_arr + n_ff_exp() accessor (upstream
  changed these from a scalar member during the rebase)
- sliding_window_pattern: get_arr, the pattern is mandatory for this arch
- partial_rotary_factor: read only from rope_parameters (base.py mirrors
  the top-level key automatically)
- document why TOKEN_EMBD/OUTPUT are forced to F16 (they are the two
  dense tensors in Maple, and the reference GGUFs ship them as F16)
- add @ModelBase.example("deepgrove/maple-preview")

* tests: add Maple to the SWA pattern array list

get_arr for maple.attention.sliding_window_pattern requires an array, but
the harness only emitted a per-layer array for the arches in its list, so
test-llama-archs -a maple failed to load the model.

Assisted-by: DeepSeek Harness

* maple: move swiglu_clamp_exp to the converter

The loader prefilled 7.0 and read the key optionally. The converter now
writes it and the loader reads it as required, because llama-graph.cpp
skips the clamp when the limit is 0 and an optional read would silently
run unclamped. The test harness provides the key for the same reason.

Also drops tensor_force_quant: base.py already forces FFN_GATE_INP to F32
and TOKEN_EMBD/OUTPUT to F16 for ternary file types.

Assisted-by: DeepSeek Harness

* convert: fix the LazyBase func signature in the Maple converter

ty flagged the stack() closure: it takes no argument, while LazyBase is
annotated with func: Callable[[Any], Any]. Pass the tensor list through
args instead of closing over it, the same way kimi_k3 does, so the
callable shape matches.

Assisted-by: DeepSeek Harness
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
* gguf-py: add Maple tensor constants

Add MODEL_ARCH.MAPLE, its "maple" name, and the tensor list for the
Maple 20B-A1B ternary MoE architecture: token embeddings, output,
attention with Q/K RMS norms, and per-expert FFN tensors.

* convert: add Maple HF->GGUF converter

Register MapleForCausalLM in the HF architecture map and add the
converter for the Maple 20B-A1B ternary MoE model: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, partial rotary factor 0.5, and
per-expert weight stacking into merged 3D tensors.

* llama: add Maple architecture (20B-A1B ternary MoE)

Add the Maple 20B-A1B ternary MoE architecture: 24 layers, 256
experts with 8 active, sliding-window attention (SWA-512) interleaved
with global attention at a 3:1 ratio, and ternary TQ1_0/TQ2_0
quantization support.

- register LLM_ARCH_MAPLE between MAMBA2 and JAMBA
- implement llama_model_maple: Q/K RMS norms after projection (GEMMA4
  style), rope applied only on SWA layers (nope_on_global_attention),
  ISWA KV cache, and MoE FFN with swiglu gate clamp at +7 (DEEPSEEK4
  style)
- mark MAPLE as unsupported by the model saver (roundtrip skipped)

* tests: mark Maple as MoE-mandatory

Maple is always-MoE: the model throws when n_expert == 0, so the
test harness must only run the MoE config for LLM_ARCH_MAPLE.

* maple: apply review feedback (n_ff_exp_arr, get_arr, rope params)

- load_arch_hparams: use n_ff_exp_arr + n_ff_exp() accessor (upstream
  changed these from a scalar member during the rebase)
- sliding_window_pattern: get_arr, the pattern is mandatory for this arch
- partial_rotary_factor: read only from rope_parameters (base.py mirrors
  the top-level key automatically)
- document why TOKEN_EMBD/OUTPUT are forced to F16 (they are the two
  dense tensors in Maple, and the reference GGUFs ship them as F16)
- add @ModelBase.example("deepgrove/maple-preview")

* tests: add Maple to the SWA pattern array list

get_arr for maple.attention.sliding_window_pattern requires an array, but
the harness only emitted a per-layer array for the arches in its list, so
test-llama-archs -a maple failed to load the model.

Assisted-by: DeepSeek Harness

* maple: move swiglu_clamp_exp to the converter

The loader prefilled 7.0 and read the key optionally. The converter now
writes it and the loader reads it as required, because llama-graph.cpp
skips the clamp when the limit is 0 and an optional read would silently
run unclamped. The test harness provides the key for the same reason.

Also drops tensor_force_quant: base.py already forces FFN_GATE_INP to F32
and TOKEN_EMBD/OUTPUT to F16 for ternary file types.

Assisted-by: DeepSeek Harness

* convert: fix the LazyBase func signature in the Maple converter

ty flagged the stack() closure: it takes no argument, while LazyBase is
annotated with func: Callable[[Any], Any]. Pass the tensor list through
args instead of closing over it, the same way kimi_k3 does, so the
callable shape matches.

Assisted-by: DeepSeek Harness
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. model Model specific testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants