Skip to content

feat: add GLM5Next MTP, optimize - #29928

Merged
pwilkin merged 7 commits into
ggml-org:masterfrom
pwilkin:glm5nextmtp
Oct 7, 2026
Merged

pwilkin merged 7 commits into
ggml-org:masterfrom
pwilkin:glm5nextmtp

Conversation

@pwilkin

@pwilkin pwilkin commented Oct 3, 2026

Copy link
Copy Markdown
Member

Overview

Implement MTP for GLM5Next, optimize the graph

Additional information

Supersedes #27917

Requirements

@pwilkin
pwilkin requested review from CISC and ggerganov as code owners October 3, 2026 22:28
@github-actions github-actions Bot added the model Model specific label Oct 3, 2026
Comment thread src/models/glm5-next.cpp Outdated
cur = build_norm(cur, layer.attn_norm, nullptr, LLM_NORM_RMS, il);
cb(cur, "mtp_attn_norm", il);

const bool headless = n_outputs == 0 && !(cparams.embeddings_nextn && !cparams.embeddings_nextn_masked);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what's this headless stuff? can you keep it consistent with other MTP?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You mean implement headless for other viable MTP models too? :)

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

headless skips the attention when you don't have any outputs to write, i.e. you're just populating the KV cache

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah sure, but make sure it doesn't cause any random graph re-allocations or graph reuse issues.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think due to previous MTP graph caching shenanigans n_outputs is already included in the graph hash code input params, so it should be safe, but ofc I'll check.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

headless skips the attention

You mean skips the FFN, right? The attention should not be skipped even for empty outputs.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On second thought: I would recommend not adding this in this PR rather in it's own PR, so that we can measure performance and any graph issues.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, that's the clever part - you actually skip attention, i.e. at least half of the attention, since you don't need the attention output for anything - you just run the part that's needed to compute the KV cache parameters.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On second thought: I would recommend not adding this in this PR rather in it's own PR, so that we can measure performance and any graph issues.

Yes, keep it simple. Currently all sparse attention graphs have some problem because they reallocate (#29867 (comment)). We need to fix this first before adding more logic.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aight, will do a standard "skip FFN" behavior and add the headless mode in a separate PR.

Comment on lines +202 to +204
if (cell.seq_id.size() > 1) {
return false;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In which cases does this happen?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn't - it's a defensive check added to surface potential silent data corruption. I can remove it or add an extra warning message if you prefer.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(but IMO this is worth keeping because common_context_seq_rm is part of the API, so the function should actually refuse to do this instead of silently corrupting the data)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It should be GGML_ASSERT I think. I don't see how we can end up with more than one sequence in a recurrent cell.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@ggerganov it happens every time on llama_memory_recurrent::seq_cp, there's no copy-on-write behavior, so the two sequences end up sharing the cell.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In "real life" this is harmless, because once any decode happens the sequences disconnect from each other, but it's obtainable from a legal sequence of API calls.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, let's keep it like this for now. Basically the condition is going to occur if one tries to rollback a sequence that was just copied - we will deny that.

The recurrent state rollback logic in general needs to be asserted anyway since there are various failure cases that currently fail silently (f.ex, trying to rollback after a decode batch). It's something in the TODO list already.

@NenadSteric

NenadSteric commented Oct 4, 2026 •

Copy link
Copy Markdown

Here is performance data from my local server - i hope it is useful
With this setup :

item value
code PR head 1e24094d7 vs master 836d57176 (the PR's merge base)
model avar6/GLM-5.3-Flash-BF16-gguf on Hugging Face, revision 7083b208 folder Q2_K-Q3_K-3.20bpw (3 shards, 119.5 GiB; routed experts gate/up Q2_K, down Q3_K)
hardware 6x NVIDIA RTX 3090 (24 GB each), PCIe 4.0 x16, no NVLink; driver 615.71.09;200 W, no clock locks
build (both) CUDA 13.2, GCC 15; cmake -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86 -DGGML_NATIVE=ON -DBUILD_SHARED_LIBS=ON -DGGML_CUDA_CCCL_VERSION=v3.4.3 (CCCL 3.4.3 as in upstream)
benchmark llama-bench -m <shard 1 of 3> -ngl 99 -sm layer -ts 9/7/7/7/7/9 -fa 1 -p 0 -n 128 -d 0,8192,32768 -r 3
environment GGML_CUDA_P2P unset (P2P off, the default); no other GGML_/LLAMA_ variables

Result: #29928 vs base

two runs per build, each -r 3, alternating master/PR/master/PR

context already in the cache master #29928 change
0 tokens 37.25 tok/s 37.24 tok/s none (path not used)
8,192 tokens 30.78 / 30.81 tok/s 36.14 / 36.19 tok/s +17.4 / +17.5%
32,768 tokens 30.40 / 30.37 tok/s 35.50 / 35.51 tok/s +16.8 / +16.9%

the tests were run and the tables were created by Claude Opus
speedup comes from sparse-attention commit (1df2e15), llama-bench does not use MTP (afaiu)

@pwilkin

pwilkin commented Oct 5, 2026

Copy link
Copy Markdown
Member Author

@ggerganov @am17an all right, removed the headless code, implemented the standard MTP strategy for now.

Comment thread src/models/glm5-next.cpp Outdated
Comment on lines +339 to +342

if (inp->gather) {
inp->gather_mask_h = ggml_repeat_4d(ctx0, inp->gather_mask, n_sel, n_head, 1, n_tokens);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rebase on latest master because this logic has changed.

@pwilkin

pwilkin commented Oct 5, 2026

Copy link
Copy Markdown
Member Author

@ggerganov done, also exposed a bug during testing (see #30020)

@ServeurpersoCom

ServeurpersoCom commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

#30042 removes the gather path (merge order is #30017, then #30042, then this one), so after rebasing you can drop deacfbf and the gather_mask_h hunk entirely. Your fix does make the gather about 5% faster than sparse FA at -ub 16, I benchmarked it, but since -ub 16 has no real use case and would stay untested, we're keeping a single sparse path. Also note gather_mask becomes sel_mask.

EDIT:

@pwilkin, I tested the rebased version on my side, with a small extra commit moving the MTP graph to the new nextn crop helpers: on GLM-5.3-Flash it doubles decode at short context (97% acceptance) and gives +52% at 32k, with no reallocation. Want me to push it to your branch?

@pwilkin

pwilkin commented Oct 6, 2026

Copy link
Copy Markdown
Member Author

@ServeurpersoCom yeah, do it :)

saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 6, 2026
An MTP head loaded from its own file joins the expert cache as a member
with its own file backing, so the finite-L2 and prefill-swap gates can
accept it like a joined draft model.

With --expert-cache-draft off, an MTP head (-md) with every layer on the GPU
and no draft tensor overrides reads no host or file expert, so the finite
tier and the prefill swap never see its work. The gate accepts it as well.
In the joint cache the head competed with the target for L1 and landed
mostly in the file (DS4F roleplay, L1 16 GiB, L2 60 GiB: 16 of 256 experts
in L1, MTP n1 -7.0 % decode); kept on the GPU it gives +14.1 % over no
speculation.

An MTP context on the target's own weights (draft-mtp without -md)
passes as well: its NextN block is one more routed layer of the target's
cache (GLM5-Next with upstream PR ggml-org#29928, finite L2, one slot).

test-expert-config covers both gates.
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 6, 2026
docs/ranma/exl3.md: the EXL3 types and the load check, the conversion (MTP layers follow the model class;
a model whose MTP layers cannot be stored as EXL3, such as Qwen3.8-Flash-Next, is converted without them
after an error message), the backends, the models checked (llama, qwen3moe, lfm2moe, qwen4exp, deepseek4,
glm5-next; loading is not limited to them), the MTP layer of glm5-next through the NextN graph of
upstream PR ggml-org#29928 with --spec-type draft-mtp and what was checked with it, the precision of the prompt
GEMM and the F32 request, the RDNA3 status (compiles, GEMV on the GPU, the WMMA GEMM is RDNA4 only,
untested) and GGML_CUDA_HC_GATED_FUSION=0. README: one section with both lines, and the ExLlamaV3 credit
next to the license.
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 6, 2026
The README now leads with what the fork is for and its three main features, and leaves the full
list of changes to a page of its own.

- What this fork is: also one person's own development work in a single session. The expert cache
  installs a plan only when every slot is idle, so sessions that send requests in turn leave it few
  chances; this is said before the advice to use upstream for many users.
- Main features: the expert cache, the smart MTP draft length and EXL3 weights, a few lines each
  with their pages. EXL3 needs GGUF files in the format of this fork, from the Hugging Face
  repositories or, for another EXL3 model, the converter, with the note that not every model is
  guaranteed to work yet; those files do not open in other tools.
- The expert cache entry, expert-cache.md and expert-cache-l2.md say to give the slot count with
  -np: without it the server runs four slots, the per-slot state grows with them, and a finite
  host tier is refused.
- Benchmarks becomes a section of its own; the numbers come with the benchmark commit, the last
  of the series.
- Roadmap: the porting paragraph is gone; the section lists plans that are not done yet:
  DeepSeek V4 Flash performance work, DeepSeek V4.1 Flash with EXL3, and experimental RDNA3 and CUDA
  builds. GLM 5.3 is no longer listed: upstream has merged it, and its EXL3 files convert and run.
- Target environment: other environments may work but are not guaranteed, as in CONTRIBUTING.md,
  with a table of the main features on RDNA4 and RDNA3 that tells maintainer tests, user reports
  and compile checks apart.
- The list of changes over upstream moves to docs/ranma/README.md, with its links made relative to
  that directory, together with the rule of one line and one page per change. It gains glm5-next in the
  EXL3 line and a GLM-5.3 Flash section for LoRA on the attention output, described in
  graph-runtime.md, and for the MTP draft of upstream PR ggml-org#29928, carried as its six commits unchanged.
- deepseek-v4.md: the selected-cell attention is a HIP path measured on RDNA4 only, not an
  RDNA4-only path; its gather has no RDNA4 condition.
Build the GLM5-Next multi-token-prediction head as graph_mtp: the NextN block
embeds enorm(tok)+hnorm(h) through eh_proj, runs one plain DSA layer and the
shared lm_head, reusing the trunk's builders through the no_build tag ctor.
llama_memory_recurrent also tolerates a partial seq_rm when the context holds
no recurrent layers, which is what the MTP draft context needs.

Assisted-by: Claude
A NextN forward with no output rows (the MTP catch-up and the draft-context
prefill) persists only through its cache writes, so the headless graph keeps
the MLA latent, indexer key|gate and pooled-key writes and drops the query
path, the indexer selection, the attention body, the FFN and the LM head. The
4-token catch-up falls from 6.9 ms to 0.33 ms of kernels; the greedy output
hashes and the draft acceptance are unchanged.

Assisted-by: Claude
…lback

Three fixes from the architectural review. The headless graph prune now also
requires that no unmasked nextn extraction is live, because that mode reads
n_tokens hidden rows regardless of the logits flags. Masked extraction
publishes the hidden rows gathered by the output ids, so a batch whose output
flags are not a prefix exports the right rows. A partial recurrent rollback
whose tail cell is shared with another sequence is now rejected instead of
silently moving that sequence's tail.

Assisted-by: Claude
…runing it

Replace the headless NextN prune with the crop pattern the other MTP
graphs use: gather the attention output and the block input at the
output ids before the position-wise FFN and the shared head. A NextN
forward with no output rows (the MTP catch-up and the draft-context
prefill) then runs the FFN and the head over zero rows. The 4-token
catch-up falls from 6.9 ms to 2.9 ms of kernels; greedy output hashes
are unchanged.

Assisted-by: Claude
@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Rebase only

Replace the local crop condition and the masked select of t_h_nextn
with crop_before_nextn and crop_after_nextn, so the MTP graph narrows
its rows the same way as the main graph and the other models.

Describe the shared cell and empty filter branches of the recurrent
partial rollback.
@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Compared against unsloth's MTP implementation on the same GGUF (GLM-5.3-Flash UD-IQ1_S, greedy): acceptance is in the same range with no systematic gap (n=2: 99/71/75/81/78% vs 99/78/74/78/74% on 5 prompts up to 32k), and on the one prompt where both base models agree, all configs of both implementations produce the exact same output. Elsewhere the base trunks themselves drift at near ties on this 1.5 bpw quant, so a token level comparison isn't possible.

@pwilkin

pwilkin commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

Okay, so I guess this is gtg?

@ggerganov

Copy link
Copy Markdown
Member

Doing a quick test on my end.

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Yes from my side, it just needs a second approval now. I'm playing with it on my RTX PRO 6000, but even the 1-bit quant is a tight fit: with MTP on, the draft context eats enough VRAM that I drop from 128K to 48K of context.

@ggerganov

Copy link
Copy Markdown
Member

Hm, trying to run the Q2_K model + Q4_0 mtp from here: https://huggingface.co/ggml-org/GLM-5.3-Flash-GGUF/tree/main

I get this error:

[57360] 0.58.053.836 I load_tensors: loading model tensors, this can take a while... (load_mode = mmap)
[57360] 0.58.066.843 E llama_model_load: error loading model: check_tensor_dims: tensor 'blk.0.attn_norm.weight' not found
[57360] 0.58.066.847 E llama_model_load_from_file_impl: failed to load model
[57360] 0.58.066.848 E common_speculative_init_result: failed to load draft model, '.cache/huggingface/hub/models--ggml-org--GLM-5.3-Flash-GGUF/snapshots/14b83f6336796fce6b282f1b5564fae290adb177/mtp-GLM-5.3-Flash-Q4_0.gguf'
[57360] 0.58.066.851 E srv    load_model: failed to load draft model, '.cache/huggingface/hub/models--ggml-org--GLM-5.3-Flash-GGUF/snapshots/14b83f6336796fce6b282f1b5564fae290adb177/mtp-GLM-5.3-Flash-Q4_0.gguf'
[57360] 0.58.066.855 I srv    operator(): operator(): cleaning up before exit...
[57360] 0.58.067.409 E srv  llama_server: exiting due to model loading error

@ggerganov

Copy link
Copy Markdown
Member

This is the command that I am trying:

./bin/llama-server -hf ggml-org/GLM-5.3-Flash-GGUF:Q2_K --spec-type draft-mtp

@pwilkin

pwilkin commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

@ggerganov testing on the quant I use now

@ServeurpersoCom

ServeurpersoCom commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

That's the separate MTP GGUF case: this PR only supports the NextN layer embedded in the main GGUF, and glm5-next doesn't have the mtp_only loading path yet (there's a TODO for it in load_arch_tensors). Unlike glm4-moe and qwen35moe, it's not just making the trunk tensors optional: an MTP-only file also needs its own decoder graph fed by the target hidden state, which is the DECODER_MTP follow-up mentioned in that TODO. -> I check

I tested with the unsloth UD-IQ1_S quant, which embeds the NextN layer, with the first shard replaced by the one from their Shard_Rewrite folder (the original uses the old glm5next arch name):

llama-server -m GLM-5.3-Flash-UD-IQ1_S-00001-of-00003.gguf -ngl 99 -fa on -c 40960 -np 1 --spec-type draft-mtp --spec-draft-n-max 2

@ggerganov

ggerganov commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Looking at src/models/mimo2.cpp as a reference, I was able to make it work with this patch:

Details
diff --git a/src/models/glm5-next.cpp b/src/models/glm5-next.cpp
index 440f90dce2..1537a2a63a 100644
--- a/src/models/glm5-next.cpp
+++ b/src/models/glm5-next.cpp
@@ -64,9 +64,12 @@ void llama_model_glm5_next::load_arch_tensors(llama_model_loader & ml) {
     const int64_t hc         = hparams.dsv4_hc_mult;
     const int64_t hc_mix_dim = (2 + hc)*hc;
 
-    // the NextN block is loaded but only used by the MTP graph.
-    // Separated trunk_only/mtp_only handling TODO with DECODER_MTP graph in the MTP follow up
-    int mtp_flags = 0;
+    const bool mtp_only = (hparams.n_layer_nextn > 0) && (ml.get_weight("blk.0.attn_norm.weight") == nullptr);
+    const std::string mtp_probe = "blk." + std::to_string(n_layer) + ".nextn.eh_proj.weight";
+    const bool trunk_only = (hparams.n_layer_nextn > 0) && (ml.get_weight(mtp_probe.c_str()) == nullptr);
+    const int trunk_flags = mtp_only ? TENSOR_NOT_REQUIRED : 0;
+    int mtp_flags         = trunk_only ? TENSOR_NOT_REQUIRED : 0;
+
     if (!ml.load_mtp) {
         mtp_flags |= TENSOR_SKIP;
     }
@@ -79,18 +82,19 @@ void llama_model_glm5_next::load_arch_tensors(llama_model_loader & ml) {
     for (int i = 0; i < n_layer_all; ++i) {
         auto & layer = layers[i];
 
-        const int flags = (i >= n_layer) ? mtp_flags : 0;
+        const bool is_nextn = i >= n_layer;
+        const int flags = is_nextn ? mtp_flags : trunk_flags;
 
         layer.attn_norm = create_tensor(tn(LLM_TENSOR_ATTN_NORM, "weight", i), {n_embd}, flags);
         layer.ffn_norm  = create_tensor(tn(LLM_TENSOR_FFN_NORM,  "weight", i), {n_embd}, flags);
 
         if (i < n_layer) {
-            layer.hc_attn_fn    = create_tensor(tn(LLM_TENSOR_HC_ATTN_FN,    "weight", i), {hc*n_embd, hc_mix_dim}, 0);
-            layer.hc_attn_base  = create_tensor(tn(LLM_TENSOR_HC_ATTN_BASE,  "weight", i), {hc_mix_dim}, 0);
-            layer.hc_attn_scale = create_tensor(tn(LLM_TENSOR_HC_ATTN_SCALE, "weight", i), {3}, 0);
-            layer.hc_ffn_fn     = create_tensor(tn(LLM_TENSOR_HC_FFN_FN,     "weight", i), {hc*n_embd, hc_mix_dim}, 0);
-            layer.hc_ffn_base   = create_tensor(tn(LLM_TENSOR_HC_FFN_BASE,   "weight", i), {hc_mix_dim}, 0);
-            layer.hc_ffn_scale  = create_tensor(tn(LLM_TENSOR_HC_FFN_SCALE,  "weight", i), {3}, 0);
+            layer.hc_attn_fn    = create_tensor(tn(LLM_TENSOR_HC_ATTN_FN,    "weight", i), {hc*n_embd, hc_mix_dim}, flags);
+            layer.hc_attn_base  = create_tensor(tn(LLM_TENSOR_HC_ATTN_BASE,  "weight", i), {hc_mix_dim}, flags);
+            layer.hc_attn_scale = create_tensor(tn(LLM_TENSOR_HC_ATTN_SCALE, "weight", i), {3}, flags);
+            layer.hc_ffn_fn     = create_tensor(tn(LLM_TENSOR_HC_FFN_FN,     "weight", i), {hc*n_embd, hc_mix_dim}, flags);
+            layer.hc_ffn_base   = create_tensor(tn(LLM_TENSOR_HC_FFN_BASE,   "weight", i), {hc_mix_dim}, flags);
+            layer.hc_ffn_scale  = create_tensor(tn(LLM_TENSOR_HC_FFN_SCALE,  "weight", i), {3}, flags);
         }
 
         const int64_t head_dim = hparams.n_embd_head_kda;
@@ -100,25 +104,25 @@ void llama_model_glm5_next::load_arch_tensors(llama_model_loader & ml) {
         if (hparams.is_recr(i)) {
             auto conv = [&](llm_tensor tid) {
                 ggml_tensor * t = create_tensor(tn(tid, "weight", i), {d_conv, 1, d_inner, 1}, TENSOR_NOT_REQUIRED);
-                return t ? t : create_tensor(tn(tid, "weight", i), {d_conv, 1, d_inner}, 0);
+                return t ? t : create_tensor(tn(tid, "weight", i), {d_conv, 1, d_inner}, flags);
             };
             layer.ssm_q_conv = conv(LLM_TENSOR_SSM_CONV1D_Q);
             layer.ssm_k_conv = conv(LLM_TENSOR_SSM_CONV1D_K);
             layer.ssm_v_conv = conv(LLM_TENSOR_SSM_CONV1D_V);
 
-            create_tensor_qkv(layer, i, n_embd, d_inner, d_inner, d_inner, 0);
+            create_tensor_qkv(layer, i, n_embd, d_inner, d_inner, d_inner, flags);
 
-            layer.ssm_f_a  = create_tensor(tn(LLM_TENSOR_SSM_F_A,  "weight", i), {n_embd, head_dim}, 0);
-            layer.ssm_f_b  = create_tensor(tn(LLM_TENSOR_SSM_F_B,  "weight", i), {head_dim, d_inner}, 0);
-            layer.ssm_beta = create_tensor(tn(LLM_TENSOR_SSM_BETA, "weight", i), {n_embd, n_head}, 0);
+            layer.ssm_f_a  = create_tensor(tn(LLM_TENSOR_SSM_F_A,  "weight", i), {n_embd, head_dim}, flags);
+            layer.ssm_f_b  = create_tensor(tn(LLM_TENSOR_SSM_F_B,  "weight", i), {head_dim, d_inner}, flags);
+            layer.ssm_beta = create_tensor(tn(LLM_TENSOR_SSM_BETA, "weight", i), {n_embd, n_head}, flags);
 
-            layer.ssm_a    = create_tensor(tn(LLM_TENSOR_SSM_A_NOSCAN, i), {n_head}, 0);
-            layer.ssm_dt_b = create_tensor(tn(LLM_TENSOR_SSM_DT, "bias", i), {d_inner}, 0);
+            layer.ssm_a    = create_tensor(tn(LLM_TENSOR_SSM_A_NOSCAN, i), {n_head}, flags);
+            layer.ssm_dt_b = create_tensor(tn(LLM_TENSOR_SSM_DT, "bias", i), {d_inner}, flags);
 
-            layer.ssm_g_a    = create_tensor(tn(LLM_TENSOR_SSM_G_A,  "weight", i), {n_embd, head_dim}, 0);
-            layer.ssm_g_b    = create_tensor(tn(LLM_TENSOR_SSM_G_B,  "weight", i), {head_dim, d_inner}, 0);
-            layer.ssm_o_norm = create_tensor(tn(LLM_TENSOR_SSM_NORM, "weight", i), {head_dim}, 0);
-            layer.wo         = create_tensor(tn(LLM_TENSOR_ATTN_OUT, "weight", i), {d_inner, n_embd}, 0);
+            layer.ssm_g_a    = create_tensor(tn(LLM_TENSOR_SSM_G_A,  "weight", i), {n_embd, head_dim}, flags);
+            layer.ssm_g_b    = create_tensor(tn(LLM_TENSOR_SSM_G_B,  "weight", i), {head_dim, d_inner}, flags);
+            layer.ssm_o_norm = create_tensor(tn(LLM_TENSOR_SSM_NORM, "weight", i), {head_dim}, flags);
+            layer.wo         = create_tensor(tn(LLM_TENSOR_ATTN_OUT, "weight", i), {d_inner, n_embd}, flags);
         } else {
             const int64_t q_lora_rank      = hparams.n_lora_q;
             const int64_t kv_lora_rank     = hparams.n_lora_kv;
@@ -155,9 +159,9 @@ void llama_model_glm5_next::load_arch_tensors(llama_model_loader & ml) {
         }
 
         if (i < (int) hparams.n_layer_dense_lead) {
-            layer.ffn_gate = create_tensor(tn(LLM_TENSOR_FFN_GATE, "weight", i), {n_embd, n_ff}, 0);
-            layer.ffn_down = create_tensor(tn(LLM_TENSOR_FFN_DOWN, "weight", i), {n_ff, n_embd}, 0);
-            layer.ffn_up   = create_tensor(tn(LLM_TENSOR_FFN_UP,   "weight", i), {n_embd, n_ff}, 0);
+            layer.ffn_gate = create_tensor(tn(LLM_TENSOR_FFN_GATE, "weight", i), {n_embd, n_ff}, flags);
+            layer.ffn_down = create_tensor(tn(LLM_TENSOR_FFN_DOWN, "weight", i), {n_ff, n_embd}, flags);
+            layer.ffn_up   = create_tensor(tn(LLM_TENSOR_FFN_UP,   "weight", i), {n_embd, n_ff}, flags);
         } else {
             const int64_t n_ff_exp        = hparams.n_ff_exp(i);
             const int64_t n_expert_shared = hparams.n_expert_shared;

PTAL and add a TODO to consolidate and reuse this logic. I think we had the same issue for a few other models already.

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Tested your patch with the split ggml-org MTP head and it works identically to the embedded one, with the same acceptance and no reallocation.

Make the trunk tensors optional when the file only holds the NextN
layer, and the NextN tensors optional when the file only holds the
trunk, so the split MTP GGUF loads as a draft model.

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
@ServeurpersoCom

Copy link
Copy Markdown
Contributor

The same trunk_only/mtp_only block is copied in 16 models, so once this one is merged I can open a follow-up that moves it into a shared helper, as the TODO says.

@pwilkin

pwilkin commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

Aight, let's merge and consolidate in a followup.

@pwilkin
pwilkin merged commit b9acf13 into ggml-org:master Oct 7, 2026
16 of 18 checks passed
@ggerganov ggerganov added the highlight Changes that will be highlighted in the next release notes label Oct 7, 2026
@ggerganov

Copy link
Copy Markdown
Member

Image input aborts when MTP is enabled:

[65142] 0.25.466.832 I decoding image batch 1/1, n_tokens_batch = 115
[65142] 0.25.476.017 I image decoded (batch 1/1) in 9 ms
[65142] 0.26.306.851 E init: the tokens of sequence 3 in the input batch have inconsistent sequence positions:
[65142]  - the last position stored in the memory module of the context (i.e. the KV cache) for sequence 3 is X = 3909
[65142]  - the tokens for sequence 3 in the input batch have a starting position of Y = 4025
[65142]  it is required that the sequence positions remain consecutive: Y = X + 1
[65142] 0.26.306.856 E decode: failed to initialize batch
[65142] 0.26.306.857 E spec      process: llama_process(ctx_dft) head=0 failed rc=-1 (pos=4025)
[65142] 0.26.306.867 E srv        decode: failed to process speculative batch
[65142] 0.26.306.898 E srv  update_slots: decode() failed: failed to process speculative batch
[65142] 0.26.306.902 E srv    send_error: task id = 0, error: decode() failed: failed to process speculative batch
[65142] 0.26.306.906 I slot      release: id  3 | task 0 | stop processing: n_tokens = 4026, truncated = 0
[65142] 0.26.306.915 W srv          stop: cancel task, id_task = 0
[65142] 0.26.307.056 I srv  update_slots: all slots are idle

I think with the new llama_batch_ext API we should be able to support this.

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

@ngxson your commit ngxson@a748c27057 looks like exactly the fix for this, do you plan to continue it? I can reproduce the crash with GLM-5.3-Flash + mmproj and test your branch whenever you want.

@amaze28

amaze28 commented Oct 8, 2026

Copy link
Copy Markdown

Validation on ROCm/gfx906 (5x MI50): long-prompt crash fixed, big speedup — and confirming the image-input abort on HIP as well

Following up on the crash I reported against #27917 (abort in the draft-mtp path on ~18K-token prompts): this merged implementation fixes it. Same rig, same model, same 18K repro prompt — now completes correctly with correct retrieval of a planted fact, server healthy after.

Environment: llama.cpp master (post-#29928), HIP backend (-DGGML_HIP=ON -DAMDGPU_TARGETS=gfx906, ROCm 7.2.1), 5x AMD Instinct MI50 32GB, tensor-split 1,1,1,1,1. Model: unsloth/GLM-5.3-Flash-GGUF UD-IQ3_XXS (with blk.45.nextn.*). Flags: -ngl 999 -fa on -ctk q8_0 -ctv q8_0 -c 196608 --spec-type draft-mtp.

Results vs 11.7 t/s undrafted baseline (256-token generations):

  • n-max=1: 17.87 t/s (acceptance 0.752)
  • n-max=2: 18.41 t/s (acceptance 0.655, mean len 2.31)
  • n-max=3: 15.70 t/s (acceptance 0.461)
  • 18K-token prompt: prefill 143.6 t/s (was 103.9 pre-MTP — nice prefill gain too), decode-after 16.26 t/s, no crash.

Fastest configuration this hardware has ever served — thanks for the clean implementation.

Re: image input aborting with MTP enabled — reproduces identically on ROCm/gfx906. With --mmproj loaded alongside --spec-type draft-mtp, any image_url request returns:

{"error":{"code":500,"message":"decode() failed: failed to process speculative batch","type":"server_error"}}

Text-only requests on the same server work fine; per-request "speculative.n_max": 0 does not avoid it. Reproducing identically on HIP supports this being the common speculative path (draft context missing the image-token positions) rather than backend kernels. Happy to test patches here — this 5x MI50 rig is a standing ROCm validation bench.

EmeraldBitTwizzler pushed a commit to EmeraldBitTwizzler/llama.cpp that referenced this pull request Oct 9, 2026
…at merge past b11453, drop unslothai#247, pin unslothai#251

Upstream K2 Horizon (ggml-org#29535), the glm5-next gather removal (ggml-org#30042) and the
GLM5-Next MTP graph (ggml-org#29928) broke the Inkling, unslothai#243 and unslothai#241 pins. Each
branch now carries a merge of upstream master. unslothai#247 is in every tag from
b11452 on. unslothai#251 adds --moe-cache-mib auto on top of ggml-org#29887.
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 9, 2026
An MTP head loaded from its own file joins the expert cache as a member
with its own file backing, so the finite-L2 and prefill-swap gates can
accept it like a joined draft model.

With --expert-cache-draft off, an MTP head (-md) with every layer on the GPU
and no draft tensor overrides reads no host or file expert, so the finite
tier and the prefill swap never see its work. The gate accepts it as well.
In the joint cache the head competed with the target for L1 and landed
mostly in the file (DS4F roleplay, L1 16 GiB, L2 60 GiB: 16 of 256 experts
in L1, MTP n1 -7.0 % decode); kept on the GPU it gives +14.1 % over no
speculation.

An MTP context on the target's own weights (draft-mtp without -md)
passes as well: its NextN block is one more routed layer of the target's
cache (GLM5-Next with upstream PR ggml-org#29928, finite L2, one slot). The server
also joins that context to the profile, so the rows of the NextN block feed
the decode bank while it generates; joined as a separate draft model only,
the block counted no selection, got no VRAM slice and every draft step read
its experts from the host tier. RANMA_MTP_EXPERT_JOIN=0 leaves the
own-weights context out of the profile.

test-expert-config covers both gates.
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 9, 2026
docs/ranma/exl3.md: the EXL3 types and the load check, the conversion (MTP layers follow the model class;
a model whose MTP layers cannot be stored as EXL3, such as Qwen3.8-Flash-Next, is converted without them
after an error message), the backends, the models checked (llama, qwen3moe, lfm2moe, qwen4exp, deepseek4,
glm5-next; loading is not limited to them), the MTP layer of glm5-next through the NextN graph of
upstream PR ggml-org#29928 with --spec-type draft-mtp and what was checked with it, the precision of the prompt
GEMM and the F32 request, the RDNA3 status (compiles, GEMV on the GPU, the WMMA GEMM is RDNA4 only,
untested) and GGML_CUDA_HC_GATED_FUSION=0. README: one section with both lines, and the ExLlamaV3 credit
next to the license.
saturnsky added a commit to saturnsky/ranma.cpp that referenced this pull request Oct 9, 2026
The README now leads with what the fork is for and its three main features, and leaves the full
list of changes to a page of its own.

- What this fork is: also one person's own development work in a single session. The expert cache
  installs a plan only when every slot is idle, so sessions that send requests in turn leave it few
  chances; this is said before the advice to use upstream for many users.
- Main features: the expert cache, the smart MTP draft length and EXL3 weights, a few lines each
  with their pages. EXL3 needs GGUF files in the format of this fork, from the Hugging Face
  repositories or, for another EXL3 model, the converter, with the note that not every model is
  guaranteed to work yet; those files do not open in other tools.
- The expert cache entry, expert-cache.md and expert-cache-l2.md say to give the slot count with
  -np: without it the server runs four slots, the per-slot state grows with them, and a finite
  host tier is refused.
- Benchmarks becomes a section of its own; the numbers come with the benchmark commit, the last
  of the series.
- Roadmap: the porting paragraph is gone; the section lists plans that are not done yet:
  DeepSeek V4 Flash performance work, DeepSeek V4.1 Flash with EXL3, and experimental RDNA3 and CUDA
  builds. GLM 5.3 is no longer listed: upstream has merged it, and its EXL3 files convert and run.
- Target environment: other environments may work but are not guaranteed, as in CONTRIBUTING.md,
  with a table of the main features on RDNA4 and RDNA3 that tells maintainer tests, user reports
  and compile checks apart.
- The list of changes over upstream moves to docs/ranma/README.md, with its links made relative to
  that directory, together with the rule of one line and one page per change. It gains glm5-next in the
  EXL3 line and a GLM-5.3 Flash section for LoRA on the attention output, described in
  graph-runtime.md, and for the MTP draft of upstream PR ggml-org#29928, carried as its six commits unchanged.
- deepseek-v4.md: the selected-cell attention is a HIP path measured on RDNA4 only, not an
  RDNA4-only path; its gather has no RDNA4 condition.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

highlight Changes that will be highlighted in the next release notes model Model specific

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants