Repository navigation
feat: add GLM5Next MTP, optimize - #29928
Conversation
| cur = build_norm(cur, layer.attn_norm, nullptr, LLM_NORM_RMS, il); | ||
| cb(cur, "mtp_attn_norm", il); | ||
|
|
||
| const bool headless = n_outputs == 0 && !(cparams.embeddings_nextn && !cparams.embeddings_nextn_masked); |
There was a problem hiding this comment.
what's this headless stuff? can you keep it consistent with other MTP?
There was a problem hiding this comment.
You mean implement headless for other viable MTP models too? :)
There was a problem hiding this comment.
headless skips the attention when you don't have any outputs to write, i.e. you're just populating the KV cache
There was a problem hiding this comment.
Yeah sure, but make sure it doesn't cause any random graph re-allocations or graph reuse issues.
There was a problem hiding this comment.
I think due to previous MTP graph caching shenanigans n_outputs is already included in the graph hash code input params, so it should be safe, but ofc I'll check.
There was a problem hiding this comment.
headless skips the attention
You mean skips the FFN, right? The attention should not be skipped even for empty outputs.
There was a problem hiding this comment.
On second thought: I would recommend not adding this in this PR rather in it's own PR, so that we can measure performance and any graph issues.
There was a problem hiding this comment.
No, that's the clever part - you actually skip attention, i.e. at least half of the attention, since you don't need the attention output for anything - you just run the part that's needed to compute the KV cache parameters.
There was a problem hiding this comment.
On second thought: I would recommend not adding this in this PR rather in it's own PR, so that we can measure performance and any graph issues.
Yes, keep it simple. Currently all sparse attention graphs have some problem because they reallocate (#29867 (comment)). We need to fix this first before adding more logic.
There was a problem hiding this comment.
Aight, will do a standard "skip FFN" behavior and add the headless mode in a separate PR.
| if (cell.seq_id.size() > 1) { | ||
| return false; | ||
| } |
There was a problem hiding this comment.
In which cases does this happen?
There was a problem hiding this comment.
It doesn't - it's a defensive check added to surface potential silent data corruption. I can remove it or add an extra warning message if you prefer.
There was a problem hiding this comment.
(but IMO this is worth keeping because common_context_seq_rm is part of the API, so the function should actually refuse to do this instead of silently corrupting the data)
There was a problem hiding this comment.
It should be GGML_ASSERT I think. I don't see how we can end up with more than one sequence in a recurrent cell.
There was a problem hiding this comment.
@ggerganov it happens every time on llama_memory_recurrent::seq_cp, there's no copy-on-write behavior, so the two sequences end up sharing the cell.
There was a problem hiding this comment.
In "real life" this is harmless, because once any decode happens the sequences disconnect from each other, but it's obtainable from a legal sequence of API calls.
There was a problem hiding this comment.
Ok, let's keep it like this for now. Basically the condition is going to occur if one tries to rollback a sequence that was just copied - we will deny that.
The recurrent state rollback logic in general needs to be asserted anyway since there are various failure cases that currently fail silently (f.ex, trying to rollback after a decode batch). It's something in the TODO list already.
|
Here is performance data from my local server - i hope it is useful
Result: #29928 vs base two runs per build, each -r 3, alternating master/PR/master/PR
the tests were run and the tables were created by Claude Opus |
|
@ggerganov @am17an all right, removed the |
|
|
||
| if (inp->gather) { | ||
| inp->gather_mask_h = ggml_repeat_4d(ctx0, inp->gather_mask, n_sel, n_head, 1, n_tokens); | ||
| } |
There was a problem hiding this comment.
Rebase on latest master because this logic has changed.
|
@ggerganov done, also exposed a bug during testing (see #30020) |
|
#30042 removes the gather path (merge order is #30017, then #30042, then this one), so after rebasing you can drop deacfbf and the gather_mask_h hunk entirely. Your fix does make the gather about 5% faster than sparse FA at -ub 16, I benchmarked it, but since -ub 16 has no real use case and would stay untested, we're keeping a single sparse path. Also note gather_mask becomes sel_mask. EDIT: @pwilkin, I tested the rebased version on my side, with a small extra commit moving the MTP graph to the new nextn crop helpers: on GLM-5.3-Flash it doubles decode at short context (97% acceptance) and gives +52% at 32k, with no reallocation. Want me to push it to your branch? |
|
@ServeurpersoCom yeah, do it :) |
An MTP head loaded from its own file joins the expert cache as a member with its own file backing, so the finite-L2 and prefill-swap gates can accept it like a joined draft model. With --expert-cache-draft off, an MTP head (-md) with every layer on the GPU and no draft tensor overrides reads no host or file expert, so the finite tier and the prefill swap never see its work. The gate accepts it as well. In the joint cache the head competed with the target for L1 and landed mostly in the file (DS4F roleplay, L1 16 GiB, L2 60 GiB: 16 of 256 experts in L1, MTP n1 -7.0 % decode); kept on the GPU it gives +14.1 % over no speculation. An MTP context on the target's own weights (draft-mtp without -md) passes as well: its NextN block is one more routed layer of the target's cache (GLM5-Next with upstream PR ggml-org#29928, finite L2, one slot). test-expert-config covers both gates.
docs/ranma/exl3.md: the EXL3 types and the load check, the conversion (MTP layers follow the model class; a model whose MTP layers cannot be stored as EXL3, such as Qwen3.8-Flash-Next, is converted without them after an error message), the backends, the models checked (llama, qwen3moe, lfm2moe, qwen4exp, deepseek4, glm5-next; loading is not limited to them), the MTP layer of glm5-next through the NextN graph of upstream PR ggml-org#29928 with --spec-type draft-mtp and what was checked with it, the precision of the prompt GEMM and the F32 request, the RDNA3 status (compiles, GEMV on the GPU, the WMMA GEMM is RDNA4 only, untested) and GGML_CUDA_HC_GATED_FUSION=0. README: one section with both lines, and the ExLlamaV3 credit next to the license.
The README now leads with what the fork is for and its three main features, and leaves the full list of changes to a page of its own. - What this fork is: also one person's own development work in a single session. The expert cache installs a plan only when every slot is idle, so sessions that send requests in turn leave it few chances; this is said before the advice to use upstream for many users. - Main features: the expert cache, the smart MTP draft length and EXL3 weights, a few lines each with their pages. EXL3 needs GGUF files in the format of this fork, from the Hugging Face repositories or, for another EXL3 model, the converter, with the note that not every model is guaranteed to work yet; those files do not open in other tools. - The expert cache entry, expert-cache.md and expert-cache-l2.md say to give the slot count with -np: without it the server runs four slots, the per-slot state grows with them, and a finite host tier is refused. - Benchmarks becomes a section of its own; the numbers come with the benchmark commit, the last of the series. - Roadmap: the porting paragraph is gone; the section lists plans that are not done yet: DeepSeek V4 Flash performance work, DeepSeek V4.1 Flash with EXL3, and experimental RDNA3 and CUDA builds. GLM 5.3 is no longer listed: upstream has merged it, and its EXL3 files convert and run. - Target environment: other environments may work but are not guaranteed, as in CONTRIBUTING.md, with a table of the main features on RDNA4 and RDNA3 that tells maintainer tests, user reports and compile checks apart. - The list of changes over upstream moves to docs/ranma/README.md, with its links made relative to that directory, together with the rule of one line and one page per change. It gains glm5-next in the EXL3 line and a GLM-5.3 Flash section for LoRA on the attention output, described in graph-runtime.md, and for the MTP draft of upstream PR ggml-org#29928, carried as its six commits unchanged. - deepseek-v4.md: the selected-cell attention is a HIP path measured on RDNA4 only, not an RDNA4-only path; its gather has no RDNA4 condition.
Build the GLM5-Next multi-token-prediction head as graph_mtp: the NextN block embeds enorm(tok)+hnorm(h) through eh_proj, runs one plain DSA layer and the shared lm_head, reusing the trunk's builders through the no_build tag ctor. llama_memory_recurrent also tolerates a partial seq_rm when the context holds no recurrent layers, which is what the MTP draft context needs. Assisted-by: Claude
A NextN forward with no output rows (the MTP catch-up and the draft-context prefill) persists only through its cache writes, so the headless graph keeps the MLA latent, indexer key|gate and pooled-key writes and drops the query path, the indexer selection, the attention body, the FFN and the LM head. The 4-token catch-up falls from 6.9 ms to 0.33 ms of kernels; the greedy output hashes and the draft acceptance are unchanged. Assisted-by: Claude
…lback Three fixes from the architectural review. The headless graph prune now also requires that no unmasked nextn extraction is live, because that mode reads n_tokens hidden rows regardless of the logits flags. Masked extraction publishes the hidden rows gathered by the output ids, so a batch whose output flags are not a prefix exports the right rows. A partial recurrent rollback whose tail cell is shared with another sequence is now rejected instead of silently moving that sequence's tail. Assisted-by: Claude
Assisted-by: Claude
…runing it Replace the headless NextN prune with the crop pattern the other MTP graphs use: gather the attention output and the block input at the output ids before the position-wise FFN and the shared head. A NextN forward with no output rows (the MTP catch-up and the draft-context prefill) then runs the FFN and the head over zero rows. The 4-token catch-up falls from 6.9 ms to 2.9 ms of kernels; greedy output hashes are unchanged. Assisted-by: Claude
dc574e2 to
5d9e4cf
Compare
|
Rebase only |
Replace the local crop condition and the masked select of t_h_nextn with crop_before_nextn and crop_after_nextn, so the MTP graph narrows its rows the same way as the main graph and the other models. Describe the shared cell and empty filter branches of the recurrent partial rollback.
|
Compared against unsloth's MTP implementation on the same GGUF (GLM-5.3-Flash UD-IQ1_S, greedy): acceptance is in the same range with no systematic gap (n=2: 99/71/75/81/78% vs 99/78/74/78/74% on 5 prompts up to 32k), and on the one prompt where both base models agree, all configs of both implementations produce the exact same output. Elsewhere the base trunks themselves drift at near ties on this 1.5 bpw quant, so a token level comparison isn't possible. |
|
Okay, so I guess this is gtg? |
|
Doing a quick test on my end. |
|
Yes from my side, it just needs a second approval now. I'm playing with it on my RTX PRO 6000, but even the 1-bit quant is a tight fit: with MTP on, the draft context eats enough VRAM that I drop from 128K to 48K of context. |
|
Hm, trying to run the Q2_K model + Q4_0 mtp from here: https://huggingface.co/ggml-org/GLM-5.3-Flash-GGUF/tree/main I get this error: |
|
This is the command that I am trying: ./bin/llama-server -hf ggml-org/GLM-5.3-Flash-GGUF:Q2_K --spec-type draft-mtp |
|
@ggerganov testing on the quant I use now |
|
I tested with the unsloth UD-IQ1_S quant, which embeds the NextN layer, with the first shard replaced by the one from their Shard_Rewrite folder (the original uses the old glm5next arch name): llama-server -m GLM-5.3-Flash-UD-IQ1_S-00001-of-00003.gguf -ngl 99 -fa on -c 40960 -np 1 --spec-type draft-mtp --spec-draft-n-max 2 |
|
Looking at Detailsdiff --git a/src/models/glm5-next.cpp b/src/models/glm5-next.cpp
index 440f90dce2..1537a2a63a 100644
--- a/src/models/glm5-next.cpp
+++ b/src/models/glm5-next.cpp
@@ -64,9 +64,12 @@ void llama_model_glm5_next::load_arch_tensors(llama_model_loader & ml) {
const int64_t hc = hparams.dsv4_hc_mult;
const int64_t hc_mix_dim = (2 + hc)*hc;
- // the NextN block is loaded but only used by the MTP graph.
- // Separated trunk_only/mtp_only handling TODO with DECODER_MTP graph in the MTP follow up
- int mtp_flags = 0;
+ const bool mtp_only = (hparams.n_layer_nextn > 0) && (ml.get_weight("blk.0.attn_norm.weight") == nullptr);
+ const std::string mtp_probe = "blk." + std::to_string(n_layer) + ".nextn.eh_proj.weight";
+ const bool trunk_only = (hparams.n_layer_nextn > 0) && (ml.get_weight(mtp_probe.c_str()) == nullptr);
+ const int trunk_flags = mtp_only ? TENSOR_NOT_REQUIRED : 0;
+ int mtp_flags = trunk_only ? TENSOR_NOT_REQUIRED : 0;
+
if (!ml.load_mtp) {
mtp_flags |= TENSOR_SKIP;
}
@@ -79,18 +82,19 @@ void llama_model_glm5_next::load_arch_tensors(llama_model_loader & ml) {
for (int i = 0; i < n_layer_all; ++i) {
auto & layer = layers[i];
- const int flags = (i >= n_layer) ? mtp_flags : 0;
+ const bool is_nextn = i >= n_layer;
+ const int flags = is_nextn ? mtp_flags : trunk_flags;
layer.attn_norm = create_tensor(tn(LLM_TENSOR_ATTN_NORM, "weight", i), {n_embd}, flags);
layer.ffn_norm = create_tensor(tn(LLM_TENSOR_FFN_NORM, "weight", i), {n_embd}, flags);
if (i < n_layer) {
- layer.hc_attn_fn = create_tensor(tn(LLM_TENSOR_HC_ATTN_FN, "weight", i), {hc*n_embd, hc_mix_dim}, 0);
- layer.hc_attn_base = create_tensor(tn(LLM_TENSOR_HC_ATTN_BASE, "weight", i), {hc_mix_dim}, 0);
- layer.hc_attn_scale = create_tensor(tn(LLM_TENSOR_HC_ATTN_SCALE, "weight", i), {3}, 0);
- layer.hc_ffn_fn = create_tensor(tn(LLM_TENSOR_HC_FFN_FN, "weight", i), {hc*n_embd, hc_mix_dim}, 0);
- layer.hc_ffn_base = create_tensor(tn(LLM_TENSOR_HC_FFN_BASE, "weight", i), {hc_mix_dim}, 0);
- layer.hc_ffn_scale = create_tensor(tn(LLM_TENSOR_HC_FFN_SCALE, "weight", i), {3}, 0);
+ layer.hc_attn_fn = create_tensor(tn(LLM_TENSOR_HC_ATTN_FN, "weight", i), {hc*n_embd, hc_mix_dim}, flags);
+ layer.hc_attn_base = create_tensor(tn(LLM_TENSOR_HC_ATTN_BASE, "weight", i), {hc_mix_dim}, flags);
+ layer.hc_attn_scale = create_tensor(tn(LLM_TENSOR_HC_ATTN_SCALE, "weight", i), {3}, flags);
+ layer.hc_ffn_fn = create_tensor(tn(LLM_TENSOR_HC_FFN_FN, "weight", i), {hc*n_embd, hc_mix_dim}, flags);
+ layer.hc_ffn_base = create_tensor(tn(LLM_TENSOR_HC_FFN_BASE, "weight", i), {hc_mix_dim}, flags);
+ layer.hc_ffn_scale = create_tensor(tn(LLM_TENSOR_HC_FFN_SCALE, "weight", i), {3}, flags);
}
const int64_t head_dim = hparams.n_embd_head_kda;
@@ -100,25 +104,25 @@ void llama_model_glm5_next::load_arch_tensors(llama_model_loader & ml) {
if (hparams.is_recr(i)) {
auto conv = [&](llm_tensor tid) {
ggml_tensor * t = create_tensor(tn(tid, "weight", i), {d_conv, 1, d_inner, 1}, TENSOR_NOT_REQUIRED);
- return t ? t : create_tensor(tn(tid, "weight", i), {d_conv, 1, d_inner}, 0);
+ return t ? t : create_tensor(tn(tid, "weight", i), {d_conv, 1, d_inner}, flags);
};
layer.ssm_q_conv = conv(LLM_TENSOR_SSM_CONV1D_Q);
layer.ssm_k_conv = conv(LLM_TENSOR_SSM_CONV1D_K);
layer.ssm_v_conv = conv(LLM_TENSOR_SSM_CONV1D_V);
- create_tensor_qkv(layer, i, n_embd, d_inner, d_inner, d_inner, 0);
+ create_tensor_qkv(layer, i, n_embd, d_inner, d_inner, d_inner, flags);
- layer.ssm_f_a = create_tensor(tn(LLM_TENSOR_SSM_F_A, "weight", i), {n_embd, head_dim}, 0);
- layer.ssm_f_b = create_tensor(tn(LLM_TENSOR_SSM_F_B, "weight", i), {head_dim, d_inner}, 0);
- layer.ssm_beta = create_tensor(tn(LLM_TENSOR_SSM_BETA, "weight", i), {n_embd, n_head}, 0);
+ layer.ssm_f_a = create_tensor(tn(LLM_TENSOR_SSM_F_A, "weight", i), {n_embd, head_dim}, flags);
+ layer.ssm_f_b = create_tensor(tn(LLM_TENSOR_SSM_F_B, "weight", i), {head_dim, d_inner}, flags);
+ layer.ssm_beta = create_tensor(tn(LLM_TENSOR_SSM_BETA, "weight", i), {n_embd, n_head}, flags);
- layer.ssm_a = create_tensor(tn(LLM_TENSOR_SSM_A_NOSCAN, i), {n_head}, 0);
- layer.ssm_dt_b = create_tensor(tn(LLM_TENSOR_SSM_DT, "bias", i), {d_inner}, 0);
+ layer.ssm_a = create_tensor(tn(LLM_TENSOR_SSM_A_NOSCAN, i), {n_head}, flags);
+ layer.ssm_dt_b = create_tensor(tn(LLM_TENSOR_SSM_DT, "bias", i), {d_inner}, flags);
- layer.ssm_g_a = create_tensor(tn(LLM_TENSOR_SSM_G_A, "weight", i), {n_embd, head_dim}, 0);
- layer.ssm_g_b = create_tensor(tn(LLM_TENSOR_SSM_G_B, "weight", i), {head_dim, d_inner}, 0);
- layer.ssm_o_norm = create_tensor(tn(LLM_TENSOR_SSM_NORM, "weight", i), {head_dim}, 0);
- layer.wo = create_tensor(tn(LLM_TENSOR_ATTN_OUT, "weight", i), {d_inner, n_embd}, 0);
+ layer.ssm_g_a = create_tensor(tn(LLM_TENSOR_SSM_G_A, "weight", i), {n_embd, head_dim}, flags);
+ layer.ssm_g_b = create_tensor(tn(LLM_TENSOR_SSM_G_B, "weight", i), {head_dim, d_inner}, flags);
+ layer.ssm_o_norm = create_tensor(tn(LLM_TENSOR_SSM_NORM, "weight", i), {head_dim}, flags);
+ layer.wo = create_tensor(tn(LLM_TENSOR_ATTN_OUT, "weight", i), {d_inner, n_embd}, flags);
} else {
const int64_t q_lora_rank = hparams.n_lora_q;
const int64_t kv_lora_rank = hparams.n_lora_kv;
@@ -155,9 +159,9 @@ void llama_model_glm5_next::load_arch_tensors(llama_model_loader & ml) {
}
if (i < (int) hparams.n_layer_dense_lead) {
- layer.ffn_gate = create_tensor(tn(LLM_TENSOR_FFN_GATE, "weight", i), {n_embd, n_ff}, 0);
- layer.ffn_down = create_tensor(tn(LLM_TENSOR_FFN_DOWN, "weight", i), {n_ff, n_embd}, 0);
- layer.ffn_up = create_tensor(tn(LLM_TENSOR_FFN_UP, "weight", i), {n_embd, n_ff}, 0);
+ layer.ffn_gate = create_tensor(tn(LLM_TENSOR_FFN_GATE, "weight", i), {n_embd, n_ff}, flags);
+ layer.ffn_down = create_tensor(tn(LLM_TENSOR_FFN_DOWN, "weight", i), {n_ff, n_embd}, flags);
+ layer.ffn_up = create_tensor(tn(LLM_TENSOR_FFN_UP, "weight", i), {n_embd, n_ff}, flags);
} else {
const int64_t n_ff_exp = hparams.n_ff_exp(i);
const int64_t n_expert_shared = hparams.n_expert_shared;PTAL and add a TODO to consolidate and reuse this logic. I think we had the same issue for a few other models already. |
|
Tested your patch with the split ggml-org MTP head and it works identically to the embedded one, with the same acceptance and no reallocation. |
Make the trunk tensors optional when the file only holds the NextN layer, and the NextN tensors optional when the file only holds the trunk, so the split MTP GGUF loads as a draft model. Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
|
The same trunk_only/mtp_only block is copied in 16 models, so once this one is merged I can open a follow-up that moves it into a shared helper, as the TODO says. |
|
Aight, let's merge and consolidate in a followup. |
|
Image input aborts when MTP is enabled: I think with the new |
|
@ngxson your commit ngxson@a748c27057 looks like exactly the fix for this, do you plan to continue it? I can reproduce the crash with GLM-5.3-Flash + mmproj and test your branch whenever you want. |
|
Validation on ROCm/gfx906 (5x MI50): long-prompt crash fixed, big speedup — and confirming the image-input abort on HIP as well Following up on the crash I reported against #27917 (abort in the draft-mtp path on ~18K-token prompts): this merged implementation fixes it. Same rig, same model, same 18K repro prompt — now completes correctly with correct retrieval of a planted fact, server healthy after. Environment: llama.cpp master (post-#29928), HIP backend ( Results vs 11.7 t/s undrafted baseline (256-token generations):
Fastest configuration this hardware has ever served — thanks for the clean implementation. Re: image input aborting with MTP enabled — reproduces identically on ROCm/gfx906. With Text-only requests on the same server work fine; per-request |
…at merge past b11453, drop unslothai#247, pin unslothai#251 Upstream K2 Horizon (ggml-org#29535), the glm5-next gather removal (ggml-org#30042) and the GLM5-Next MTP graph (ggml-org#29928) broke the Inkling, unslothai#243 and unslothai#241 pins. Each branch now carries a merge of upstream master. unslothai#247 is in every tag from b11452 on. unslothai#251 adds --moe-cache-mib auto on top of ggml-org#29887.
An MTP head loaded from its own file joins the expert cache as a member with its own file backing, so the finite-L2 and prefill-swap gates can accept it like a joined draft model. With --expert-cache-draft off, an MTP head (-md) with every layer on the GPU and no draft tensor overrides reads no host or file expert, so the finite tier and the prefill swap never see its work. The gate accepts it as well. In the joint cache the head competed with the target for L1 and landed mostly in the file (DS4F roleplay, L1 16 GiB, L2 60 GiB: 16 of 256 experts in L1, MTP n1 -7.0 % decode); kept on the GPU it gives +14.1 % over no speculation. An MTP context on the target's own weights (draft-mtp without -md) passes as well: its NextN block is one more routed layer of the target's cache (GLM5-Next with upstream PR ggml-org#29928, finite L2, one slot). The server also joins that context to the profile, so the rows of the NextN block feed the decode bank while it generates; joined as a separate draft model only, the block counted no selection, got no VRAM slice and every draft step read its experts from the host tier. RANMA_MTP_EXPERT_JOIN=0 leaves the own-weights context out of the profile. test-expert-config covers both gates.
docs/ranma/exl3.md: the EXL3 types and the load check, the conversion (MTP layers follow the model class; a model whose MTP layers cannot be stored as EXL3, such as Qwen3.8-Flash-Next, is converted without them after an error message), the backends, the models checked (llama, qwen3moe, lfm2moe, qwen4exp, deepseek4, glm5-next; loading is not limited to them), the MTP layer of glm5-next through the NextN graph of upstream PR ggml-org#29928 with --spec-type draft-mtp and what was checked with it, the precision of the prompt GEMM and the F32 request, the RDNA3 status (compiles, GEMV on the GPU, the WMMA GEMM is RDNA4 only, untested) and GGML_CUDA_HC_GATED_FUSION=0. README: one section with both lines, and the ExLlamaV3 credit next to the license.
The README now leads with what the fork is for and its three main features, and leaves the full list of changes to a page of its own. - What this fork is: also one person's own development work in a single session. The expert cache installs a plan only when every slot is idle, so sessions that send requests in turn leave it few chances; this is said before the advice to use upstream for many users. - Main features: the expert cache, the smart MTP draft length and EXL3 weights, a few lines each with their pages. EXL3 needs GGUF files in the format of this fork, from the Hugging Face repositories or, for another EXL3 model, the converter, with the note that not every model is guaranteed to work yet; those files do not open in other tools. - The expert cache entry, expert-cache.md and expert-cache-l2.md say to give the slot count with -np: without it the server runs four slots, the per-slot state grows with them, and a finite host tier is refused. - Benchmarks becomes a section of its own; the numbers come with the benchmark commit, the last of the series. - Roadmap: the porting paragraph is gone; the section lists plans that are not done yet: DeepSeek V4 Flash performance work, DeepSeek V4.1 Flash with EXL3, and experimental RDNA3 and CUDA builds. GLM 5.3 is no longer listed: upstream has merged it, and its EXL3 files convert and run. - Target environment: other environments may work but are not guaranteed, as in CONTRIBUTING.md, with a table of the main features on RDNA4 and RDNA3 that tells maintainer tests, user reports and compile checks apart. - The list of changes over upstream moves to docs/ranma/README.md, with its links made relative to that directory, together with the rule of one line and one page per change. It gains glm5-next in the EXL3 line and a GLM-5.3 Flash section for LoRA on the attention output, described in graph-runtime.md, and for the MTP draft of upstream PR ggml-org#29928, carried as its six commits unchanged. - deepseek-v4.md: the selected-cell attention is a HIP path measured on RDNA4 only, not an RDNA4-only path; its gather has no RDNA4 condition.
Overview
Implement MTP for GLM5Next, optimize the graph
Additional information
Supersedes #27917
Requirements