Skip to content

llama: share the nextn tensor flags between models - #30097

Merged
ServeurpersoCom merged 2 commits into
ggml-org:masterfrom
ServeurpersoCom:nextn-flags-helper
Oct 7, 2026
Merged

ServeurpersoCom merged 2 commits into
ggml-org:masterfrom
ServeurpersoCom:nextn-flags-helper

Conversation

@ServeurpersoCom

@ServeurpersoCom ServeurpersoCom commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Follow-up of the TODO in glm5-next: move the trunk-only and MTP-only detection that each model copied into a nextn_flags helper of llama_model_base. It probes the first trunk layer and the first NextN layer, and adds TENSOR_SKIP when MTP is not loaded. qwen4exp probes hc_attn_norm since it has no attn_norm.

deepseek4, nemotron-h, qwen35, qwen35moe, qwen3next and qwen4exp now also accept a trunk-only file, like the other models.

Additional information

Follow-up #29928

Requirements

Follow-up of the TODO in glm5-next: move the trunk-only and MTP-only
detection that each model copied into a nextn_flags helper of
llama_model_base. It probes the first trunk layer and the first NextN
layer, and adds TENSOR_SKIP when MTP is not loaded. qwen4exp probes
hc_attn_norm since it has no attn_norm.

deepseek4, nemotron-h, qwen35, qwen35moe, qwen3next and qwen4exp now
also accept a trunk-only file, like the other models.
@ServeurpersoCom
ServeurpersoCom requested a review from CISC as a code owner October 7, 2026 11:27
@github-actions github-actions Bot added the model Model specific label Oct 7, 2026
@CISC

CISC commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Huh, GCC15 didn't like that:
https://github.com/ggml-org/llama.cpp/actions/runs/37614260360/job/112768687955

Edit: Guess you'll have to return {} directly.

Lambdas that capture structured bindings need C++20, and GCC 15
rejects them under -Werror, so the models read the trunk and MTP
flags into plain variables.
@ServeurpersoCom

Copy link
Copy Markdown
Contributor Author

Huh, GCC15 didn't like that: https://github.com/ggml-org/llama.cpp/actions/runs/37614260360/job/112768687955

Edit: Guess you'll have to return {} directly.

Ha, Thanks! I flagged this exact one on another PR and still walked right into it. Fixed: the lambdas in step35 and the qwen35 family captured the structured bindings, so the models now read the flags into plain variables, and the helper returns them directly.

@CISC

CISC commented Oct 7, 2026

Copy link
Copy Markdown
Member

Huh, GCC15 didn't like that: https://github.com/ggml-org/llama.cpp/actions/runs/37614260360/job/112768687955
Edit: Guess you'll have to return {} directly.

Ha, Thanks! I flagged this exact one on another PR and still walked right into it. Fixed: the lambdas in step35 and the qwen35 family captured the structured bindings, so the models now read the flags into plain variables, and the helper returns them directly.

Pretty sure you can still do const auto [trunk_flags, mtp_flags] = nextn_flags(ml);.

@ServeurpersoCom

Copy link
Copy Markdown
Contributor Author

Pretty sure you can still do const auto [trunk_flags, mtp_flags] = nextn_flags(ml);.

The binding itself is what GCC 15 rejects once a lambda uses it: in step35 and the qwen35 family mtp_flags is read inside a [&] lambda, and the log points at that capture, not at the return. I kept plain variables in all 16 models so they stay uniform.

@CISC

CISC commented Oct 7, 2026

Copy link
Copy Markdown
Member

Pretty sure you can still do const auto [trunk_flags, mtp_flags] = nextn_flags(ml);.

The binding itself is what GCC 15 rejects once a lambda uses it: in step35 and the qwen35 family mtp_flags is read inside a [&] lambda, and the log points at that capture, not at the return. I kept plain variables in all 16 models so they stay uniform.

Ahhh, I see.

@CISC CISC left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like llama-install will be hogging the runners for the next 24 hours. :P

@ggerganov

Copy link
Copy Markdown
Member

I think it's OK to merge - this is unlikely to affect the AMD-based workflows.

@ServeurpersoCom
ServeurpersoCom merged commit 448147d into ggml-org:master Oct 7, 2026
17 of 18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model Model specific

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants