Repository navigation
llama: share the nextn tensor flags between models - #30097
Conversation
Follow-up of the TODO in glm5-next: move the trunk-only and MTP-only detection that each model copied into a nextn_flags helper of llama_model_base. It probes the first trunk layer and the first NextN layer, and adds TENSOR_SKIP when MTP is not loaded. qwen4exp probes hc_attn_norm since it has no attn_norm. deepseek4, nemotron-h, qwen35, qwen35moe, qwen3next and qwen4exp now also accept a trunk-only file, like the other models.
|
Huh, GCC15 didn't like that: Edit: Guess you'll have to |
Lambdas that capture structured bindings need C++20, and GCC 15 rejects them under -Werror, so the models read the trunk and MTP flags into plain variables.
Ha, Thanks! I flagged this exact one on another PR and still walked right into it. Fixed: the lambdas in step35 and the qwen35 family captured the structured bindings, so the models now read the flags into plain variables, and the helper returns them directly. |
Pretty sure you can still do |
The binding itself is what GCC 15 rejects once a lambda uses it: in step35 and the qwen35 family mtp_flags is read inside a [&] lambda, and the log points at that capture, not at the return. I kept plain variables in all 16 models so they stay uniform. |
Ahhh, I see. |
CISC
left a comment
There was a problem hiding this comment.
Looks like llama-install will be hogging the runners for the next 24 hours. :P
|
I think it's OK to merge - this is unlikely to affect the AMD-based workflows. |
Overview
Follow-up of the TODO in glm5-next: move the trunk-only and MTP-only detection that each model copied into a nextn_flags helper of llama_model_base. It probes the first trunk layer and the first NextN layer, and adds TENSOR_SKIP when MTP is not loaded. qwen4exp probes hc_attn_norm since it has no attn_norm.
deepseek4, nemotron-h, qwen35, qwen35moe, qwen3next and qwen4exp now also accept a trunk-only file, like the other models.
Additional information
Follow-up #29928
Requirements