Skip to content

models : consolidate nextn row cropping into shared helpers - #30017

Merged
ggerganov merged 2 commits into
masterfrom
gg/out-ids-consolidate
Oct 6, 2026
Merged

ggerganov merged 2 commits into
masterfrom
gg/out-ids-consolidate

Conversation

@ggerganov

@ggerganov ggerganov commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Overview

  • replaces the duplicated nextn output-row crop conditions and the per-model flags (narrow_early, crop_before_ffn, crop_last_layer, emit_h_nextn) with two helpers on llm_graph_context: crop_before_nextn() / crop_after_nextn()
  • exactly one of them is true whenever inp_out_ids != nullptr, so the pair documents where the narrowing happens relative to the t_h_nextn capture
  • t_h_nextn is now set unconditionally in mimo2, qwen4exp and deepseek4, as in the other nextn-capable models; host reads stay gated by cparams.embeddings_nextn
  • the models that only tested embeddings_nextn_masked now share the same condition: with extraction off the last layer is cropped before the capture, which saves its FFN rows - outputs are unchanged
  • 20 files (src/llama-graph.h + 19 models)

Additional information

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. pi:llama.cpp/MiMo-V2.6-Flash-MOPD

the other nextn-capable models set it unconditionally

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
- replace the duplicated crop conditions and the per-model flags (narrow_early,
  crop_before_ffn, crop_last_layer, emit_h_nextn) with two helpers on llm_graph_context:
  crop_before_nextn() / crop_after_nextn()
- models that only tested embeddings_nextn_masked now share the same condition, so they
  crop the last layer before the nextn capture whenever extraction is off
- t_h_nextn is now set unconditionally in mimo2, qwen4exp and deepseek4 (as in the other
  nextn-capable models); host-side reads stay gated by cparams.embeddings_nextn

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
@ggerganov
ggerganov force-pushed the gg/out-ids-consolidate branch from 57f2156 to bba2180 Compare October 6, 2026 09:56
@ggerganov ggerganov changed the title models : consolidate nextn row cropping and make graph topology independent of embeddings_nextn models : consolidate nextn row cropping into shared helpers Oct 6, 2026
@ggerganov
ggerganov marked this pull request as ready for review October 6, 2026 10:28
@ggerganov
ggerganov requested a review from CISC as a code owner October 6, 2026 10:28

@ServeurpersoCom ServeurpersoCom left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tested on CUDA: ctest -L main passes, and GLM-4.5-Air greedy outputs are bit-identical to master both without and with draft-mtp (same 59.4% acceptance), with no reallocation under GGML_SCHED_DEBUG_REALLOC=1. The 8 models that now crop late by default show no measurable prefill cost (GLM-4.5-Air and GLM-4.7-Flash, pp512 and pp4096 within noise). LGTM.

@ggerganov
ggerganov merged commit f0c41e0 into master Oct 6, 2026
18 of 19 checks passed
@ggerganov
ggerganov deleted the gg/out-ids-consolidate branch October 6, 2026 11:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model Model specific

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants