Skip to content

src : add n_expert_used_max function - #28323

Merged
danbev merged 3 commits into
ggml-org:masterfrom
danbev:n_experts_used_max
Sep 4, 2026
Merged

danbev merged 3 commits into
ggml-org:masterfrom
danbev:n_experts_used_max

Conversation

@danbev

@danbev danbev commented Sep 3, 2026

Copy link
Copy Markdown
Member

Overview

With Commit c61b98b ("model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444)") it is now possible for each layer to have a specific number of experts but there are a few checks that need to be updated to handle this upon model loading. For example:

llama_model_load: error loading model: model has expert layers but no expert layers are used

And later:

/llama.cpp/src/llama-model-loader.cpp:955: GGML_ASSERT(n_ids_used > 0) failed

This commit adds the n_expert_used_max function so that these checks can use it.

Additional information

Refs: #25444 (comment)

Requirements

With Commit c61b98b ("model: add
NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (ggml-org#25444)") it
is now possible for each layer to have a specific number of experts but
there are a few checks that need to be updated to handle this upon model
loading. For example:
```console
llama_model_load: error loading model: model has expert layers but no expert layers are used
```
And later:
```console
/llama.cpp/src/llama-model-loader.cpp:955: GGML_ASSERT(n_ids_used > 0) failed
```

This commit adds the n_expert_used_max function so that these checks
can use it.

Refs: ggml-org#25444 (comment)
@bartowski1182

Copy link
Copy Markdown
Contributor

Fixes the issue on my end!

@CISC CISC left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also replace this one:

llama.cpp/src/llama-model.cpp

Lines 1256 to 1260 in 7dc826e

// models may route a different number of experts per layer, so validate the maximum
uint32_t n_expert_used_max = 0;
for (uint32_t il = 0; il < hparams.n_layer_all; ++il) {
n_expert_used_max = std::max(n_expert_used_max, hparams.n_expert_used(il));
}

Comment thread src/llama-hparams.cpp Outdated
@bartowski1182

Copy link
Copy Markdown
Contributor

Still works on my end, just tested and it fixes the nemotron issue

@bartowski1182

Copy link
Copy Markdown
Contributor

any concerns with merging @CISC ?

@danbev
danbev merged commit 9a4843c into ggml-org:master Sep 4, 2026
24 of 26 checks passed
@danbev
danbev deleted the n_experts_used_max branch September 4, 2026 04:37
@CISC

CISC commented Sep 4, 2026

Copy link
Copy Markdown
Member

any concerns with merging @CISC ?

Nope. 😆

MarkShark2 added a commit to MarkShark2/llama.cpp that referenced this pull request Sep 5, 2026
64 upstream commits. The bulk of the 57 conflicted files is one seam: upstream
adopted the per-layer n_ff_exp_arr / n_expert_used_arr accessors and dropped
the scalar _impl fallbacks (ggml-org#28323), so those hunks take upstream's side.
Hand-resolved: the DSV4 media routing (upstream's exp_probs_b_vl path kept,
the fork's routing_id hash gather kept for palette batches), build_attn_mha's
new n_kv_max argument in the fork's sparse-attention else branch, the loader's
parallel RPC population kept under upstream's non-host-first ordering, the
Puzzle converter keeping its MTP export plus upstream's name normalization.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
x1250 pushed a commit to x1250/llama.cpp that referenced this pull request Sep 9, 2026
* src : add n_expert_used_max function

With Commit c61b98b ("model: add
NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (ggml-org#25444)") it
is now possible for each layer to have a specific number of experts but
there are a few checks that need to be updated to handle this upon model
loading. For example:
```console
llama_model_load: error loading model: model has expert layers but no expert layers are used
```
And later:
```console
/llama.cpp/src/llama-model-loader.cpp:955: GGML_ASSERT(n_ids_used > 0) failed
```

This commit adds the n_expert_used_max function so that these checks
can use it.

Refs: ggml-org#25444 (comment)

* src : use hparams.n_expert_used_max in llama_model_base::load_hparams

* src : use 0 as initial value for n_expert_used_max
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* src : add n_expert_used_max function

With Commit 3615d1b ("model: add
NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (ggml-org#25444)") it
is now possible for each layer to have a specific number of experts but
there are a few checks that need to be updated to handle this upon model
loading. For example:
```console
llama_model_load: error loading model: model has expert layers but no expert layers are used
```
And later:
```console
/llama.cpp/src/llama-model-loader.cpp:955: GGML_ASSERT(n_ids_used > 0) failed
```

This commit adds the n_expert_used_max function so that these checks
can use it.

Refs: ggml-org#25444 (comment)

* src : use hparams.n_expert_used_max in llama_model_base::load_hparams

* src : use 0 as initial value for n_expert_used_max
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* src : add n_expert_used_max function

With Commit c61b98b ("model: add
NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (ggml-org#25444)") it
is now possible for each layer to have a specific number of experts but
there are a few checks that need to be updated to handle this upon model
loading. For example:
```console
llama_model_load: error loading model: model has expert layers but no expert layers are used
```
And later:
```console
/llama.cpp/src/llama-model-loader.cpp:955: GGML_ASSERT(n_ids_used > 0) failed
```

This commit adds the n_expert_used_max function so that these checks
can use it.

Refs: ggml-org#25444 (comment)

* src : use hparams.n_expert_used_max in llama_model_base::load_hparams

* src : use 0 as initial value for n_expert_used_max
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
* src : add n_expert_used_max function

With Commit c61b98b ("model: add
NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (ggml-org#25444)") it
is now possible for each layer to have a specific number of experts but
there are a few checks that need to be updated to handle this upon model
loading. For example:
```console
llama_model_load: error loading model: model has expert layers but no expert layers are used
```
And later:
```console
/llama.cpp/src/llama-model-loader.cpp:955: GGML_ASSERT(n_ids_used > 0) failed
```

This commit adds the n_expert_used_max function so that these checks
can use it.

Refs: ggml-org#25444 (comment)

* src : use hparams.n_expert_used_max in llama_model_base::load_hparams

* src : use 0 as initial value for n_expert_used_max
Te-eMster pushed a commit to Te-eMster/mx-llama.cpp that referenced this pull request Sep 18, 2026
* src : add n_expert_used_max function

With Commit c61b98b ("model: add
NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (ggml-org#25444)") it
is now possible for each layer to have a specific number of experts but
there are a few checks that need to be updated to handle this upon model
loading. For example:
```console
llama_model_load: error loading model: model has expert layers but no expert layers are used
```
And later:
```console
/llama.cpp/src/llama-model-loader.cpp:955: GGML_ASSERT(n_ids_used > 0) failed
```

This commit adds the n_expert_used_max function so that these checks
can use it.

Refs: ggml-org#25444 (comment)

* src : use hparams.n_expert_used_max in llama_model_base::load_hparams

* src : use 0 as initial value for n_expert_used_max
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants