Name and Version
Hit on build b10782 (commit 0ba6499c); the affected line is still present on current master (95ef7fc1, 2026-09-03): src/models/nemotron-h.cpp#L148
Operating systems
Linux (crash is platform-independent — integer division by zero at model load)
Which llama.cpp modules do you know to be affected?
libllama (core library)
Command line
llama-cli -m Nemotron-3-Puzzle-75B-A9B-Q4_K_M-00001-of-00002.gguf -p "hi"
Problem description & steps to reproduce
#25444 added nemotron_h_moe (NemotronHPuzzle) with genuinely per-layer expert_used_count / expert_feed_forward_length arrays that legitimately hold 0 on non-MoE layers. #28323 fixes the two loader checks that read hparams.n_expert_used() (= layer 0) and error out or hit GGML_ASSERT on this family. There is one more layer-0-style bug in the same family that #28323 does not cover:
In llama_model_nemotron_h::load_arch_tensors, the NextN/MTP tail loop computes:
const int64_t n_ff_exp = hparams.n_ff_exp(i) ? (int64_t)hparams.n_ff_exp(i) : n_ff / (int64_t)hparams.n_expert_used(i);
For a predict-layer index i where both per-layer arrays hold 0 (expert_feed_forward_length[i] == 0 and expert_used_count[i] == 0), the fallback is an integer division by zero → SIGFPE, killing the process instead of failing with an error message. The loop runs whenever the GGUF declares nextn_predict_layers > 0, even when MTP is not being loaded (load_mtp == false only sets TENSOR_SKIP; the shape expressions are still evaluated).
Reachability. The current in-tree converter cannot emit such a file (the Puzzle class drops the MTP head unconditionally, and the non-Puzzle NemotronH class writes a scalar expert_used_count), but real files in the wild do trigger it: pre-merge community conversions of the Puzzle FP8 checkpoint that kept the MTP head, e.g. YanissAmz/Nemotron-3-Puzzle-75B-A9B-GGUF (Q4_K_M), whose blk.88 (attention half of the split predict layer) has zeros at index 88 in both arrays. More generally, any GGUF with nemotron_h_moe, nextn_predict_layers > 0, and a zero at an MTP index crashes the process at load — same class of load-time SIGFPE-on-unusual-metadata as #26366.
Steps to reproduce (with the fixes from #28323 applied — without them the load stops earlier at model has expert layers but no expert layers are used):
- Download the Q4_K_M shards from the HF repo above.
llama-cli -m <shard 1> → SIGFPE in load_arch_tensors at the line quoted above, first predict-layer iteration (i = 88).
(Note for anyone reproducing with that specific file: even with the division guarded, the file still fails cleanly later with a missing-tensor error, because it predates the merge and splits each predict layer across two blocks — mainline expects both halves per block and requires TENSOR_SKIP tensors to exist. That part is a pre-merge-layout compatibility gap in the community file, not a mainline bug; the SIGFPE is the mainline bug.)
Suggested fix. Guard the fallback the same way the trunk loop's equivalent expression should also be guarded, e.g.:
const int64_t n_ff_exp = hparams.n_ff_exp(i) ? (int64_t)hparams.n_ff_exp(i)
: hparams.n_expert_used(i) ? n_ff / (int64_t)hparams.n_expert_used(i)
: 0;
or validate up front and throw std::runtime_error so a malformed GGUF produces an error instead of a crash. Happy to turn this into a PR if preferred — verified locally that with this guard (plus fixes equivalent to #28323) the load proceeds past this point.
Disclosure: this report was drafted with AI assistance (Claude); the crash, root cause, and the verification against master are from a real debugging session on b10782.
First Bad Commit
c61b98b8 (#25444, b10776) — the line was introduced with the architecture.
Relevant log output
Floating point exception (core dumped)
# SIGFPE, no llama error output; faulting statement is the n_ff / n_expert_used(i)
# fallback in llama_model_nemotron_h::load_arch_tensors, first MTP-tail iteration.
Name and Version
Hit on build b10782 (commit
0ba6499c); the affected line is still present on current master (95ef7fc1, 2026-09-03): src/models/nemotron-h.cpp#L148Operating systems
Linux (crash is platform-independent — integer division by zero at model load)
Which llama.cpp modules do you know to be affected?
libllama (core library)
Command line
llama-cli -m Nemotron-3-Puzzle-75B-A9B-Q4_K_M-00001-of-00002.gguf -p "hi"Problem description & steps to reproduce
#25444 added
nemotron_h_moe(NemotronHPuzzle) with genuinely per-layerexpert_used_count/expert_feed_forward_lengtharrays that legitimately hold 0 on non-MoE layers. #28323 fixes the two loader checks that readhparams.n_expert_used()(= layer 0) and error out or hitGGML_ASSERTon this family. There is one more layer-0-style bug in the same family that #28323 does not cover:In
llama_model_nemotron_h::load_arch_tensors, the NextN/MTP tail loop computes:For a predict-layer index
iwhere both per-layer arrays hold 0 (expert_feed_forward_length[i] == 0andexpert_used_count[i] == 0), the fallback is an integer division by zero → SIGFPE, killing the process instead of failing with an error message. The loop runs whenever the GGUF declaresnextn_predict_layers > 0, even when MTP is not being loaded (load_mtp == falseonly setsTENSOR_SKIP; the shape expressions are still evaluated).Reachability. The current in-tree converter cannot emit such a file (the Puzzle class drops the MTP head unconditionally, and the non-Puzzle NemotronH class writes a scalar
expert_used_count), but real files in the wild do trigger it: pre-merge community conversions of the Puzzle FP8 checkpoint that kept the MTP head, e.g. YanissAmz/Nemotron-3-Puzzle-75B-A9B-GGUF (Q4_K_M), whoseblk.88(attention half of the split predict layer) has zeros at index 88 in both arrays. More generally, any GGUF withnemotron_h_moe,nextn_predict_layers > 0, and a zero at an MTP index crashes the process at load — same class of load-time SIGFPE-on-unusual-metadata as #26366.Steps to reproduce (with the fixes from #28323 applied — without them the load stops earlier at
model has expert layers but no expert layers are used):llama-cli -m <shard 1>→ SIGFPE inload_arch_tensorsat the line quoted above, first predict-layer iteration (i = 88).(Note for anyone reproducing with that specific file: even with the division guarded, the file still fails cleanly later with a missing-tensor error, because it predates the merge and splits each predict layer across two blocks — mainline expects both halves per block and requires
TENSOR_SKIPtensors to exist. That part is a pre-merge-layout compatibility gap in the community file, not a mainline bug; the SIGFPE is the mainline bug.)Suggested fix. Guard the fallback the same way the trunk loop's equivalent expression should also be guarded, e.g.:
or validate up front and
throw std::runtime_errorso a malformed GGUF produces an error instead of a crash. Happy to turn this into a PR if preferred — verified locally that with this guard (plus fixes equivalent to #28323) the load proceeds past this point.Disclosure: this report was drafted with AI assistance (Claude); the crash, root cause, and the verification against master are from a real debugging session on b10782.
First Bad Commit
c61b98b8(#25444, b10776) — the line was introduced with the architecture.Relevant log output