Skip to content

Misc. bug: SIGFPE (integer divide-by-zero) loading nemotron_h_moe GGUFs whose NextN/MTP blocks have zeroed per-layer expert arrays #28345

Description

@loenhart

Name and Version

Hit on build b10782 (commit 0ba6499c); the affected line is still present on current master (95ef7fc1, 2026-09-03): src/models/nemotron-h.cpp#L148

Operating systems

Linux (crash is platform-independent — integer division by zero at model load)

Which llama.cpp modules do you know to be affected?

libllama (core library)

Command line

llama-cli -m Nemotron-3-Puzzle-75B-A9B-Q4_K_M-00001-of-00002.gguf -p "hi"

Problem description & steps to reproduce

#25444 added nemotron_h_moe (NemotronHPuzzle) with genuinely per-layer expert_used_count / expert_feed_forward_length arrays that legitimately hold 0 on non-MoE layers. #28323 fixes the two loader checks that read hparams.n_expert_used() (= layer 0) and error out or hit GGML_ASSERT on this family. There is one more layer-0-style bug in the same family that #28323 does not cover:

In llama_model_nemotron_h::load_arch_tensors, the NextN/MTP tail loop computes:

const int64_t n_ff_exp = hparams.n_ff_exp(i) ? (int64_t)hparams.n_ff_exp(i) : n_ff / (int64_t)hparams.n_expert_used(i);

For a predict-layer index i where both per-layer arrays hold 0 (expert_feed_forward_length[i] == 0 and expert_used_count[i] == 0), the fallback is an integer division by zero → SIGFPE, killing the process instead of failing with an error message. The loop runs whenever the GGUF declares nextn_predict_layers > 0, even when MTP is not being loaded (load_mtp == false only sets TENSOR_SKIP; the shape expressions are still evaluated).

Reachability. The current in-tree converter cannot emit such a file (the Puzzle class drops the MTP head unconditionally, and the non-Puzzle NemotronH class writes a scalar expert_used_count), but real files in the wild do trigger it: pre-merge community conversions of the Puzzle FP8 checkpoint that kept the MTP head, e.g. YanissAmz/Nemotron-3-Puzzle-75B-A9B-GGUF (Q4_K_M), whose blk.88 (attention half of the split predict layer) has zeros at index 88 in both arrays. More generally, any GGUF with nemotron_h_moe, nextn_predict_layers > 0, and a zero at an MTP index crashes the process at load — same class of load-time SIGFPE-on-unusual-metadata as #26366.

Steps to reproduce (with the fixes from #28323 applied — without them the load stops earlier at model has expert layers but no expert layers are used):

  1. Download the Q4_K_M shards from the HF repo above.
  2. llama-cli -m <shard 1> → SIGFPE in load_arch_tensors at the line quoted above, first predict-layer iteration (i = 88).

(Note for anyone reproducing with that specific file: even with the division guarded, the file still fails cleanly later with a missing-tensor error, because it predates the merge and splits each predict layer across two blocks — mainline expects both halves per block and requires TENSOR_SKIP tensors to exist. That part is a pre-merge-layout compatibility gap in the community file, not a mainline bug; the SIGFPE is the mainline bug.)

Suggested fix. Guard the fallback the same way the trunk loop's equivalent expression should also be guarded, e.g.:

const int64_t n_ff_exp = hparams.n_ff_exp(i) ? (int64_t)hparams.n_ff_exp(i)
                       : hparams.n_expert_used(i) ? n_ff / (int64_t)hparams.n_expert_used(i)
                       : 0;

or validate up front and throw std::runtime_error so a malformed GGUF produces an error instead of a crash. Happy to turn this into a PR if preferred — verified locally that with this guard (plus fixes equivalent to #28323) the load proceeds past this point.

Disclosure: this report was drafted with AI assistance (Claude); the crash, root cause, and the verification against master are from a real debugging session on b10782.

First Bad Commit

c61b98b8 (#25444, b10776) — the line was introduced with the architecture.

Relevant log output

Floating point exception (core dumped)
# SIGFPE, no llama error output; faulting statement is the n_ff / n_expert_used(i)
# fallback in llama_model_nemotron_h::load_arch_tensors, first MTP-tail iteration.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions