Skip to content

common: record NaN logits instead of only aborting, with opt-in continuation (engine#315) - #93

Merged
bong-water-water-bong merged 1 commit into
1bit/cxx26-cifrom
1bit/hrx-nan-continue
Oct 6, 2026
Merged

bong-water-water-bong merged 1 commit into
1bit/cxx26-cifrom
1bit/hrx-nan-continue

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

Fixes the evidence loss in engine#315 / engine#301.

Problem

The HRX NaN guard refuses to sample a silently-wrong token — correct — but it calls GGML_ABORT on the first NaN. An intermittent fault therefore kills the run along with the evidence, and every recorded occurrence reports only vocab index 0 because that is merely the first index it checks.

Measured impact: 1 fault in 5 runs on GLM-4.7-Flash (25 of 50 mixed requests lost), and the same abort killed a full ZAYA1-8B Markovian RSA evaluation after 494 completions with 0 of 30 problems recorded.

The known cause is already fixed in the pinned tree (the Loom GFX11 wave64 lane-mask backport f1b558f191, parent of hrx-system 98d05d94), so this is a different defect. Narrowed so far: not a memory ceiling (survivors reached higher GTT), not a fixed workload step (arbitrary point, no preceding anomaly), and the corruption always includes index 0 — a torn-page signature rather than scattered bad elements.

Change

Measure the corruption before acting on it: count, index range, and whether the run is contiguous. That is the discriminator between a migrated/torn page and a single bad element — the fact an earlier probe could not obtain because the process died holding it.

With GGML_HRX_NAN_CONTINUE set, substitute -inf for the NaN entries and continue. -inf makes those entries unsampleable (exactly what masking already uses) while leaving the rest of the distribution intact, so a long evaluation completes instead of dying at the first fault.

Off by default — the guard's existing hard-abort behaviour is unchanged unless an operator opts in.

What this does and does not do

  • Does: preserves the evidence that localises the fault, and lets long runs finish under an env var.
  • Does not: fix the root cause. That needs the NaN pattern this records (contiguous page vs single element), which is why the recording half matters even to someone who never sets the continuation flag.

…nuation

engine#315 measured 1 fault in 5 runs on GLM-4.7-Flash and killed a full
ZAYA1-8B evaluation with 0 of 30 problems recorded. The guard below refuses to
sample a silently-wrong token, which is correct, but it aborts on the FIRST NaN
-- so the run and the evidence die together, and every recorded occurrence so
far reports only 'vocab index 0'.

Measure the corruption before acting on it: count, index range, and whether the
run is contiguous. That is the discriminator between a migrated/torn page and a
single bad element, and it is the fact an earlier probe could not obtain.

With GGML_HRX_NAN_CONTINUE set, substitute -inf for the NaN entries and
continue. -inf makes those entries unsampleable -- exactly what masking already
uses -- while leaving the rest of the distribution intact, so a long evaluation
completes instead of dying at the first fault. Off by default, so the guard's
existing behaviour is unchanged unless an operator asks for it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant