Repository navigation
common: record NaN logits instead of only aborting, with opt-in continuation (engine#315) - #93
Merged
Conversation
…nuation engine#315 measured 1 fault in 5 runs on GLM-4.7-Flash and killed a full ZAYA1-8B evaluation with 0 of 30 problems recorded. The guard below refuses to sample a silently-wrong token, which is correct, but it aborts on the FIRST NaN -- so the run and the evidence die together, and every recorded occurrence so far reports only 'vocab index 0'. Measure the corruption before acting on it: count, index range, and whether the run is contiguous. That is the discriminator between a migrated/torn page and a single bad element, and it is the fact an earlier probe could not obtain. With GGML_HRX_NAN_CONTINUE set, substitute -inf for the NaN entries and continue. -inf makes those entries unsampleable -- exactly what masking already uses -- while leaving the rest of the distribution intact, so a long evaluation completes instead of dying at the first fault. Off by default, so the guard's existing behaviour is unchanged unless an operator asks for it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the evidence loss in engine#315 / engine#301.
Problem
The HRX NaN guard refuses to sample a silently-wrong token — correct — but it calls
GGML_ABORTon the first NaN. An intermittent fault therefore kills the run along with the evidence, and every recorded occurrence reports onlyvocab index 0because that is merely the first index it checks.Measured impact: 1 fault in 5 runs on GLM-4.7-Flash (25 of 50 mixed requests lost), and the same abort killed a full ZAYA1-8B Markovian RSA evaluation after 494 completions with 0 of 30 problems recorded.
The known cause is already fixed in the pinned tree (the Loom GFX11 wave64 lane-mask backport
f1b558f191, parent ofhrx-system 98d05d94), so this is a different defect. Narrowed so far: not a memory ceiling (survivors reached higher GTT), not a fixed workload step (arbitrary point, no preceding anomaly), and the corruption always includes index 0 — a torn-page signature rather than scattered bad elements.Change
Measure the corruption before acting on it:
count, indexrange, and whether the run iscontiguous. That is the discriminator between a migrated/torn page and a single bad element — the fact an earlier probe could not obtain because the process died holding it.With
GGML_HRX_NAN_CONTINUEset, substitute-inffor the NaN entries and continue.-infmakes those entries unsampleable (exactly what masking already uses) while leaving the rest of the distribution intact, so a long evaluation completes instead of dying at the first fault.Off by default — the guard's existing hard-abort behaviour is unchanged unless an operator opts in.
What this does and does not do