Repository navigation
hrx: order the decode-split q8 pack after the reduce's global output stores (engine#123, #140) - #28
Merged
bong-water-water-bong merged 1 commit intoSep 27, 2026
Conversation
…stores (engine#123, ggml-org#140) Every decode-split reduce_fused variant (direct, cooperative, multipass, in both the ggml.* and qwen3_moe corpora) writes the reduced attention output to global memory and then packs it into next_q8 with pack_completed_q8, which reads that output back with a different wave/lane mapping. The barrier between them was kernel.barrier<workgroup>, which orders workgroup (LDS) memory only; it compiled to s_waitcnt lgkmcnt(0); s_barrier, with no store-completion wait before it and no buffer_gl0_inv after it. A pack wave could therefore read output before another wave's stores landed, or from a stale L0 line (WGP mode, or after a CWSR restore onto another CU), and quantize the arena slot's previous contents. The f32 output stays correct but next_q8, which the following matmul consumes, is wrong, and NaN when the stale bytes decode as NaN. That is the source of the NaN router logits behind engine#123 and of the run-to-run divergence in engine#140. Queue preemption from any process's KFD eviction opens the window, which is why both issues tracked page migration and GPU contention. Use kernel.barrier<global> scope(workgroup) ordering(acq_rel) at the five sites. Measured on gfx1151 with a separate process that only forces KFD queue evictions: - Qwen3-0.6B, 24-token greedy: 0 of 354 divergent (before: 13-34 per ~177) - Qwen3-0.6B, forced compaction + mlockall: 0 of 118 (before: 42-62 %) - Qwen3-Coder-30B, 2113-token prompt, GGML_HRX_FA_PARTIAL_ALIGN=256: 16 of 16 correct, 0 NaN (before: 1 of 8 correct, NaN guard fired 7 times) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
merged commit Sep 27, 2026
9ca9603
into
1bit/hrx-vulkan-patched
1 check passed
bong-water-water-bong
pushed a commit
to 1bit-MONSTER/engine
that referenced
this pull request
Sep 27, 2026
…bal barrier (#123, #140) Brings 1bit-MONSTER/llama.cpp#28 and #29: in every decode-split reduce_fused variant the barrier between the reduce's global output stores and pack_completed_q8's output loads was kernel.barrier<workgroup> (LDS only), so the next_q8 copy could be packed from stale output. It is now kernel.barrier<global> followed by the LDS barrier (5 sites, ops and qwen3_moe corpora; #29 restores the LDS fence #28 dropped and adds the missing global barrier to the undispatched standalone reduce_f32 kernels). The pin moves fa226f9 -> 00adc2b, fast-forward. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
added a commit
to 1bit-MONSTER/engine
that referenced
this pull request
Sep 27, 2026
…bal barrier (#123, #140) (#176) Brings 1bit-MONSTER/llama.cpp#28 and #29: in every decode-split reduce_fused variant the barrier between the reduce's global output stores and pack_completed_q8's output loads was kernel.barrier<workgroup> (LDS only), so the next_q8 copy could be packed from stale output. It is now kernel.barrier<global> followed by the LDS barrier (5 sites, ops and qwen3_moe corpora; #29 restores the LDS fence #28 dropped and adds the missing global barrier to the undispatched standalone reduce_f32 kernels). The pin moves fa226f9 -> 00adc2b, fast-forward. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Root cause of 1bit-MONSTER/engine#123 (NaN router logits → wild expert read) and 1bit-MONSTER/engine#140 (run-to-run HRX0 divergence).
Bug
Every decode-split
reduce_fusedvariant (direct, cooperative, multipass, in bothloom-libs/opsandqwen_moe/qwen3_moeflash_attention_decode_split_f32_f16_wmma.loom) writes the reduced output to global memory and then runspack_completed_q8, which reads that output back with a different wave/lane mapping. The barrier between the two was:It compiles to
s_waitcnt lgkmcnt(0); s_barrier, with nos_waitcnt_vscnt 0before it and nobuffer_gl0_invafter it. A pack wave can therefore readoutputbefore another wave's stores land, or from a stale L0 line (WGP mode, or after a CWSR restore onto another CU), and quantize the arena slot's previous contents intonext_q8. The f32 output is correct; the q8 copy the next matmul consumes is not, and it is NaN when the stale bytes decode as NaN. Any process's KFD queue eviction preempts in-flight dispatches and opens the window. That is why both issues tracked page migration and GPU contention, and why arena alignment changed the fault rate.Fix
kernel.barrier<global> scope(workgroup) ordering(acq_rel)at the five sites. 5 lines, no other change.How it was found
hipHostRegister+mprotectchurn, no kernels) reproduces the corruption in HRX. The HIP llama.cpp backend is immune, and plain HIP kernels keep their register/LDS/WMMA state across the same preemptions.output, differentq8_output.Measured (gfx1151, strixhalo)
mlockallGGML_HRX_FA_PARTIAL_ALIGN=256Not in this PR
The standalone two-dispatch
*_decode_split_reduce_f32kernels rewritepartial_maxin place across the same kind of LDS-only barrier. They are not dispatched today, so this PR leaves them alone, since it can't test them end to end.🤖 Generated with Claude Code