Skip to content

hrx: order the decode-split q8 pack after the reduce's global output stores (engine#123, #140) - #28

Merged
bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-decode-split-pack-barrier
Sep 27, 2026
Merged

bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-decode-split-pack-barrier

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

Root cause of 1bit-MONSTER/engine#123 (NaN router logits → wild expert read) and 1bit-MONSTER/engine#140 (run-to-run HRX0 divergence).

Bug

Every decode-split reduce_fused variant (direct, cooperative, multipass, in both loom-libs/ops and qwen_moe/qwen3_moe flash_attention_decode_split_f32_f16_wmma.loom) writes the reduced output to global memory and then runs pack_completed_q8, which reads that output back with a different wave/lane mapping. The barrier between the two was:

kernel.barrier<workgroup> scope(workgroup) ordering(acq_rel)   // LDS ordering only

It compiles to s_waitcnt lgkmcnt(0); s_barrier, with no s_waitcnt_vscnt 0 before it and no buffer_gl0_inv after it. A pack wave can therefore read output before another wave's stores land, or from a stale L0 line (WGP mode, or after a CWSR restore onto another CU), and quantize the arena slot's previous contents into next_q8. The f32 output is correct; the q8 copy the next matmul consumes is not, and it is NaN when the stale bytes decode as NaN. Any process's KFD queue eviction preempts in-flight dispatches and opens the window. That is why both issues tracked page migration and GPU contention, and why arena alignment changed the fault rate.

Fix

kernel.barrier<global> scope(workgroup) ordering(acq_rel) at the five sites. 5 lines, no other change.

How it was found

  • A second process that only forces KFD queue evictions (hipHostRegister + mprotect churn, no kernels) reproduces the corruption in HRX. The HIP llama.cpp backend is immune, and plain HIP kernels keep their register/LDS/WMMA state across the same preemptions.
  • Serial execution with every command's bindings hashed: good runs are bit-identical (0 of 32,559 lines differ). In all 4 corrupted requests the first differing command is this kernel, with identical inputs, identical f32 output, different q8_output.

Measured (gfx1151, strixhalo)

test before after
Qwen3-0.6B, 24-token greedy, foreign evictions 13–34 divergent per ~177 requests (3 runs) 0 / 354
Qwen3-0.6B, forced compaction + mlockall 42–62 % 0 / 118
Qwen3-Coder-30B, 2113-token prompt, GGML_HRX_FA_PARTIAL_ALIGN=256 1/8 correct, NaN guard fired 7× 16/16 correct, 0 NaN

Not in this PR

The standalone two-dispatch *_decode_split_reduce_f32 kernels rewrite partial_max in place across the same kind of LDS-only barrier. They are not dispatched today, so this PR leaves them alone, since it can't test them end to end.

🤖 Generated with Claude Code

…stores (engine#123, ggml-org#140)

Every decode-split reduce_fused variant (direct, cooperative, multipass, in both the
ggml.* and qwen3_moe corpora) writes the reduced attention output to global memory and
then packs it into next_q8 with pack_completed_q8, which reads that output back with a
different wave/lane mapping. The barrier between them was kernel.barrier<workgroup>,
which orders workgroup (LDS) memory only; it compiled to s_waitcnt lgkmcnt(0);
s_barrier, with no store-completion wait before it and no buffer_gl0_inv after it. A
pack wave could therefore read output before another wave's stores landed, or from a
stale L0 line (WGP mode, or after a CWSR restore onto another CU), and quantize the
arena slot's previous contents. The f32 output stays correct but next_q8, which the
following matmul consumes, is wrong, and NaN when the stale bytes decode as NaN. That
is the source of the NaN router logits behind engine#123 and of the run-to-run
divergence in engine#140. Queue preemption from any process's KFD eviction opens the
window, which is why both issues tracked page migration and GPU contention.

Use kernel.barrier<global> scope(workgroup) ordering(acq_rel) at the five sites.

Measured on gfx1151 with a separate process that only forces KFD queue evictions:
- Qwen3-0.6B, 24-token greedy: 0 of 354 divergent (before: 13-34 per ~177)
- Qwen3-0.6B, forced compaction + mlockall: 0 of 118 (before: 42-62 %)
- Qwen3-Coder-30B, 2113-token prompt, GGML_HRX_FA_PARTIAL_ALIGN=256: 16 of 16
  correct, 0 NaN (before: 1 of 8 correct, NaN guard fired 7 times)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the ggml label Sep 27, 2026
@bong-water-water-bong
bong-water-water-bong merged commit 9ca9603 into 1bit/hrx-vulkan-patched Sep 27, 2026
1 check passed
bong-water-water-bong pushed a commit to 1bit-MONSTER/engine that referenced this pull request Sep 27, 2026
…bal barrier (#123, #140)

Brings 1bit-MONSTER/llama.cpp#28 and #29: in every decode-split reduce_fused variant the barrier
between the reduce's global output stores and pack_completed_q8's output loads was
kernel.barrier<workgroup> (LDS only), so the next_q8 copy could be packed from stale output.
It is now kernel.barrier<global> followed by the LDS barrier (5 sites, ops and qwen3_moe
corpora; #29 restores the LDS fence #28 dropped and adds the missing global barrier to
the undispatched standalone reduce_f32 kernels). The pin moves
fa226f9 -> 00adc2b, fast-forward.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit to 1bit-MONSTER/engine that referenced this pull request Sep 27, 2026
…bal barrier (#123, #140) (#176)

Brings 1bit-MONSTER/llama.cpp#28 and #29: in every decode-split reduce_fused variant the barrier
between the reduce's global output stores and pack_completed_q8's output loads was
kernel.barrier<workgroup> (LDS only), so the next_q8 copy could be packed from stale output.
It is now kernel.barrier<global> followed by the LDS barrier (5 sites, ops and qwen3_moe
corpora; #29 restores the LDS fence #28 dropped and adds the missing global barrier to
the undispatched standalone reduce_f32 kernels). The pin moves
fa226f9 -> 00adc2b, fast-forward.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant