Skip to content

hrx: fence both global and LDS at the decode-split barriers (engine#123 follow-up) - #29

Merged
bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-decode-split-barriers-both-spaces
Sep 27, 2026
Merged

bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-decode-split-barriers-both-spaces

Conversation

@bong-water-water-bong

@bong-water-water-bong bong-water-water-bong commented Sep 27, 2026 •

Copy link
Copy Markdown

Follow-up to #28 (engine#123 / ggml-org#140).

Two changes

1. Pack sites: fence LDS again, not only global. A Loom kernel.barrier fences only the memory space it names (loom docs, Synchronization names both rendezvous and memory). #28 replaced the barrier<workgroup> before pack_completed_q8 with barrier<global>. That orders the reduce's global output stores, but it drops the LDS fence the cooperative and multipass reducers had for their staging buffers. All 5 pack sites now use both, i.e. HIP __syncthreads semantics:

kernel.barrier<global> scope(workgroup) ordering(acq_rel)
kernel.barrier<workgroup> scope(workgroup) ordering(acq_rel)

2. Standalone two-dispatch reducer: same bug as #28, now fixed. In ggml_flash_attention_decode_split_reduce_f32 and its qwen3_moe copy, wave 0 rewrites partial_max in global memory with the exp scales. Every wave then reads it back after a barrier<workgroup> only. Wave 0 had read those same lines during the max pass, so other waves could hit stale GL1/L0 copies (the block maxima instead of the scales). This kernel is not dispatched today; it gets the <global> barrier too.

Compiled ISA (gfx1151)

  • Pack handoff, from the JIT dumps of the real engine run: direct (short ctx), cooperative (~600 tokens) and multipass (30B, 2113 tokens) all compile to vmcnt(0), vscnt(0), s_barrier, buffer_gl1_inv, buffer_gl0_inv, s_barrier, then the pack global_load_b128.
  • Standalone reducer, via loom-compile: before, 2 barriers and no invalidates; after, 3 barriers with buffer_gl1_inv + buffer_gl0_inv.

Tests

result
Qwen3-0.6B, 24-token greedy, with a process forcing KFD queue evictions 0 of 118 divergent (only the warm-up request differs)
Qwen3-Coder-30B, 2113-token prompt, GGML_HRX_FA_PARTIAL_ALIGN=256, evictions 8 of 8 correct, 0 NaN
Standalone reducer, ad-hoc on-device check case run with iree-test-loom (scale 1 vs stale max 0; not committed), 2000 runs before the fix 0 failures, so the race is not reproducible there; the fix rests on the ISA and the memory model

Decode speed vs #28 is within run-to-run noise on this shared box (2 interleaved rounds):

tok/s #28 this PR
0.6B d0 296–300 276–287
0.6B d1000 199–207 204–215
30B d2100 58.9–62.0 57.0–60.6

🤖 Generated with Claude Code

…23 follow-up)

A Loom kernel.barrier fences only the memory space it names. #28 changed the
barrier before pack_completed_q8 from <workgroup> to <global>, which orders the
reduce's global output stores but drops the LDS fence the cooperative and
multipass reducers relied on for their staging buffers. Keep both: a
<global> barrier followed by the original <workgroup> barrier (HIP
__syncthreads semantics) at the five pack sites.

The standalone two-dispatch reducer (ggml_flash_attention_decode_split_reduce_f32
and its qwen3_moe copy) had the same bug as #28: wave 0 rewrites partial_max in
global memory with the exp scales and every wave reads it back after a
<workgroup>-only barrier, so the reads could hit stale GL1/L0 lines holding the
block maxima. It is not dispatched today. Add the <global> barrier there too.

Compiled ISA, gfx1151: every pack handoff (direct, cooperative, multipass) is now
vmcnt(0), vscnt(0), s_barrier, buffer_gl1_inv, buffer_gl0_inv, s_barrier; the
standalone reducer gains buffer_gl1_inv/buffer_gl0_inv (it had none).

Measured with a separate process forcing KFD queue evictions: Qwen3-0.6B 0 of
118 divergent; Qwen3-Coder-30B, 2113-token prompt, partial alignment 256: 8 of 8
correct, 0 NaN. Decode speed vs #28 within run-to-run noise (0.6B d0 276-287 vs
296-300, d1000 204-215 vs 199-207; 30B d2100 57-61 vs 59-62 tok/s).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the ggml label Sep 27, 2026
@bong-water-water-bong
bong-water-water-bong merged commit 00adc2b into 1bit/hrx-vulkan-patched Sep 27, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant