Repository navigation
hrx: fence both global and LDS at the decode-split barriers (engine#123 follow-up) - #29
Merged
bong-water-water-bong merged 1 commit intoSep 27, 2026
Conversation
…23 follow-up) A Loom kernel.barrier fences only the memory space it names. #28 changed the barrier before pack_completed_q8 from <workgroup> to <global>, which orders the reduce's global output stores but drops the LDS fence the cooperative and multipass reducers relied on for their staging buffers. Keep both: a <global> barrier followed by the original <workgroup> barrier (HIP __syncthreads semantics) at the five pack sites. The standalone two-dispatch reducer (ggml_flash_attention_decode_split_reduce_f32 and its qwen3_moe copy) had the same bug as #28: wave 0 rewrites partial_max in global memory with the exp scales and every wave reads it back after a <workgroup>-only barrier, so the reads could hit stale GL1/L0 lines holding the block maxima. It is not dispatched today. Add the <global> barrier there too. Compiled ISA, gfx1151: every pack handoff (direct, cooperative, multipass) is now vmcnt(0), vscnt(0), s_barrier, buffer_gl1_inv, buffer_gl0_inv, s_barrier; the standalone reducer gains buffer_gl1_inv/buffer_gl0_inv (it had none). Measured with a separate process forcing KFD queue evictions: Qwen3-0.6B 0 of 118 divergent; Qwen3-Coder-30B, 2113-token prompt, partial alignment 256: 8 of 8 correct, 0 NaN. Decode speed vs #28 within run-to-run noise (0.6B d0 276-287 vs 296-300, d1000 204-215 vs 199-207; 30B d2100 57-61 vs 59-62 tok/s). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
merged commit Sep 27, 2026
00adc2b
into
1bit/hrx-vulkan-patched
1 check passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #28 (engine#123 / ggml-org#140).
Two changes
1. Pack sites: fence LDS again, not only global. A Loom
kernel.barrierfences only the memory space it names (loom docs, Synchronization names both rendezvous and memory). #28 replaced thebarrier<workgroup>beforepack_completed_q8withbarrier<global>. That orders the reduce's globaloutputstores, but it drops the LDS fence the cooperative and multipass reducers had for their staging buffers. All 5 pack sites now use both, i.e. HIP__syncthreadssemantics:2. Standalone two-dispatch reducer: same bug as #28, now fixed. In
ggml_flash_attention_decode_split_reduce_f32and its qwen3_moe copy, wave 0 rewritespartial_maxin global memory with the exp scales. Every wave then reads it back after abarrier<workgroup>only. Wave 0 had read those same lines during the max pass, so other waves could hit stale GL1/L0 copies (the block maxima instead of the scales). This kernel is not dispatched today; it gets the<global>barrier too.Compiled ISA (gfx1151)
vmcnt(0),vscnt(0),s_barrier,buffer_gl1_inv,buffer_gl0_inv,s_barrier, then the packglobal_load_b128.loom-compile: before, 2 barriers and no invalidates; after, 3 barriers withbuffer_gl1_inv+buffer_gl0_inv.Tests
GGML_HRX_FA_PARTIAL_ALIGN=256, evictionsiree-test-loom(scale 1 vs stale max 0; not committed), 2000 runs before the fixDecode speed vs #28 is within run-to-run noise on this shared box (2 interleaved rounds):
🤖 Generated with Claude Code