Repository navigation
cuda: preserve weighted RMS reduction order with vector loads - #157
Open
GenerelSchwerz wants to merge 28 commits into
Open
GenerelSchwerz wants to merge 28 commits into
GenerelSchwerz wants to merge 28 commits into
Conversation
Carry the reviewed kernel sources from PR125, PR129 and PR132 as the matching control for the separate tensor-affine HC follow-up. Assisted-by: Codex
Preserve gate intermediates, original rounding and live HC outputs. Retain strict input alias checks and standard allocation dependencies. Assisted-by: Codex
Use the exact tested control sources from published PR130, PR129 and PR132. Assisted-by: Codex
Reuse the allocated affine match to find the planned HC_POST and retain prepared emission priority. Reject expanded normalization cuts whose terminal is not written by this kernel. Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Use four original logical RMS threads per physical thread to load/store float4 values with a 256-thread block while preserving the original 1024-thread accumulation and reduction tree. Weighted RMS keeps normalized-before-gamma multiplication and its existing parameter broadcast rules. No temporary gamma-first stores or changed sum association are needed.
The new path requires plain F32 weighted RMS, no ADD/scale, physical warp size 32, width >= 1024 divisible by 4, and aligned source/destination rows and strides. Small/odd/unaligned cases, unweighted RMS, live-output and explicit in-place fallbacks, HC pending/custom reads and quantized producer emission retain their existing paths. No model names, fixed hidden widths, new graph matcher or allocator policy.
This is a one-file leaf, 41 additions and 1 deletion, atop PR156 prerequisite commit
08b1d9bb1d1fa59996052591f3197782871030b7. The PR base remainsreference/upstream-kernels-20261001, officialdcd387a412ca54e172a8d60eb71ef6753850c8ca. The full diff includes earlier prerequisite commits; review the one-file leaf commit for the RMS change.Measurements
RTX 5070 Ti, SM120a, matching PR156 control, normal upstream scheduler/allocation/dispatch and GPU events around 2000 captured graph replays per measurement.
Other standalone norms, F16/F32 multirow components and live-intermediate controls are flat. These are isolated component measurements, not model-serving, MTP, pristine-upstream or additive series gains. Older-GPU speed is unmeasured.
Validation
Evidence and exact commands/source/module/raw-array/profile hashes:
/home/gencoolpc/moe-cache-tests/results/upstream-kernel-series-20261001/HYPERCONNECTION/HC-ORDER-ASSESSMENT-20261008, especiallyFINAL-NORM-REGRESSION-RESULTS-V1.json,NORM-AUDIT-REPORT-V1.json,ORDER-COMPARISON-RESULTS-V1.mdandNORM-ROOT-SOURCE-REVIEW-V1.md.Related work and scope
Upstream PR20520 and PR29720 already propose vectorized RMS with a different sum tree. This fork alternative preserves the original tree and existing callback/shader interface. No speed comparison against those branches is claimed. Future upstream integration should reconcile these approaches with the existing authors.
Pending HC reads, quantized emission, shared activation staging, prefetch and paired down/injection are separate work and are not covered by this leaf.
Requirements