Repository navigation
ci(bump-hrx): resolve ggml/src/ggml-hrx conflicts toward our line, keep failing elsewhere - #335
Conversation
…ep failing elsewhere The scheduled sync fails every run at the merge of AMD's pin into 1bit/hrx-vulkan-patched: 15 conflicting files, all under ggml/src/ggml-hrx/ (the workflow header says so, and I confirmed the ancestry guard passes: 1bit/hrx-vulkan is an ancestor of AMD's b802a507b with no extra commits, so the failure is the merge, not the guard). Measured 2026-10-06 in fresh build directories with this workflow's own ExternalProject arguments, three requests of a 269-token prompt at -ub 512: AMD core + our ggml-hrx full gate 1073/1073 (0 FAIL) 0 NaN, reply " Paris." AMD core + AMD ggml-hrx full gate 1019/1071 (54 FAIL) 18 NaN, reply "??????" So resolving these files toward AMD's side imports engine#315's all-NaN logits into the pinned line. The merge now takes ours for conflicts inside ggml/src/ggml-hrx/ and keeps AMD's side everywhere else, and still aborts loudly on any conflict outside that path so a real divergence cannot be auto-resolved by accident.
Careful with
|
| provider set | kDecodeSplitMaxKeyValueTokenCapacity |
|---|---|
ours (monolithic FA kernel: direct_f32 64-256, cooperative_f32 257-2048, multipass_f32 2049-262144) |
32768 is correct |
AMD's refactor (ops/flash_attention/* + motifs/flash_attention/*: only direct_f32 64-256 and cooperative_f32 257-2048 — multipass removed) |
must be 2048 |
Measured on the AMD-side merge: with only those two providers, a capacity of 2049-32768 offers the decode-split dispatch with no provider, so the kernel selector rejects every candidate (all_rejected) and the whole decode fails — which is exactly what the comment above that constant warns about. I capped it to 2048 there for that reason.
So: if this PR restores our provider set, keep 32768 and do not carry the 2048 cap over. The constant and the corpus must agree, and nothing in CI catches the mismatch — it only shows up as a decode failure on a long context.
Verified tree-wide that with the AMD corpus the only two providers are the ones listed (reduce_fused_direct_f32, reduce_fused_cooperative_f32, both in motifs/flash_attention/completion_counter_reduce.loom) and reduce_fused_multipass_f32 exists nowhere.
PR Reviewer Guide 🔍(Review updated until commit 15f4beb)Here are some key observations to aid the review process:
|
|
Persistent review updated to latest commit 15f4beb |
Fixes the failing scheduled HRX pin sync (engine#332).
What failed
The scheduled
bump-hrxrun has failed on every run from 2026-10-01 to 2026-10-05 (the pin silently froze; #328 fixed the merge strategy, this fixes the content resolution). The step "Sync the fork and merge AMD's pin into our line" exits 1 at the merge of AMD'sb802a507binto1bit/hrx-vulkan-patched: 15 conflicting files, all underggml/src/ggml-hrx/. The ancestry guard passes (1bit/hrx-vulkanis an ancestor ofb802a507bwith 0 extra commits), so the failure is the merge, not the guard.Why the conflicts recur
Both lines rewrote
ggml/src/ggml-hrx/— AMD's kernel-family refactor against our backend. It will conflict on every scheduled run until one side'sggml-hrxis chosen once and committed as the merge resolution.Which side — measured, not assumed
Fresh build directories, the engine's own
ExternalProjectarguments (cmake/hrx.cmake: fresh-B, Release,amdclang,-DHRX_SOURCE_DIR=…), fulltest-backend-ops -b HRX0, and a GLM-4.7-Flash Q4_K_M 269-token prompt at-ub 512:test-backend-ops -b HRX0ggml-hrx' Paris.'ggml-hrx??????The 40
MUL_MATfailures (iq1_s/iq1_m/iq3_xxs/mxfp4/pq2_0/ptq1_0/q2_*) and the engine#315 all-NaN logits are properties of AMD'sggml-hrx, not of the core sync. Resolving these files toward AMD's side would import #315's NaN into the pinned line.What this does
ggml/src/ggml-hrx/resolve toward ours (the branch being merged into) — the patched line.ggml/src/ggml-hrx/still aborts loudly (::error::), so a real divergence cannot be auto-resolved by accident.::notice::naming the files it took.Verification
Conflict resolution logic A/B-simulated on a scratch repo: only-
ggml-hrxconflict → merge commit with our side, AMD side elsewhere; mixed conflict → abort. Patch applies cleanly on currentmain(ef04beb,bump-hrx.ymlblob90e36c0).Credit: the measurement and resolution were developed in the notes on engine#332; this PR lands them. Blocked until now only by the
workflowscope on the pushing credential.