Repository navigation
hrx: pin llama.cpp to the hybrid (AMD core + our ggml-hrx) - fixes MUL_MAT failures and #315 NaN - #339
hrx: pin llama.cpp to the hybrid (AMD core + our ggml-hrx) - fixes MUL_MAT failures and #315 NaN#339bong-water-water-bong wants to merge 1 commit into
Conversation
AMD core sync with our HRX backend kept whole — the configuration measured to fix
both the MUL_MAT failures and the all-NaN logits:
test-backend-ops -b HRX0 AMD core + our ggml-hrx 1073/1073 (0 FAIL)
AMD core + AMD ggml-hrx 1019/1071 (54 FAIL)
GLM-4.7-Flash, 269-token, -ub 512
our backend 0 NaN, reply " Paris."
AMD backend 18 NaN, reply "??????"
third_party/hrx-system is unchanged at 98d05d94 (our loom: AMD's does not export
loomc_amdgpu_runtime_global_flags_t, which our loom-jit.cpp needs), so this moves
one gitlink: llama.cpp 522dab47 -> f2099e9b.
|
Superseded by #336, which makes the same change — llama.cpp to The one thing worth preserving from here, in case it is useful later: the hybrid branch was assembled and verified locally ( |
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
Pins
third_party/llama.cppto AMD's core sync with our HRX backend kept whole (f2099e9b), which is the configuration measured to fix both open blockers.third_party/hrx-systemis unchanged at98d05d94(our loom — AMD's does not exportloomc_amdgpu_runtime_global_flags_t, which ourloom-jit.cppneeds), so this moves one gitlink.Measured — fresh build directories, the engine's own
ExternalProjectarguments (cmake/hrx.cmake: fresh-B, Release, amdclang,-DHRX_SOURCE_DIR=…), never through an incrementally rebuilt nested directory (round 77's lesson):test-backend-ops -b HRX0-ub 512ggml-hrx' Paris.'ggml-hrx(this PR)' Paris.'ggml-hrx??????The 54 failures are 40
MUL_MAT(iniq1_s/iq1_m/iq3_xxs/mxfp4/pq2_0/ptq1_0/q2_*) plus others; the NaN is engine#315.Gates:
tools/check_pins.py origin/main-> ok third_party/llama.cpp: 522dab478 -> f2099e9b7 is ahead;tools/registry_build.py --check-pinspasses and the registry records the new source.Why our backend and not AMD's: replacing our
ggml-hrxwith AMD's refactor is what carries both defects — see engine#315 and the evidence table. Thehrx-vulkan-patchedline already carries our patchedggml-hrx; this is that choice made explicit on the core sync.Note this is offered as the alternative to #329 (which pins the pure sync
56c3c8a3). Both are verified; the difference is whether AMD'sggml-hrxor ours is in the pinned tree, and that is the call this PR is asking for rather than assuming.