Skip to content

Bump the HRX pair to the full AMD sync (llama.cpp b04c4e95 + hrx-system 563afce1) - #329

Closed
bong-water-water-bong wants to merge 8 commits into
mainfrom
1bit/hrx-pinsync-full
Closed

bong-water-water-bong wants to merge 8 commits into
mainfrom
1bit/hrx-pinsync-full

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Moves both pinned submodules onto AMD's tested pair, merged with our commits. Opened as a draft on purpose — see the validation gate.

from to
third_party/llama.cpp 522dab47 b04c4e95 = AMD b802a507 merged with our 159 commits, plus llama.cpp#94 so it also carries the engine#315 fix
third_party/hrx-system 98d05d94 563afce1 = AMD 40a1a36c merged with our libhrx knobs

hrx-system conflicts were 3 Loom AMDGPU target-info files, resolved to AMD: our commit there was f1b558f19 "[Loom] GFX11 wave64: drain ALU dependencies … (backport)" — a backport of a fix AMD has since landed. Our 98d05d94a libhrx knobs commit does not conflict and is preserved.

llama.cpp carried 75 conflict hunks: 11 files took AMD's side per-conflict with the auto-merge remainder preserved, loom-libs/manifest.json became a semantic union (277 files / 215 exports, keeping 61 + 55 of ours, plus AMD's new sources/kernels/link_modules/plan_cases), and 3 superseded monoliths gave way to AMD's ops/flash_attention/* + motifs/flash_attention/* refactor (symbols still exist at the new paths).

Validation gate — do not un-draft until this passes

Neither merged tip has been built or run. GitHub-hosted CI builds this without HRX, so CI cannot catch an HRX breakage. On Strix Halo:

cmake -B build -G Ninja -DONEBIT_HRX=ON && cmake --build build --target onebit
tests/serve_e2e.sh build/1bit <Qwen3-0.6B Q4_K_M .gguf> hrx
tests/serve_e2e.sh build/1bit <Qwen3-0.6B Q4_K_M .gguf> cpu
test-backend-ops -b HRX0

Then this is the pin the release gate must be re-run against — it measures the build we would ship, which the currently-running pass does not.

Moves both pinned submodules onto AMD's tested pair, merged with our commits.

third_party/llama.cpp  522dab47 -> b04c4e95
  = AMD b802a507 merged with our 159 commits, then llama.cpp#94 on top, so it
    carries the engine#315 NaN fix as well as AMD's kernel-corpus refactor.
third_party/hrx-system 98d05d94 -> 563afce1
  = AMD 40a1a36c merged with our libhrx device knobs (HRX_AQL_BLOCK_SIZE,
    HRX_COMMAND_BUFFER_MODE). The three Loom target-info conflicts resolved to
    AMD's side: our commit there was a backport of a fix AMD has since landed.

DO NOT MERGE UNTIL VALIDATED ON STRIX HALO. GitHub-hosted CI builds this WITHOUT
HRX, so CI cannot catch an HRX breakage:
  cmake -B build -G Ninja -DONEBIT_HRX=ON && cmake --build build --target onebit
  tests/serve_e2e.sh build/1bit <Qwen3-0.6B Q4_K_M .gguf> hrx
  tests/serve_e2e.sh build/1bit <Qwen3-0.6B Q4_K_M .gguf> cpu
  test-backend-ops -b HRX0
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Pre-validation check: registry_pins will fail as this stands

This moves third_party/llama.cpp 522dab47 -> b04c4e95 but does not touch registry/architectures.json, which records the pin it was generated from. ctest registry_pins (tools/registry_build.py --check-pins) compares that record with the committed gitlink and exits 1 on a mismatch. It is added under if(Python3_Interpreter_FOUND), not under ONEBIT_HRX, so the GitHub-hosted build job runs it.

Verified locally on main: current tree passes; with the gitlink at b04c4e95 and the record still at 522dab47 it fails. So scripts/registry-regen.sh is needed before un-drafting.

The content should not change: I diffed the registry's inputs between the two pins - convert_hf_to_gguf.py, conversion/, src/llama-arch.cpp, gguf-py/gguf/constants.py - and they are byte-identical, so the architecture set is the same and only sources["llama.cpp (hrx)"] moves. (sources records no hrx-system pin, so 563afce1 does not enter the registry.)

Two things checked while in there, in case they save a run:

  • The decode-split dispatcher at the new head references exactly one symbol, ggml_flash_attention_decode_split_f32_f16_wmma_next_q8, and the merged loom-libs/manifest.json (215 exports) still exports it. The removed ops/flash_attention_decode_split_f32_f16_wmma.loom monolith therefore has no dangling caller.
  • That is also why engine#314 changes shape: the 1,529-line kernel carrying the two kernel.subgroup.vote ops - and its ggml_flash_attention_decode_split_f32_f16_wmma export - is gone at this pin, and kernel.subgroup.vote appears 0 times across all 12 motifs|ops/flash_attention/ sources. Masking is per-score in motifs/flash_attention/decode_partials.loom now.

The bump moved third_party/llama.cpp to b04c4e95 but left
registry/architectures.json recording 522dab4, so ctest registry_pins - which the
build job runs, not gated on ONEBIT_HRX - fails on the pin mismatch.

The registry's inputs are byte-identical between the two pins (convert_hf_to_gguf.py,
conversion/, gguf-py/gguf/constants.py, src/llama-arch.cpp, src/llama-model.cpp), so
only sources["llama.cpp (hrx)"] moves; the architecture set and counts are unchanged.

Co-authored-by: agent <agent@local>
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Added the registry regen in e943577 (on this branch, so the PR is now 3 files): sources["llama.cpp (hrx)"] moves 522dab47 -> b04c4e95. Verified on the branch:

  • tools/registry_build.py --check-pins -> passes
  • --check-gaps, tests/registry_gaps_test.py, tests/census_watch_test.py, tools/copyright.py --check -> pass

I made it a one-line edit rather than running scripts/registry-regen.sh because the generator's inputs are byte-identical between the pins - same diff set as the previous comment (convert_hf_to_gguf.py, conversion/, gguf-py/gguf/constants.py, src/llama-arch.cpp, src/llama-model.cpp) - so the regenerated file differs from the current one only in that field. If you would rather fold it into the bump commit, scripts/registry-regen.sh amortises the same change with git commit --amend.

The HRX validation gate in the description is untouched: still a draft, still needs the Strix Halo build plus serve_e2e and test-backend-ops.

AMD's refactor removed the multipass reduce provider, so the merged corpus only
covers key_value_token_capacity up to 2048. Our inherited constant was still
32768, which would offer the decode-split dispatch for capacities with no
provider, making the selector reject every candidate ("all_rejected") and
failing the whole decode. Capped at 2048.
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Static check while the HRX gate waits: nothing is left dangling by the corpus refactor

The riskiest part of this merge is the corpus refactor (three monoliths replaced by motifs), because dispatchers reference kernels by string through GGML_HRX_KERNEL_REF("<catalog>", "<name>") - a dropped name would pass the build and fail at runtime. I extracted every ref from ggml/src/ggml-hrx/dispatch_registration/** at b04c4e95 and compared it with the merged manifests:

catalog referenced exported missing
loom_libs 185 215 none
qwen3_moe 14 14 none
qwen 1 4 none
hrx 2 2 none

202 references, all present. 51e538fa changes only dispatch-flash-attention.cpp, so it does not move any of these.

That cap is the other half of the same refactor - the capacity range rather than the names: the merged corpus keeps direct_f32 (64-256) and cooperative_f32 (257-2048) for decode-split but not the multipass provider above 2048, so offering the dispatch past 2048 would reject every candidate. With kDecodeSplitMaxKeyValueTokenCapacity = 2048 the longer contexts fall through to flash_attention_f32_f16_wmma, which is what the note above the constant asks for. Worth one line in the description if you want the HRX gate to know why >2048 now takes the general kernel.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

The bump does not build: the manifest union kept 47 pre-refactor paths

I ran the validation gate from the description on Strix Halo (cmake -B build -G Ninja -DONEBIT_HRX=ON -DONEBIT_HRX_TOOLCHAIN=/opt/rocm-therock ...), and it fails in the HRX dependency build, before any GPU test:

FAILED: ggml/src/ggml-hrx/kernel-corpus-sources.inc ...
generate_kernel_corpus.py: failed to read
  .../kernel-corpus/kernels/loom-libs/ops/unary_f32.loom: [Errno 2]

ops/unary_f32.loom does not exist at this pin - AMD's refactor moved it to ops/elementwise/unary_f32.loom. The merged loom-libs/manifest.json still lists it under files, and generate_kernel_corpus.py reads every files entry, so the first stale path stops the build.

Measured against the tree the PR pins (51e538fa), from ggml/src/ggml-hrx/kernel-corpus/kernels/loom-libs:

list entries missing on disk
manifest.json -> files 277 47
export compile_recipe + link_modules refs 241 24

First missing files entries (25 of 47):

ops/unary_f32.loom, ops/scale_f32.loom, ops/scale_bias_f32.loom, ops/copy_f32.loom,
ops/binary_f32.loom, ops/binary_bc_f32.loom, ops/rmsnorm_f32.loom,
ops/rmsnorm_binary_f32.loom, ops/rmsnorm_binary_q8_1_x4.loom,
ops/llm_attention_q_matmul_rope_f32_f32_wmma.loom,
ops/llm_attention_k_matmul_rope_set_rows_f32_f32_wmma.loom,
ops/llm_attention_v_matmul_set_rows_f32_f32_wmma.loom,
ops/llm_attention_q_matmul_rope_decode_f32_f32.loom,
ops/llm_attention_k_matmul_rope_set_rows_decode_f32_f32.loom,
ops/llm_attention_v_matmul_set_rows_decode_f32_f32.loom,
ops/mul_mat_id_f32_f32_wmma.loom, ops/mul_mat_id_f16_f16_wmma.loom,
ops/mul_mat_id_swiglu_f32_f32_wmma.loom, ops/mul_mat_id_swiglu_f16_f16_wmma.loom,
motifs/mul_mat_id_swiglu_f16_f16_wmma_body.loom, motifs/mul_mat_id_swiglu_f16_f16_accumulate.loom,
ops/mul_mat_id_postops_f32_f32_wmma.loom, ops/mul_mat_id_postops_next_rmsnorm_f32_f32_wmma.loom,
ops/mul_mat_bias_f32_f32_wmma.loom, ops/mul_mat_add_f32_f32_wmma.loom

and the missing recipe refs include motifs/unary_f32_apply.loom, motifs/rmsnorm_f32.loom, motifs/rope_f32.loom, motifs/binary_f32_apply.loom, motifs/mul_mat_id_f32_f32_postops.loom, motifs/mul_mat_id_f32_f32_wmma_core.loom plus the ops/llm_attention_* and ops/mul_mat_* families above.

So the semantic union kept our pre-refactor entries alongside AMD's new ones. What is needed: drop or repath our stale files/recipe entries to AMD's new locations, then confirm every export's symbols still resolve there. Happy to finish that mapping and hand over a corrected manifest - I did not push to the fork's branch while you may be building it.

One other thing worth knowing, not a pin problem: a first configure aborted in tools/check_hrx_revision.py with fatal: Not a valid object name 32d0c76a..., because git submodule update --init --depth 1 leaves third_party/hrx-system shallow and the checker cannot walk ancestry in a shallow clone. 32d0c76a is an ancestor of your 563afce1 (compare: ahead 339, behind 0), so git fetch --unshallow in that submodule clears it.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Follow-up: the manifest union is not consumable by its own generator (it is not just stale paths)

I repaired the 47 stale files paths locally by basename (14 repath cleanly, e.g. ops/unary_f32.loom -> ops/elementwise/unary_f32.loom) and dropped the rest. The generator then got past the missing file and failed on the next invariant:

generate_kernel_corpus.py: conflicting dependency list for ops/rows/set_rows.loom

collect_sources() requires exactly one library_sources list per primary source (it raises otherwise). On the manifest as merged:

check result
exports 215
distinct primary sources 117 (after repair)
primary sources with conflicting library lists 2 in the original (ops/mul_mat_f32_f32_wmma.loom, ops/mul_mat_q6_k_packed_f16_wmma.loom); the repath then created a third (ops/rows/set_rows.loom)
exports whose primary source file does not exist at the pin 19

The 19 lost primaries are ours, not AMD's:

ggml_rmsnorm_binary_q8_1_x4, ggml_rmsnorm_gate_f32_f16, ggml_rmsnorm_gate_f32_q8_1_x4,
llm_attention_{q,k,v}_matmul_rope_*_{wmma,decode_f32_f32},
ggml_mul_mat_id_{f32_f32,swiglu_f32_f32,postops_f32_f32,postops_next_rmsnorm_f32_f32}_wmma,
ggml_mul_mat_{add,bias,bias_add,bias_add_next_rmsnorm,add_next_rmsnorm}_f32_f32_wmma,
ggml_set_rows_scatter

AMD's refactor restructured exactly those areas (ops/llm/llm_attention_*_{legacy,tiled,vector}_*, ops/rows/, ops/matmul/), so the union kept manifest entries whose files the merge no longer carries. The files list therefore describes 277 files while only 230 exist, 14 exist elsewhere and 31 exist nowhere.

Conclusion: this needs reconciliation, not a repath. Either take AMD's corpus wholesale and deliberately re-port the kernels we still need (the -DONEBIT_GPU_PRIVATE HIP add-on now covers several of those paths anyway), or port each of the 19 into AMD's layout and regenerate files/recipes consistently. My throwaway repair proves the first failure is the manifest and not the toolchain - with it the build proceeds to the next manifest invariant - but dropping 22 exports is not an answer either.

Reproduce (Strix Halo, in a throwaway worktree, no fork push):

cmake -B build -G Ninja -DONEBIT_HRX=ON -DONEBIT_HRX_TOOLCHAIN=/opt/rocm-therock \
  -DONEBIT_HRX_LIBHSA=/opt/rocm-therock/.../libhsa-runtime64.so.1
cmake --build build/hrx/llama --target llama-bench -j4

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

The corpus repair is pushed and the HRX build is green

Branch 1bit/amd-sync-corpus-repair = 82a14638e, ahead 1 / behind 0 of 1bit/hrx-vulkan-patched, so it lands cleanly on the pinned line. 67 files: AMD's corpus restored wholesale, our kernels kept beside it, the manifest rebuilt, two includes restored.

What changed

  • Corpus base. loom-libs/ is AMD's b802a507 tree. The auto-merge had folded our kernel edits into AMD's refactored files, which left duplicate SSA names in shared motifs (motifs/dequant.loom had 6 %iq1_s_format= definitions against AMD's 3, two of them in one scope) - the corpus could not be parsed at all.
  • Our kernels kept, beside AMD's. 54 sources under loom-libs/ours/, taken from the pre-merge fork, so nothing of ours is lost and no path collides with AMD's refactor. (24 of the 54 had genuinely been dropped by the merge; the other 29 still existed but are used from ours/ so our exports have their own primaries.)
  • manifest.json rebuilt as AMD's 160 exports plus our 55, recipes repathed to ours/. That removes the 47 missing files paths, the 24 missing recipe refs, and the two shared primaries that carried two different library_sources lists (ops/mul_mat_f32_f32_wmma.loom, ops/mul_mat_q6_k_packed_f16_wmma.loom).
  • One C++ merge artifact: dispatch-mul-mat.cpp had our body but AMD's include block, so common_append_mul_mat_token_tail and the IQ3_XXS matcher were undeclared. The two includes are back.

Build evidence (Strix Halo, throwaway worktree, thermal-run -j4 so the running pass was not disturbed)

cmake -B build -G Ninja -DONEBIT_HRX=ON -DONEBIT_HRX_TOOLCHAIN=/opt/rocm-therock \
  -DONEBIT_HRX_LIBHSA=/opt/rocm-therock/lib/python3.14/site-packages/_rocm_sdk_core/lib/libhsa-runtime64.so.1
cmake --build build --target onebit -j4     # OK

0 FAILED steps; bin/llama-bench and build/1bit both link. The re-run reports [3/3] Completed 'llama_hrx', i.e. nothing stale.

Remaining for the PR's gate is the GPU half (serve_e2e.sh hrx+cpu, test-backend-ops -b HRX0); the box is still running the pre-fix pass.

One config note repeated here because it cost a cycle: git submodule update --init --depth 1 leaves hrx-system shallow, and check_hrx_revision.py then cannot walk ancestry (fatal: Not a valid object name 32d0c76a). 32d0c76a is an ancestor of 563afce1, so git fetch --unshallow in that submodule clears it.

Once this branch is on the pinned line, the engine gitlink in this PR needs to move 51e538fa -> 82a14638e (and registry/architectures.json's sources with it - I verified the registry's generator inputs are unchanged, so only that field moves).

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

GPU half run: the build is green, and test-backend-ops -b HRX0 now runs - 989 pass, 84 fail

The gate that blocked the bump has stopped on the box, so I could finally run the hardware half in the repaired worktree (Strix Halo, gfx1151, under thermal-run -j4).

Building it took three more merge artifacts, all now in the branch (amended, so the branch is 420b4e6f, still ahead 1 / behind 0 of 1bit/hrx-vulkan-patched):

  1. dispatch/transient-reuse-guard.cpp was missing from ggml/src/ggml-hrx/CMakeLists.txt. The file survived the merge, the source list did not, so libggml-hrx.so carried ggml::hrx::pack_transients_without_false_dependencies as an undefined symbol and test-backend-ops died at the first op that used it.
  2. common_q8_prefill_relaxed() lost its definition - dispatch-mul-mat.cpp kept the declaration in the header and both callers, but the definition was inside a block that resolved to AMD's side. Another undefined symbol, hit at MUL_MAT_HADAMARD.
  3. The two includes from my earlier report.

onebit and test-backend-ops both build with 0 errors.

The runs

Backend HRX0: FAIL        989 OK / 84 FAIL
op failures shape
MUL_MAT 40 type_a in iq1_s, iq1_m, iq3_xxs, iq2_xxs, mxfp4, q2_*, pq2_0, ptq1_0; ERR = inf
MUL_MAT_ID 30 f32/f16, m=512,n=129,k=256; ERR ~ 1.0 (i.e. wrong, not noise)
GET_ROWS 7 same IQ/MXFP4/Prism types; ERR = inf
ADD 5

ERR = inf across exactly the formats our corpus adds and AMD's does not is the pattern to look at: at this pin, AMD's motifs/dequant.loom carries 45 IQ references where ours carries 81, and our ours/ops/kquant_decode_f32.loom is the file that enumerates IQ1_S/IQ1_M/IQ2_XXS/IQ2_XS/IQ3_XXS. So the sync's dispatch is very likely selecting a kernel that has no lane for those formats.

I do not want to claim that from the pattern alone, so I have started the control: the same test-backend-ops -b HRX0 build at the pre-sync pin (522dab478), to separate "the sync regressed these" from "HRX already failed these". I will report which it is before anyone acts on this.

Everything else - the corpus reconciliation, the kept kernels under ours/, the manifest rebuild - is unchanged from the previous comment.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Control run: the 84 failures are a regression from the sync, not pre-existing

I built the same test-backend-ops at the pre-sync pin (522dab478, in ~/wt/pin-fork-hrx) and ran the identical command:

pre-sync 522dab478 synced + repaired 420b4e6f
OK 1075 989
FAIL 0 84
not supported 19687 19689
verdict 2/2 backends passed Backend HRX0: FAIL

Same box, same toolchain, same -b HRX0, same IREE_HAL_AMDGPU_LIBHSA_PATH, both under thermal-run -j4. The pre-sync tree passes everything; the synced tree fails 84.

So this is not something my corpus reconciliation introduced - the reconciliation is what makes the synced tree build and run at all - and it is not a pre-existing HRX weakness. It is the sync: at this pin, the ops handling the formats our corpus adds and AMD's does not (iq1_s, iq1_m, iq2_xxs, iq3_xxs, mxfp4, q2_*, and the Prism pq2_0/ptq1_0) return inf, and MUL_MAT_ID at m=512,n=129,k=256 returns a result off by ~1.0.

The shape of it matters for the merge policy: the "auto-merged remainder preserved" rule kept our callers and matchers while the merge resolved the files that carry the format lanes to AMD's side. Building a corpus from AMD's tree plus our kernels keeps our files intact, but it does not bring back a dispatch path that AMD's refactor rewrote.

This PR should not be un-drafted or merged until these 84 are back to 0 - MUL_MAT_ID alone is every MoE model's expert matmul, and the IQ/Prism formats are what our models use. I am running serve_e2e.sh on Qwen3-0.6B (HRX + CPU) now and will report it.

Branch is 420b4e6f (1bit/amd-sync-corpus-repair, ahead 1 / behind 0).

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

serve_e2e.sh on the repaired pin: CPU passes, HRX cannot load its kernels

Same repaired worktree, Qwen3-0.6B-Q4_K_M:

  • cpu: PASS - /health 200, /v1/models, answers "Paris", streams 32 chunks. So the engine binary and the non-HRX path are fine at this pin.
  • hrx: FAIL - at the first decode:
E load_kernel_executable: compiled ABI does not match manifest for gfx1151|6442987762ea...
E load_kernel_executable: compiled ABI does not match manifest for gfx1151|503d476a18...
E load_kernel_executable: compiled ABI does not match manifest for gfx1151|11699ba2f3...
E graph_compute: failed to prepare command 587 kind=Kernel kernel_id=8042800487544822715 bindings=5
E llama_decode: failed to decode, ret = -3

(My first attempt passed the device as HRX0, which 1bit serve rejects with --device HRX0 cannot run a .gguf (hrx, cpu or ds4); with hrx it gets past device selection and fails at kernel load.)

runtime/kernel-executable-cache.cpp:226 is the check: it compares the compiled export's binding_count, constant_byte_length and parameter_count against dispatch.bindings and definition.launch_parameters, i.e. against the manifest's recorded ABI. So at this pin the manifest's ABI metadata and what the pinned loom actually compiles disagree, and the kernel is refused before it runs.

That is a different symptom from the 84 test-backend-ops failures - there the kernels loaded and returned inf or a result off by 1.0; here one is refused outright. Both say the same thing though: the corpus metadata and the pinned hrx-system loom are not a matched pair. The thing I would check first is whether b802a507's paired hrx-system is actually our 563afce1; hrx-minimum-commit.txt only expresses a floor (32d0c76a, which 563afce1 satisfies 339 commits later), so a satisfying pin is not necessarily AMD's tested one. If AMD's pair is a specific commit, that would explain both symptoms at once.

I have not changed anything else. Branch 420b4e6f, ahead 1 / behind 0.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Split into two causes: one is our hrx-system loom delta, the other is the hybrid dispatch

I ran the two controls.

First: the pins are not the problem

ROCm/ggml-staging-automation is where bump-hrx.yml reads AMD's pair from, and right now it pins exactly:

llama.cpp   b802a507b57c5a5019c38f3b4498c1b7a14dec08
hrx-system  40a1a36c0082e3049932de44776f64d39ad15610

which is precisely the base our 51e538fa and 563afce1 are built on. So "wrong hrx-system pin" is ruled out - I was wrong to suspect it.

Cause 1 - our loom delta makes every kernel fail to load

Our 563afce1 differs from AMD's 40a1a36c in three files, one of them substantial:

loom/src/loom/target/arch/amdgpu/planning/wait_plan.c   +222  -5
libhrx/src/libhrx/runtime.c                              +25
iree/.../hal/drivers/amdgpu/logical_device.c             +29

That wait_plan.c delta is our "drain ALU dependencies after overwriting lane-mask SGPRs" work (loom_amdgpu_wait_plan_needs_mask_write_state, loom_amdgpu_wait_plan_track_mask_reads). I put AMD's stock 40a1a36c in the worktree and rebuilt (the loom binaries are rebuilt, mtime confirms), and:

  • serve_e2e.sh ... hrx: PASS - 0 compiled ABI does not match manifest, 32 SSE chunks.
  • The three kernels it had refused were all AMD's, not ours (ggml_mul_mat_vector_f32_f32, ggml_rmsnorm_binary_f32, ggml_mul_mat_swiglu_q4_q8_1_x4_lowtoken_dot_q8_output), with the same binding/launch counts in our manifest and AMD's.

So our loom delta changes the compiled export ABI (or the parameter count) relative to what the corpus records, and every kernel the model needs is then refused. It needs re-expressing on top of AMD's plan, or dropping where AMD's newer code already covers the hazard.

Cause 2 - the 84 op failures are not the loom

With AMD's stock loom the run is identical: 989 OK / 84 FAIL, same ops. So that half comes from the hybrid dispatch - our matchers and callers auto-merged into AMD's tree while the kernel definitions now come from AMD's corpus. Next control: revert ggml/src/ggml-hrx/dispatch_registration and runtime to b802a507 and re-run; if the 84 become "not supported" instead of failing, the dispatch is the whole of cause 2.

The worktree currently sits at AMD's stock loom (the passing state for serve_e2e); branch 420b4e6f is unchanged.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Cause 1 narrowed: it is the merged wait_plan.c, and the message may not be about the ABI at all

Two things I should correct from my last comment.

The two other hrx-system deltas are inert. libhrx/src/libhrx/runtime.c (+25) and runtime/src/iree/hal/drivers/amdgpu/logical_device.c (+29) only read HRX_AQL_BLOCK_SIZE / HRX_COMMAND_BUFFER_MODE from the environment and pass them as device parameters; unset (they are unset in every run of mine), device_param_count is 0 and the code path is AMD's. So the loom file is the whole of cause 1.

And wait_plan.c differs at all three pins:

98d05d94 (pre-sync 1bit/main)  362735d7e220
563afce1 (post-merge 1bit/main) 8c02c014caae
40a1a36c (AMD's pin)            a6f663f47231

So the post-merge file is neither ours nor AMD's - it is the auto-merge of the two, which is the same failure mode as motifs/dequant.loom (our backport spliced into AMD's restructured planner).

The message may be misleading. Look at the condition in kernel-executable-cache.cpp:226:

if (executable->launch_program == nullptr ||
    executable->export_info.binding_count != dispatch.bindings.size() ||
    !constant_abi_valid ||
    executable->export_info.parameter_count != dispatch.bindings.size() + definition.launch_parameters.size()) {
    error_message = "compiled ABI does not match manifest for " + key;

launch_program == nullptr - i.e. the planner produced nothing - raises the very same text. Given that the file that fixes it is planning/wait_plan.c, and that the counts in our manifest and AMD's agree for all three refused kernels, "our merged planner fails to plan these kernels" is the better reading than "the ABI disagrees". I would not name it an ABI problem again until someone checks launch_program directly.

So the fix on the hrx-system side is to resolve wait_plan.c to AMD's (our backport is a duplicate of machinery AMD now has in planning/mask_write_wait.c / wait_frontier.c), or re-apply it deliberately on top of AMD's plan. I am building exactly that now - 563afce1 with only wait_plan.c taken from 40a1a36c - and will report the serve_e2e result, then push it as a branch if it holds.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Cause 1 fixed and verified: 1bit/amd-sync-wait-plan = 3d0cca89a0

One file, and it is a clean fast-forward on 1bit/main (ahead 1 / behind 0):

loom/src/loom/target/arch/amdgpu/planning/wait_plan.c   5 insertions(+), 222 deletions(-)

563afce1's wait_plan.c is replaced by 40a1a36c's. Our libhrx device knobs and the two inert environment-driven files are untouched, so the only behavioural change is the planner.

Verified on Strix Halo (rebuild + run, thermal-run -j4):

serve_e2e.sh build/1bit Qwen3-0.6B-Q4_K_M.gguf hrx
ok   hrx: /health 200 ... ok hrx: answers Paris ... ok hrx: streams (32 chunks)   PASS
0 × "compiled ABI does not match manifest"

Before the change the same command failed at the first decode with three refusals and llama_decode ret = -3; now it answers. build/1bit and test-backend-ops build with 0 errors.

That confirms the reading: the auto-merged planner produces no launch program for those kernels, and kernel-executable-cache.cpp reports launch_program == nullptr under the "ABI does not match manifest" text. AMD's loom already has the lane-mask hazard machinery (planning/mask_write_wait.c, wait_frontier.c), so our backport was a duplicate spliced into a restructured file - the same failure mode as motifs/dequant.loom on the corpus side.

hrx-system has no AGENTS.md, and I left 1bit/main itself alone - the branch is there to fast-forward when you want it (or to point this PR at it).

Still open

Cause 2 is untouched: test-backend-ops -b HRX0 is still 989 OK / 84 FAIL, identical with either loom, so the gate still cannot pass. That half is the hybrid dispatch - our matchers and callers against AMD's kernel definitions - and it is the part that needs the actual port.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Cause 2 characterised: the failures are shape-specific, so it is specialisation selection, not missing formats

I bisected with test-backend-ops' own filters (each run is seconds, so these are single-op repros).

MUL_MAT_ID, m=512,k=256, by token count:

n result
128 0 fail / 2
129 20 fail / 20
256 0 fail / 2
512 0 fail / 2

At n=129 every variant fails - f32, f16, q1_0, q4_0, q4_1, q4_K, q8_0, iq2_xxs, mxfp4 - with ERR ~ 1.0, which is what one missing row of the output looks like. A 129-token batch is a 128-row tile plus a one-token tail, so the tail is not being computed.

MUL_MAT(type_a=iq1_s, m=16, k=256), by token count: n=1 OK, n=2..8 FAIL (ERR = inf), n=9 OK, and so on. So it is not "the IQ format has no lane" - the same format passes at n=1 and n=9 and fails in between. It is the choice of kernel for those token counts.

One flag ruled out: GGML_HRX_DISABLE_DEPRECATED_MUL_MAT_DISPATCH=1 makes it worse (all 20 MUL_MAT_ID cases at n=129 fail instead of 20 already failing but the rest of the op set changing too), so our deprecated path is not the culprit.

Taken with the pre-sync tree passing all of these (1075 OK / 0 FAIL), the mechanism is clear: our dispatch selects a specialisation for the token count, and the set of specialisations it can choose from is now AMD's corpus instead of ours, so the selected kernel is not the one the shape needs. That is the port - our dispatchers re-based onto AMD's kernel definitions - and it is bounded: it is MUL_MAT and MUL_MAT_ID token-count selection, not the whole 130-file delta.

Nothing else changed; both fix branches stand (llama.cpp 420b4e6f, hrx-system 3d0cca89a0).

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Correction to the flag sentence in the previous comment: on the MUL_MAT_ID -p "m=512,n=129,k=256" filter, GGML_HRX_DISABLE_DEPRECATED_MUL_MAT_DISPATCH=1 gives 20 fail / 20 tested - exactly the same as without it. So the flag changes nothing at this shape; it simply does not rescue it. My "makes it worse" was a misreading of two runs whose filters differed.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Cause 2 is mutual inconsistency: neither dispatcher works against the other side's kernels

I traced the MUL_MAT_ID tail failure into the code.

The merge took AMD's dispatcher wholesale for this file. dispatch-mul-mat-id.cpp:

522dab4 / ebe16621 (ours)   1bee27ec8a95   -> ggml_mul_mat_id_f32_f32_wmma, ..._postops_..._wmma, ..._postops_next_rmsnorm_..._wmma
b802a507 / 420b4e6f (AMD)   5c9822c5aa9e   -> ggml_mul_mat_id_tiled_input_f32_publish_f32, ..._skinny_input_...

but the header next to it, dispatch-mul-mat-id-common.h, is ours (merged 467fda4e vs AMD 843956d9). So the file pair is itself half-merged, and with AMD's dispatcher the tail shape fails 20/20 at n=129.

Restoring our dispatcher does not fix it either. With dispatch-mul-mat-id.cpp taken from ebe16621 and rebuilt, our kernels are selected - and the failure moves one stage earlier:

load_kernel_executable: compiled ABI does not match manifest for
  gfx1151|9a32b179...|ggml_moe_build_expert_table|...|token_count=129|route_count=4|route_stride=4|expert_count=4|...
load_kernel_executable: compiled ABI does not match manifest for
  gfx1151|9a32b179...|ggml_moe_build_expert_partition_table|...

Those are AMD's definitions (ops/moe_routing_tables.loom, libs: 0), being called with our dispatch's binding layout. Our ours/ exports that do carry library lists are fine (ggml_mul_mat_id_f32_f32_wmma has 11 of our own motifs); the routing tables are the ones where our caller meets AMD's definition.

So it is not "our kernels are stale" or "AMD's kernels are buggy" - it is that our caller layout and AMD's kernel definitions disagree in both directions, and each side only works with its own. That is the port: per op, pick one side's dispatcher and its kernel definitions, rather than the half-merge that is there now. I have reverted my experiment and left the tree as it was.

Running now: is any of this AMD's own?

The one control that settles the boundary is a pure AMD tree - ggml/src/ggml-hrx at b802a507, hrx-system at 40a1a36c, none of our 130 files - running the same test-backend-ops -b HRX0. If that also reports 84 failures, the regressions are AMD's on gfx1151 and the port is about our features, not about fixing AMD. If it reports 0, every failure is the half-merge. It is building now and I will report it.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

The pure-AMD control was invalid - please disregard that line of enquiry

Reverting ggml/src/ggml-hrx to b802a507 in place builds, but the binary it produces cannot dispatch:

47,065 x  graph_compute: unsupported HRX node 0: ...
~6.6k failures in the ops it does run (ADD, RMS_NORM, CPY, SET_ROWS, MUL, CLAMP, ...)

That is a revert inside our tree, not AMD's shipped pair, so it says nothing about whether AMD's own b802a507 + 40a1a36c passes on gfx1151. I stopped the run, reverted, and rebuilt the tree back to the repaired state (420b4e6f, ours/ intact). A real control needs its own worktree at AMD's exact tree, and I would only spend that build if someone wants the answer.

Everything in the previous comment stands and does not depend on this: with AMD's dispatcher the MUL_MAT_ID tail fails 20/20 at n=129, and restoring our dispatcher moves the refusal to ggml_moe_build_expert_table / ..._partition_table (AMD's definitions, our binding layout). The two sides are mutually inconsistent, and the port is per-op.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

The merge discarded our side of 59 files - that is the whole of cause 2

I classified every file under ggml/src/ggml-hrx by blob hash across our pre-merge tip (ebe16621), AMD's pin (b802a507) and the merged tree (420b4e6f):

class files
AMD-only (new in AMD) 259
ours-only (AMD lacks) 153
identical in both 100
ours discarded - both had it, they differed, merged took AMD's 59
auto-merged (both sides' edits) 21
kept ours 14

So of the 94 files both trees had and that differed, the merge took AMD's side in 59 (63%), auto-merged 21 and kept ours in 14. The PR description's "11 content-conflicted files ... auto-merged remainder preserved (so our non-conflicting 159 commits survive)" does not hold inside ggml/src/ggml-hrx: our side was discarded, not preserved.

The 59 by area: dispatch_registration 14, kernel-corpus 21, tools 11, dispatch 7, benchmarks 2, runtime 2, loom-jit.cpp/h 2.

And they map exactly onto the failures:

discarded file op that fails
dispatch_registration/common/dispatch-get-rows.cpp GET_ROWS (7)
dispatch_registration/common/dispatch-mul-mat-id.cpp MUL_MAT_ID tail (30)
dispatch-binary.cpp, dispatch-rmsnorm.cpp, dispatch-scale.cpp, dispatch-unary.cpp, dispatch-copy.cpp, dispatch-gather-add.cpp ADD (5) and the rest of the op set
kernel-corpus/.../ops/mul_mat_f32_f32_decode.loom, mul_mat_f32_f32_wmma.loom, mul_mat_q5_k_q8_plane_wmma.loom, mul_mat_swiglu_f32_f32_wmma.loom MUL_MAT (40)
kernel-corpus/.../motifs/dequant.loom, motifs/mul_mat_f32_f32_wmma_core.loom the quant dequant paths
the whole kernel-corpus/kernels/qwen_moe/** corpus and its manifest.json MoE / MUL_MAT_ID
dispatch/command-plan.h, command-program.cpp, dispatch/dispatch.h, dispatch-registry.h, loom-jit.cpp, runtime/kernel-executable-cache.h the ABI/planner side of cause 1

That is why the callers and callees disagree: we kept our dispatcher in some files while the kernel it targets and the dispatch registry that describes it were replaced by AMD's.

Experiment running: git checkout ebe16621 -- <the 59> ("prefer our side for exactly the files the merge threw away"), rebuild, and run the full test-backend-ops -b HRX0. If it holds, the port is a 59-file restore plus whatever the two sides genuinely need to reconcile; if it does not, the build will say which shared file cannot simply be taken back.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Worse than the 59: the merge also deleted 83 files we had and AMD does not

The "restore the 59 from ebe16621" build fails in the corpus generator:

generate_kernel_corpus.py: failed to read .../kernel-corpus/kernels/qwen_moe/ggml/linear_q6k_f32.loom: [Errno 2]

because one of the 59 is our kernel-corpus/kernels/qwen_moe/manifest.json, which lists our qwen_moe kernels - and those kernels are not in the merged tree at all. So I counted:

files
files only we had (not in AMD's pin) 153
of those, entirely missing from the merged tree 83
- under kernel-corpus/kernels/loom-libs 61
- under kernel-corpus/kernels/qwen_moe 15
- under benchmarks/loom 6

So the merge's effect on ggml/src/ggml-hrx is 59 files where our side was discarded + 83 files deleted outright = 142 of our files. My earlier corpus reconciliation recovered the 54 loom-libs sources that our 55 surviving exports reference (they are the ours/ tree), which is why the tree builds at all - but the qwen_moe corpus (15 kernels plus its manifest) and 29 other files are simply gone.

That also explains why our restored dispatchers then refused AMD's routing kernels: the whole MoE corpus they were written against was deleted, so the names resolve to AMD's replacements with AMD's layouts.

Where that leaves it. "Keep our side" means restoring 142 files, not 59 - and the qwen_moe corpus (with its manifest) is the largest coherent piece of that. That is also the MoE/HIP workstream, so the two threads meet there. I have not attempted the 142-file restore: at that point it is a decision about which corpus wins per model family, and it should be made deliberately rather than by an agent's git checkout.

Tree reverted to the repaired state (420b4e6f, clean) and rebuilding.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Bigger than the 84 ops: real models do not run at the synced pin

I ran llama-bench -ngl 99 -fa 1 -p 512 -n 0 -r 1 -dev HRX0 on the two models the HRX decode work targets, against the repaired tree and against the pre-sync build, same box, same lock, same env.

Pre-sync (522dab478): both fine

gpt-oss 20B MXFP4 MoE            | HRX | HRX0 | pp512 | 958.75 t/s
deepseek2 30B.A3B Q4_K - Medium  | HRX | HRX0 | pp512 | 536.35 t/s

Synced + repaired (420b4e6f + hrx-system 3d0cca89a0): both die before the first measurement

gpt-oss:  CONFIG/INVALID: symbol '@ggml.mul_mat_id.input_size' value 2880 violates constraint 'mul...'
          test_prompt: failed to decode prompt batch, res = -3  -> failed to run prompt warmup
GLM:      CONFIG/INVALID: symbol '@ggml.flash_attention.qk_head_size' value 576 violates constraint 'range'
          test_prompt: failed to decode prompt batch, res = -3  -> failed to run prompt warmup

Both failures are Loom root linking failures on the corpus's own symbol constraints - mul_mat_id.input_size has a multiple-of-something constraint that 2880 violates, and flash_attention.qk_head_size has a range constraint that GLM's 576 falls outside. So the corpus refuses the shapes these models actually use.

Two controls:

  • It is not the private add-on. The same two failures occur with -DONEBIT_GPU_PRIVATE and without it (I built both; identical errors).
  • It is not the model files or the box. The pre-sync build reads the same two GGUFs on the same GPU and gets almost 1000 t/s on gpt-oss.

And it is not just the two models: serve_e2e on Qwen3-0.6B passes at the synced pin, which is why the earlier report looked as good as "989 OK / 84 FAIL". In practice no model that uses MoE expert matmuls at a non-multiple input size, or MLA attention with a 576 head size, runs at all - gpt-oss-20b and GLM-4.7-Flash are exactly the models the HRX kernel work exists for.

So the count of things this bump needs is larger than the 84 op regressions: the port has to satisfy the corpus's constraints, not merely resolve kernel names. A repaired-and-building tree that cannot start the target models is still not landable, and I would keep this PR a draft.

(Round-10 note for context: the gate pass itself measured pin 522dab47, pre-sync and pre-#315, so nothing it produced speaks to this.)

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

The two model failures are corpus constraints, and relaxing them is necessary but not sufficient

I traced round-19's failures to exact lines. AMD's corpus tightened two symbols against ours:

@ggml.flash_attention.qk_head_size
  ours, 522dab4  ops/flash_attention_f32_f16_wmma.loom:26          range(16, 36864), mul 16
  AMD,  420b4e6f motifs/flash_attention/decode_partials.loom:5     range(16, 512),   mul 16

@ggml.mul_mat_id.input_size
  ours, 522dab4  ops/mul_mat_id_f32_f32_wmma.loom:14               range(256, 32768), mul 32
  AMD,  420b4e6f ops/matmul_id/{skinny,tiled}_input_f32_publish_f32.loom:14
                                                                   range(256, 32768), mul 256

GLM-4.7-Flash needs qk_head_size 576 (over AMD's 512) and gpt-oss-20b uses mul_mat_id.input_size
2880 (32x90, not a multiple of 256), so both die in Loom root linking. The pre-sync build reads the same
two GGUFs on the same GPU at 958.75 t/s and 536.35 t/s.

I relaxed both in 7 files and rebuilt:

  • gpt-oss prompt now runs: 1906.16 t/s pp512 (pre-sync 958.75).
  • GLM still fails, one stage later: TARGET/059: target 'amdgpu-rdna3-5' export 'ggml_flash_attention_f32_f16_wmma' config 'amdgpu.rdna3_5.core' rejected 'index.assume'. So the range
    was guarding a real kernel limitation - 576 needs our attention kernel (ours accepted 36864) or a widened
    one, not a relaxed check.
  • Generation still fails even where the prompt works. /completion on gpt-oss returns 500:
    materialize: compiled ABI does not match manifest ... |ggml_binary_f32|recipe=direct|.... So
    wait_plan.c fixed the kernels it covered, not all of them; ggml_binary_f32 is refused at decode.

Net: the constraints are a real defect class of this sync (they exclude shapes real models use), but they
are not the whole of it, and I am not claiming the models work - prompt-only is not "running". Both
changes are uncommitted in the validation worktree so they cannot be mistaken for a landed fix.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Correction: relaxing the constraints is a false fix, not a partial one

My previous comment reported "gpt-oss prompt now runs: 1906.16 t/s" after relaxing mul_mat_id.input_size from mul 256 to mul 32. That number is real but the change is not a fix, and I want to correct the framing before anyone builds on it.

The prompt path passing means the constraint check no longer rejects 2880 - it does not mean a kernel supports 2880. The decode that follows is refused:

materialize: compiled ABI does not match manifest for ... |ggml_binary_f32|recipe=direct|...
materialize: ... |ggml_get_rows_f32|workload=dynamic|ggml.get_rows_f32.hidden_capacity=2880|...
materialize: ... |ggml_rmsnorm_f32|workload=dynamic|ggml.rmsnorm_f32.hidden_size=2880|...
graph_compute: failed to prepare command 0 ... -> llama_decode ret = -3

So AMD's kernels are compiled for the shapes their constraints declare, and letting 2880 past the guard just moves the failure to kernel materialisation. The range/mul bounds are load-bearing, not over-strict.

What that means for the port: mul_mat_id.input_size at 2880 needs our mul_mat_id kernels (ours allowed mul 32) and our dispatcher for that op - which is exactly the file the merge discarded (dispatch-mul-mat-id.cpp, of the 59). Restoring the dispatcher alone is not enough, because the MoE routing kernels it then calls (ggml_moe_build_expert_table, ggml_moe_build_expert_partition_table) are AMD's definitions and refuse our binding layout - which is why round 15's experiment failed the way it did.

I have reverted the constraint edits; the validation worktree is back to 420b4e6f exactly. Nothing in this changes the earlier structural findings - the 142-file half-merge, the model-level failures, and the two verified fix branches (420b4e6f, 3d0cca89a0) all stand.

Engine#329 as it stands does not build: the AMD sync auto-merged our kernels
into AMD's refactored corpus, which left duplicate SSA names in shared motifs,
47 manifest paths with no file, 19 exports with no source, and
dispatch-mul-mat.cpp with our body and AMD's include block.

Point third_party/llama.cpp at 1bit/amd-sync-corpus-repair (420b4e6f) and
third_party/hrx-system at 1bit/amd-sync-wait-plan (3d0cca89a0), and move the
registry source pin with them. Both are one commit ahead of what this PR pinned,
so the pin check still only moves forward.

With these, cmake --build build --target onebit succeeds (0 errors) and
serve_e2e.sh passes on HRX for Qwen3-0.6B. test-backend-ops -b HRX0 still
reports 989 OK / 84 FAIL and gpt-oss/GLM do not run, so this makes the gate
measurable, not green.
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

I pinned the repaired trees into this branch - it builds now, and here is what the gate will show

This PR's head (4847858) does not build: the AMD sync auto-merged our kernels into AMD's refactored corpus, leaving duplicate SSA names in shared motifs, 47 manifest paths with no file, 19 exports with no source, and dispatch-mul-mat.cpp with our body and AMD's include block. Running validate-pinsync.sh against it would have failed at the build stage.

So I added one commit on top, 4e775e3 (fast-forward, nothing rewritten):

third_party/llama.cpp   51e538fa -> 420b4e6fe  (1bit/amd-sync-corpus-repair)
third_party/hrx-system  563afce1 -> 3d0cca89a  (1bit/amd-sync-wait-plan)
registry/architectures.json  sources["llama.cpp (hrx)"] moved with them

tools/check_pins.py 4847858 says both are ahead (forward-only, no rollback), and the registry edit is the one line.

What the gate will now show, measured on this box (all with -DONEBIT_HRX=ON, thermal-run -j4):

stage result
cmake --build build --target onebit PASS, 0 errors
serve_e2e.sh ... hrx (Qwen3-0.6B) PASS (0 ABI refusals, 32 SSE chunks)
serve_e2e.sh ... cpu PASS
test-backend-ops -b HRX0 FAIL - 989 OK / 84 FAIL (pre-sync: 1075/0)
gpt-oss-20b-MXFP4 / GLM-4.7-Flash llama-bench -p 512 FAIL - Loom root linking: mul_mat_id.input_size 2880 needs a multiple of 256, flash_attention.qk_head_size 576 exceeds the corpus's 512 cap

So the validation is now measurable rather than dead at the build step, and it will fail on the last two rows - those are the port (our dispatchers and kernels against AMD's kernel definitions and constraints), not something this commit can fix.

Two notes for whoever runs it: git submodule update --init --depth 1 leaves hrx-system shallow, which trips check_hrx_revision.py with fatal: Not a valid object name 32d0c76a; git fetch --unshallow in that submodule clears it. And the gate pass is finished, so the only thing holding the box is the 17 standing hyperloom processes.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Correction: "compiled ABI does not match manifest" is not about the ABI - it is launch_program == nullptr

I have been reporting the decode failures as an "ABI/specialisation mismatch". That was wrong, and I instrumented the check to find out.

kernel-executable-cache.cpp:239 is the only site of that message, and its condition is four disjuncts:

if (executable->launch_program == nullptr ||
    executable->export_info.binding_count != dispatch.bindings.size() ||
    !constant_abi_valid ||
    executable->export_info.parameter_count != dispatch.bindings.size() + definition.launch_parameters.size())

I printed all of those numbers on every load, re-applied the repoint so GLM reaches decode, and ran the server. For every failing kernel the counts agree:

ggml_mul_mat_vector_bias_residual_f32_f32:
    launch_program=0  compiled{bindings=6 params=7 const_bytes=4}  manifest{bindings=6 launch=1}
ggml_mul_mat_vector_q6_f32_f32:
    launch_program=0  compiled{bindings=3 params=4 const_bytes=4}  manifest{bindings=3 launch=1}
ggml_rmsnorm_binary_f32:
    launch_program=0  compiled{bindings=3 params=4 const_bytes=4}  manifest{bindings=3 launch=1}

6 == 6 and 7 == 6 + 1, 4 == 1 * 4; same for the others. launch_program = 0 is the entire failure - the planner produced no program and the code reports it with the ABI text. (This is what I suspected in round 13 and then set aside once swapping wait_plan.c made serve_e2e pass for Qwen - I should not have dropped it.)

So the port's second layer is Loom planning failing for those kernels at those shapes, not a metadata or workload-specialisation disagreement. It is the same class as the round-14 wait_plan.c merge artifact, but it survives with AMD's stock planner in the tree, so it lives in the kernels/planning for those shapes rather than in our loom delta. That changes where to look: not at bindings/launch_parameters reconciliation, but at why the planner emits nothing.

Instrumentation reverted; the tree is back to 420b4e6f exactly.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

The decode blocker is 15 kernels with an empty launch program - and 9 of them are AMD's

Following the launch_program=0 finding: the compile is not failing. compiled_ref->resolve() succeeds (its failure path would report compiled_ref->error_message() instead), the executable loads and export_info is populated - and all three ABI counts agree. What is empty is the host-side program the planner builds (ggml_hrx_loom_jit_launch_program_evaluate).

Counted the distinct kernels that fail this way in one GLM run: 15, split 6 ours / 9 AMD's:

AMD's:  ggml_binary_f32, ggml_binary_bc_f32, ggml_mul_mat_vector_f32_f32,
        ggml_mul_mat_vector_bias_residual_f32_f32, ggml_mul_mat_vector_q6_f32_f32,
        ggml_moe_build_expert_table, ggml_moe_build_expert_partition_table,
        ggml_mul_mat_id_skinny_pair_input_f32_swiglu_publish_f32, ...

So this is not the ours/ reconciliation and not a metadata disagreement: AMD's own kernels compile but the planner emits no program for them at GLM's shapes, with AMD's stock planner in the tree.

That is precisely the class our lane-mask drain backport (f1b558f19, wait_plan.c, +222 lines) exists to fix. My round-14 fix dropped it, because the merge's version of that file was itself broken (8c02c014 - neither ours nor AMD's) and made serve_e2e fail for Qwen. Both things were true, and I only reported one of them: the merged file was broken and AMD's planner does not schedule these kernels. The right move is what FULL-SYNC-WITH-AMD.md says was resolved "to AMD's side" and what I flagged as needing re-expression - port the drain fix onto AMD's planner rather than deleting it.

Current map, all measured on this box:

layer state
corpus build / manifest fixed (420b4e6f)
planner, Qwen path fixed (3d0cca89a0)
GLM attention symbol cap fixed by repointing (prompt 872.47 t/s)
15 kernels, empty launch program open - blocks decode
84 test-backend-ops failures open

Nothing changed in the tree; it is still 420b4e6f exactly.

…g launch programs

The previous pin (420b4e6f) still had runtime/loom-jit-disk-cache.cpp storing
entries without the host-side launch program, so every cache hit was refused as
"compiled ABI does not match manifest". 39588ef9 skips caching those results and
bumps the cache version.

With it, GLM-4.7-Flash-Q4_K_M runs on HRX end to end: pp512 896.84 t/s,
tg32 28.05 t/s. serve_e2e on Qwen3-0.6B still passes. test-backend-ops -b HRX0
is unchanged at 989 OK / 84 FAIL.
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

GLM-4.7-Flash now runs on HRX at these pins - pp512 896.84 t/s, tg32 28.05 t/s

I found the cause of the "compiled ABI does not match manifest" wall, fixed it, and moved this PR's pin to the fix (99a7a75; check_pins.py 4e775e3 and registry_build.py --check-pins both pass).

The bug

runtime/loom-jit-disk-cache.cpp serialises magic, version, launch_config, workload_argument_count, the hsaco and the manifest - and never the host-side launch program. launch_program does not appear in that file at all, so loom_jit_disk_cache_load() returns a result whose launch_program is null - and kernel-executable-cache.cpp:239 rejects it through the launch_program == nullptr disjunct while printing the ABI message. A cache hit on any kernel that needs a launch program was therefore refused, with an error naming the wrong subsystem.

The fix (39588ef9)

  • do not store a result that carries a launch program - the format cannot hold it;
  • bump kVersion 1 -> 3 so entries written by the broken code are ignored;
  • two manifest corrections without which GLM cannot run at all: point ggml_flash_attention_f32_f16_wmma at our pre-merge kernel (AMD's caps qk_head_size at 512 and GLM needs 576) and give the kquant exports that share ours/ops/kquant_decode_f32.loom our own motif files, so they stop resolving against AMD's root corpus.

Measured on Strix Halo

GLM-4.7-Flash-Q4_K_M  HRX  HRX0  pp512  896.84 t/s
GLM-4.7-Flash-Q4_K_M  HRX  HRX0  tg32    28.05 t/s

before: 57 compiled ABI does not match manifest refusals across 15 distinct kernels and no run at all; after: AMD's kernels plan, serve_e2e.sh ... hrx on Qwen3-0.6B still passes (32 chunks, "Paris"), and the same host runs GLM end to end.

What this does not fix

test-backend-ops -b HRX0 is unchanged at 989 OK / 84 FAIL, so those are a separate cause from the cache bug. And I have not checked GLM's output quality - it runs, which is not the same as being right, and #310 is precisely a decode-accuracy issue.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

The bump regresses 32 MUL_MAT_ID shapes that passed at the current pin - measured, with a one-minute repro

I have the pre-sync build and the synced build side by side on the same box, so this is a direct A/B of the same test:

pre-sync  (~/wt/pin-fork-hrx, pin 522dab4):   test-backend-ops -b HRX0 -o MUL_MAT_ID  ->  108/108 passed
this PR   (39588ef9 = 420b4e6f corpus + our planner/cache fixes):                        ->   78/108 passed

The 32 failures are all numerical, all ERR ≈ 1.0 (the output is entirely wrong, not partly corrupted):

[MUL_MAT_ID] ERR = 1.000004080 > 0.000500000   MUL_MAT_ID(type_a=f32,type_b=f32,n_mats=4,n_used=4,b=0,m=512,n=129,k=256): FAIL
[MUL_MAT_ID] ERR = 0.999995787 > 0.000500000   MUL_MAT_ID(type_a=f16,...): FAIL

Failing token counts across every type_a and both b: n = 5, 17, 32, 129. Passing: n = 1, 4, 16, 128, 256, 512, 1024, 8192. (Shapes with k=16/64 come back not supported, correctly: ggml.mul_mat_id.input_size is declared range(256, 32768), mul(256).)

It is the kernel, not the dispatch

I instrumented match_mul_mat_id_dispatch to print what each shape asks for. Failing and passing shapes select the same kernel with the same parameters apart from the token count:

HRX-MMID token=5   route=4 input_route=4 expert=4 route_stride=4 skinny=1
HRX-MMID token=17  route=4 input_route=4 expert=4 route_stride=4 skinny=0
HRX-MMID token=129 route=4 input_route=4 expert=4 route_stride=4 skinny=0     <- FAIL

skinny=1 is token_count <= 5, so 5 goes to ggml_mul_mat_id_skinny_input_f32_publish_f32 and the rest to ggml_mul_mat_id_tiled_input_f32_publish_f32 - and both routes fail for their respective token counts. Nothing in the dispatch differs between n=128 (passes) and n=129 (fails) except that number, so this is AMD's kernel handling of certain token counts at these pins, not a selection or binding problem.

Why this matters for the PR

These shapes worked at 522dab4 and do not work at the AMD sync. That is the per-op-family question the PR has been waiting on, and it now has a number attached: one op family, 32 regressed shapes, a one-minute repro (test-backend-ops -b HRX0 -o MUL_MAT_ID, no server, no model load), and the same ERR = 1.0 / "output never written" signature as the GLM all-NaN logits in #315 - which makes it plausible that one fix closes both.

My own commits on this branch (420b4e6f, 39588ef9) make the bump build, plan and run; they do not fix this, and they cannot - it is in the kernel set, not in the corpus manifest or the JIT cache.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

The MUL_MAT_ID regression is a two-line dispatch change - 78/108 -> 108/108, and GLM's NaN goes with it

Following the regression I reported above, I found the fix, and it is smaller than the port I was bracing for.

Our pre-merge dispatcher (ebe16621) had no tiled/skinny split - one
ggml_mul_mat_id_f32_f32_wmma for every token count. AMD's matcher instead names their tiled_*/skinny_*
kernels, which mishandle n = 5, 17, 32, 129. Both constants in
dispatch_registration/common/dispatch-mul-mat-id.cpp now name our kernel again:

 static constexpr KernelCatalogRef kMulMatIdTiledF32F32Kernel =
-    GGML_HRX_KERNEL_REF("loom_libs", "ggml_mul_mat_id_tiled_input_f32_publish_f32");
+    GGML_HRX_KERNEL_REF("loom_libs", "ggml_mul_mat_id_f32_f32_wmma");
 static constexpr KernelCatalogRef kMulMatIdSkinnyF32F32Kernel =
-    GGML_HRX_KERNEL_REF("loom_libs", "ggml_mul_mat_id_skinny_input_f32_publish_f32");
+    GGML_HRX_KERNEL_REF("loom_libs", "ggml_mul_mat_id_f32_f32_wmma");

Our mul_mat_id_* exports are already in the corpus (the reconciliation kept them beside AMD's), so this is a
dispatcher change only - no manifest edit, no kernel port:

check before after
test-backend-ops -b HRX0 -o MUL_MAT_ID 78 / 108 (32 x ERR ~ 1.0) 108 / 108
GLM-4.7-Flash -p 512 -n 32, default -ub 512 test_gen: failed to decode generation batch pp512 902.21 t/s, tg32 27.40 t/s

That second row is #315: the all-NaN logits used to fire for every ubatch of >= 256 tokens, which is why the
-ub 128 workaround existed. At the default ubatch GLM now generates, so the workaround is no longer
needed either.

One intermediate result worth keeping: pointing only the tiled constant at our kernel gave 90/108 and fixed
n=17/32/129, but left n=5 at ERR ~ 85 - garbage rather than an unwritten output - because the skinny
constant was still on a different interface. Both must name the same kernel, exactly as our pre-merge
dispatcher did.

I will commit this onto the repair branch and move the pin, so the branch carries the fix rather than just
pointing at it.

… kernel again

39588ef9 dropped our MoE matmul in favour of AMD's tiled/skinny pair, which
mishandles n = 5, 17, 32 and 129. 56c3c8a3 points both matcher constants back at
our ggml_mul_mat_id_f32_f32_wmma.

  test-backend-ops -b HRX0 -o MUL_MAT_ID   78/108 -> 108/108
  GLM-4.7-Flash, default -ub 512           decode failure -> pp512 902.21 t/s, tg32 27.40 t/s

The GLM row is engine#315: the all-NaN logits are gone and the -ub 128 workaround
is no longer needed.
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Pinned: 56c3c8a3 is on 1bit/amd-sync-corpus-repair and this branch now pins it (c718497). check_pins.py says ahead, registry_build.py --check-pins passes.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Gate after the fix: 989/1073 -> 1019/1071, failures 84 -> 54

Full test-backend-ops -b HRX0 with 56c3c8a3 pinned:

passed failed
before (39588ef9) 989 84
after (56c3c8a3) 1019 54

Exactly the 30 MUL_MAT_ID failures. What is left, by op:

MUL_MAT  40      e.g. m=16,n=1..6,9,16,k=256 (type_a=q2_K and type_b=f32)
GET_ROWS  7      type=pq2_0/ptq1_0/mxfp4/q2_K/iq2_*, n=256, m=5   -> ERR = inf
ADD       5      MUL_MAT_VEC_FUSION(type=mxfp4/iq2_xxs, m=1, n=32) -> ERR 44 .. 1388

ERR = inf is 25 of the 54 - and every one of those is at a small token count (n=1..6, 9, or m=5). That is the same low-token-count class as the mul_mat_id skinny defect I just fixed: a kernel that leaves its output unwritten when the batch is small.

Which suggests the same lever applies again, and it is worth checking before anyone plans a kernel port: several of these names (ggml_mul_mat_f32_f32_wmma, ggml_get_rows_f32, ggml_binary_f32) are names both trees use. If the merge kept AMD's file for a shared name, then the corpus has AMD's narrow implementation under a name our dispatcher trusts - and repointing that export at our pre-merge source (as I did for ggml_flash_attention_f32_f16_wmma) may be the whole fix. I will check that rather than assume it.

ggml_binary_f32 is a name both trees use and the merged corpus held AMD's file
under it, so the ADD path ran AMD's implementation. 89d4f17c repoints that export
at our pre-merge source.

  test-backend-ops -b HRX0 -o ADD   5 failures -> 22/22 passed
  GLM-4.7-Flash pp512/tg32          885.17 / 27.30 t/s, unchanged

Not fixed this way, and recorded as such: ggml_get_rows_f32 (the same repoint
makes GET_ROWS worse, 7 -> 24) and MUL_MAT (40 -> 42 alone).
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Correction: the binary_f32 change is neutral, and my "5 -> 0" was a case-set mismatch

I claimed in the previous comment (and in commit 89d4f17c) that repointing ggml_binary_f32 fixed the ADD failures: "5 failures -> 22/22 passed". That comparison was invalid, and the full gate says so.

-o ADD does not select the [ADD]-class cases. Measured:

full run, [ADD]-class cases:            5   (all 5 fail, all are MUL_MAT_VEC_FUSION)
-o ADD run,  [ADD]-class cases:         0   (it ran 22 cases of a different class, all passing)
full run, before the change:  1019 passed / 54 failed
full run, after the change:   1019 passed / 54 failed

So the binary_f32 repoint is neutral: the same five failures are still there, and I was comparing the 5 failing vec-fusion cases against 22 unrelated ones.

What is true: the corpus did hold AMD's file under our ggml_binary_f32 name (f8e74b98 against our 1064d5cc), repointing it does restore our pre-merge kernel, -o ADD passes 22/22, and GLM is unaffected (pp512 885.17, tg32 27.30). What is not true is that it fixes anything measurable.

I am moving the pin back to 56c3c8a3 so this branch carries only changes with a measured effect - the JIT-cache fix and the mul_mat_id fix. The extra commit stays on 1bit/amd-sync-corpus-repair rather than being rewritten away, so the record is visible.

The real lesson is about my own method: -o <op> is not the same selection as the [class] in the full log, and only the full run is a valid A/B here. The two negative results from the previous round stand and were measured the right way - GET_ROWS gets worse with the same repoint (7 -> 24), and MUL_MAT does not improve (40 -> 42 alone, 37 with the constants swapped).

…neutral

The full gate is 1019 passed / 54 failed both before and after 89d4f17c, and my
'5 -> 22/22' comparison was invalid: -o ADD selects 22 cases of a different class
while the 5 [ADD]-class failures are the MUL_MAT_VEC_FUSION ones, which still
fail. Keeping the pin on the two changes that have a measured effect (the
JIT-cache fix and the mul_mat_id routing fix).
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Correction: the mul_mat_id fix did not remove the NaN - it raised the limit

I wrote that after 56c3c8a3 "GLM now generates at the default ubatch, so the -ub 128 workaround is no longer needed". That is too strong, and here is the measurement on the current pin.

llama-server, GLM-4.7-Flash-Q4_K_M, HRX0, default -ub 512, GGML_HRX_NAN_CONTINUE=1, two generated tokens each:

prompt tokens NaN events
512 (llama-bench -p 512, single ubatch) 0
558 2 (every generated token)
666 2
774 2
882 2
839 (needle prompt) 18 over three requests

So the fix raised the batch limit - pre-fix a 269-token prompt already failed, post-fix 512 works - but every prompt from 558 up still produces all-NaN logits. #315 is not fixed by this pin, and the workaround still has a purpose.

Note the boundary coincides with the -ub 512 split (558 = 512 + 46): either the continuation batch with a non-empty KV cache is the trigger, or a >512 single batch is. I have not separated those two, and I am not going to claim which.

What the fix genuinely did, and what I am still claiming: test-backend-ops -b HRX0 -o MUL_MAT_ID 78/108 -> 108/108, the full gate 84 -> 54 failures, and a 512-token prompt that failed pre-fix now generates (pp512 902, tg32 27.4). That is real; "the NaN is gone" was not.

An attempted -ub 128 re-run at 839 tokens never reached the server (my readiness gate waited 1200s and the server had already exited), so whether the workaround still works post-fix is unmeasured rather than assumed either way.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Second correction, and a much better lead: the NaN is content-dependent, not length-dependent

Two measurements this round overturn what I said in the last two comments.

1. The same prompt, before and after the mul_mat_id fix.

269-token real-text prompt, -ub 512, 12 requests, one NaN event counted per request
  pre-fix  (39588ef9) : 11 / 12
  post-fix (56c3c8a3) : 12 / 12

So the fix did not address the NaN at all. My "GLM now generates at the default ubatch" (round 44) and "the fix raised the limit from 269 to 512" (round 48) were both wrong: the 512-token case I measured with llama-bench -p 512 uses a synthetic prompt, and comparing it to a real-text prompt is not a comparison.

2. The trigger is the prompt's content, not its length or the ubatch size.

llama-bench with a synthetic 839-token prompt, GGML_HRX_NAN_CONTINUE=1, all four ubatch settings:

-ub 128, 256, 512, 1024 at -p 839   ->  0 NaN events, every configuration completes

Same model, same pins, same length: no NaN at any -ub. The real-text prompts at 269, 558, 666, 774, 839 and 882 tokens fail, and they fail at every -ub I have tried except 128. A synthetic prompt samples the vocabulary roughly uniformly; repeated English prose drives the MoE routing to a small set of experts.

So the lead is routing concentration, not length. That fits everything: content-dependence, the apparent length-dependence (longer prose piles more assignments onto the same hot expert), why random prompts never fire it, and why the mul_mat_id swap changed the op test but not this.

What still stands from the fix, measured: -o MUL_MAT_ID 78/108 -> 108/108, full gate 84 -> 54 failures. Those are op-level and unaffected by any of the above. The NaN is a separate, still-open bug, and #315 should not be closed on the strength of that fix.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Verified state of this bump, and what it still costs

For the record, since this PR has been open a while and its title predates the last few commits.

Landed on this branch and verified: the corpus/manifest repair (420b4e6f), the HRX planner fix (3d0cca89a0 in hrx-system), the JIT launch-program cache fix (39588ef9) and the mul_mat_id routing fix (56c3c8a3, which takes -o MUL_MAT_ID from 78/108 to 108/108). tools/check_pins.py and tools/registry_build.py --check-pins pass.

What the bump still costs — measured in fresh build directories with the engine's own ExternalProject arguments (round 77's lesson: never measure HRX through an incrementally rebuilt nested build directory; that mistake cost several rounds):

build full gate test-backend-ops -b HRX0 GLM-4.7-Flash, 269-token, -ub 512
pre-sync core + our ggml-hrx 1073/1073 0 NaN, ' Paris.'
AMD's core + our ggml-hrx 1073/1073 — 0 FAIL 0 NaN, ' Paris.'
AMD's core + AMD's ggml-hrx 1019/1071 — 54 FAIL 18 NaN, ??????

The core sync is sound; AMD's ggml-hrx carries the 54 failures (40 MUL_MAT, in iq1_s/iq1_m/iq3_xxs/mxfp4/pq2_0/ptq1_0/q2_*) and the all-NaN logits of engine#315. Putting our backend on the same core removes both.

Ready but unpushable: branch 1bit/amd-core-our-hrx (local 3d8dbcd7a) = AMD's core + our ggml-hrx, hrx-system at 98d05d94. GitHub refuses it with "Unable to determine if workflow can be created or updated ... workflows scope may be required", and a workflow edit gives the exact form: "refusing to allow an OAuth App to create or update workflow ... without workflow scope". Ordinary files push fine — I pushed a registry branch today — so one workflow-scoped token unblocks both that branch and the bump-hrx resolution in engine#332.

@bong-water-water-bong
bong-water-water-bong marked this pull request as ready for review October 6, 2026 15:31
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Corrected pin branch pushed — this PR's staged pin imports AMD's ggml-hrx, which is release-blocking

Picked up the branch from the #332 notes and pushed it with the credential that has the workflow scope:

  • 1bit-MONSTER/llama.cpp branch 1bit/amd-core-our-hrx = f2099e9b7 — verified exact by tree, not by eye:
    git diff --cached 522dab47 -- ggml/src/ggml-hrx is empty (the backend tree hash is 919cd999f in both),
    and git diff --cached 56c3c8a3 -- . ':(exclude)ggml/src/ggml-hrx' is empty (nothing outside the backend moves).
    So it is 56c3c8a3's core with 522dab47's backend, which is the branch the bump-hrx: the scheduled HRX pin sync is failing #332 measurement names.

Why this matters here: the staged pin takes AMD's ggml-hrx, measured at 1019/1071 on the full
test-backend-ops -b HRX0 gate (54 FAIL) with 18 NaN events
on a GLM-4.7-Flash 269-token prompt at -ub 512.
On the same core, our ggml-hrx gives 1073/1073, 0 FAIL, 0 NaN. So #329 as it stands should be treated as
superseded by the corrected bump, not merged.

The corrected engine bump is: third_party/llama.cpp → f2099e9b7 (1bit/amd-core-our-hrx), third_party/hrx-system
unchanged at 98d05d94 — AMD's hrx-system does not export loomc_amdgpu_runtime_global_flags_t, which our
loom-jit.cpp needs.

Also up: #335 (fix/bump-hrx-resolve-ours) resolves conflicts inside ggml/src/ggml-hrx/ toward our line in the
scheduled sync, so this stops recurring rather than being fixed once by hand.

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Corrected pin is now up as a draft: #336 — third_party/llama.cpp → f2099e9b7 (1bit/amd-core-our-hrx), third_party/hrx-system unchanged at 98d05d94. Both pin gates pass (registry_build.py --check-pins; check_pins.py origin/main → 522dab478 -> f2099e9b7 is ahead). Treat this PR as superseded by #336 — its staged pin takes AMD's ggml-hrx (1019/1071, 54 FAIL, 18 NaN) rather than ours (1073/1073, 0 NaN).

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Closing as superseded by #336. Do not land this PR.

Your measurement is the decisive one, and it inverts this PR's premise: resolving the 15 conflicting
ggml/src/ggml-hrx/ files toward AMD's side imports AMD's backend into the release pin, which is #315.

I verified #336's composition independently, by tree hash rather than by eye:

ggml/src/ggml-hrx @ 522dab47   = 919cd999f5709c16ebb3b856e26aff695bf704c5
ggml/src/ggml-hrx @ f2099e9b7  = 919cd999f5709c16ebb3b856e26aff695bf704c5   <- byte-identical
files differing from 522dab47 within ggml-hrx   : 0
files differing from 56c3c8a3 outside ggml-hrx  : 0

So f2099e9b7 is genuinely AMD's core with our ggml-hrx, and this PR is genuinely AMD's core with AMD's.
I also had a correction to make on my own record: I claimed test-backend-ops -b HRX0 could not run in this
repo. That was wrong in the way that matters — it has no target in the engine's top-level ninja graph (which is
where my search stopped), but it exists and runs in the llama.cpp ExternalProject build tree with
LLAMA_BUILD_TESTS=ON, which is how you got 1073/1073. The item in bump-hrx.yml's PR body is legitimate and
decisive, and it should stay there.

Two things from this branch that are worth keeping, in case they are not already in #336's line:

Full write-up and the six harness bugs I hit (all mine, none of them the pins) are in
pm-salvage/PIN329-VALIDATION.md.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant