Repository navigation
ggml-hrx: decode-split FA always-stage fallback (engine#314) - #96
Closed
bong-water-water-bong wants to merge 381 commits into
Closed
bong-water-water-bong wants to merge 381 commits into
bong-water-water-bong wants to merge 381 commits into
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
convert: register sibling HF names for architectures the runtime already has
… reducer (engine#115) The multipass reducer wrote the per-block normalization scale back into the global partial_max transient and re-read it in the output pass. Replace that round-trip with a recompute in the output pass (scale = exp(block_max - max) from the intact partial_max) and add the unroll/schedule hints the reference reduce_f32 output pass carries. Removes the prime suspect for the >2048 silent corruption + intermittent HSA fault.
…ntion, converter)
Zyphra ZAYA1-VL-8B is the ZAYA1 language model with a Qwen2.5-VL vision tower
and two additions:
- vision-only LoRA, used on image tokens: rank 8 on CCA's q, k, both value
projections and o_proj; rank 32 on every expert's fc1 and fc2. mtmd decodes
an image as its own ubatch of embeddings, and graphs are only reused between
ubatches of the same kind, so a ubatch of embeddings takes the LoRA for every
token. The experts' LoRA needs the hidden state between fc1 and the SwiGLU,
so image ubatches run the experts explicitly (same top-1 choice and weight
as build_moe_ffn); o_proj's LoRA reads the attention output before wo.
- bidirectional attention within an image: the mmproj key
clip.vision.decode_non_causal makes mtmd decode image chunks non-causally,
for any vision projector (Gemma 3/4 decide it by projector type).
The converter reads ZAYA1-VL's checkpoint (layers.{i}.attn / layers.{i}.mlp,
the legacy Megatron-style tensors, rotary_base, rope_pct), stacks the
experts' LoRA, writes the ranks, fills the vision config's Qwen2.5-VL
defaults, and writes the mmproj (temporal patch 1: the second patch conv is
zeros). Its chat template takes llama.cpp's media markers (<__media_<id>__>,
random per server) as well as Zyphra's list form, and puts images in front of
the user turn as Zyphra's does.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zaya: ZAYA1-VL (vision LoRA on image tokens, bidirectional image attention, converter)
HRX: multi-pass KV-block reduction for flash-attention decode-split (engine#115)
… it on the GPU The grouped conv ran as a batched matmul with the sequences in dim 3. Vulkan cannot broadcast src0 over dim 3, so whenever a ubatch can hold several sequences (llama-server reserves for its parallel slots) both tap matmuls of every layer fell back to the CPU: 162 graph splits. With the sequences folded into the token dimension the weight broadcasts over dim 2 only: 2 splits. ZAYA1-8B Q4_K_M, llama-server -np 4, four concurrent 128-token requests: 50-66 -> 135-160 tok/s total on Vulkan0 (one request: unchanged). Unchanged: wikitext perplexity 21.5731 (Vulkan0, -b 512) and 21.5518 (HRX0); CPU 1 vs 4 sequences per ubatch identical (25.1372, 8 chunks); F16 vs transformers FP32 95/96; test-llama-archs -a zaya OK on Vulkan and CPU.
zaya: fold sequences into tokens in the grouped conv, so Vulkan keeps it on the GPU
…, argsort, narrow get_rows, strided copy, broadcast repeat) New loom kernels with their dispatch registrations, for ops HRX sent to the CPU: - ggml_grouped_mul_mat_f16_f32: MUL_MAT with a batched F16 weight (one matrix per group, no broadcast), such as ZAYA's grouped convolution. This also lets the loader place such weights on HRX. - ggml_softmax_rows_f32 (no mask, scale 1), ggml_sum_rows_f32, ggml_argsort_rows_f32 (rank per element, ties by index), ggml_get_rows_small_f32 (rows narrower than 4 or not a multiple of 4; strided ids, as top-k views are), ggml_copy_strided_f32 (CONT of strided views, and REPEAT that only broadcasts, with zero strides). Each matcher claims only cases the existing kernels do not take. AMD files get one-line hookups: CMakeLists, dispatch-common, REPEAT in the declared ops. test-backend-ops -b HRX0, all against the CPU: MUL_MAT (batched F16 cases), SOFT_MAX 10/10, SUM_ROWS 6/6, ARGSORT 48/48, GET_ROWS 17/17, CONT 2/2, CPY 53/53, REPEAT 5/5.
- One token per sequence: the conv input transpose is a reshape, not a copy. - MEAN over the query-head group as SUM_ROWS + SCALE. - Qpre/Kpre reshape the contiguous matmul outputs instead of copying them. With the block entirely on HRX the copy made prefill wrong (perplexity 71.8). With the HRX kernels of the previous commit, ZAYA's decode graph on HRX0 goes from 641 graph splits to 1.
ggml-hrx: kernels for ZAYA's router and grouped conv; ZAYA1-8B decodes entirely on HRX (25.5 → 47.9 tok/s)
router_top8_f32 seeds each lane's argmax with best_id = 0x7FFFFFFF and only replaces it when a candidate is strictly greater (ordered ogt) or wins the equal-value tie-break (oeq). When a lane's candidate logits are unordered (NaN) or below -FLT_MAX, neither fires and the sentinel is published verbatim into route_ids. Consumers (e.g. qwen3_moe_routed_gate_up_swiglu_q4k_q8) treat the id as bounded via index.assume -- an optimizer hint, not a check -- and compute expert * weight_expert_bytes in 32 bits, so 0x7FFFFFFF wraps to 0xFFF28000 (~4.29 GB) and the kernel issues a global read past the end of the weights buffer: HSA_STATUS_ERROR_MEMORY_FAULT / llama_decode ret=-3. Seed with the lane's own first expert id instead. lane_expert_base is always < expert_count for a real lane, the result is bit-identical whenever a lane does find a winner (the seed only participates in the first comparison), and a no-winner lane now publishes a valid id whose selected logit is -FLT_MAX, so its route weight normalizes to 0 and contributes nothing. Verified on gfx1151 (Strix Halo), Qwen3-Coder-30B-A3B-Instruct-Q4_K_M, HRX0, llama-bench -p 0 -n 8: -d 2100 0/5, -d 3000 0/3, -d 4800 0/3, <=2048 path 0/3, and greedy output identical to the CPU backend of the same build. Fixes 1bit-MONSTER/engine#123 (the fault is not in the decode-split multipass path; disabling that dispatch only changed the layout enough to mask it).
hrx: never publish the MoE router no-winner sentinel as an expert id
…g expert engine#123's residual, after #25 stopped the 0x7FFFFFFF sentinel from faulting: when the driver migrates a page behind in-flight HRX work (the engine#140-class nondeterminism), a row's router logits go all-NaN. The argmax seed is a finite -FLT_MAX, so the row's softmax still produced a plausible uniform 1/route_count and the token decoded with a valid-but-wrong expert, silently. - router_top8_f32.loom: when no lane of the row found an ordered candidate, publish the row's own NaN instead of that masked uniform weight, so the corruption reaches the logits rather than being hidden behind a plausible one. - common/sampling.cpp: refuse to sample NaN logits - log and abort loudly. NaN is never a legitimate logit, unlike -inf, which masking legitimately uses. - dispatch-flash-attention.cpp: GGML_HRX_FA_PARTIAL_ALIGN selects the decode-split partial transients' alignment (default 4096, production unchanged); 256 reproduces the engine#123/ggml-org#140 rig oracle without editing stress literals. Verified on the MoE repro (Qwen3-Coder-30B-A3B-Instruct Q4_K_M, 2113 ctx, fresh server per sample, under memory pressure): before, 6/13 samples silently decoded ' Paris???????????????'; after, the corrupted samples abort with "HRX returned NaN logits ... refusing to sample a silently wrong token (engine#123)".
hrx: make the all-NaN router-logit case loud instead of a silent wrong expert (engine#123)
LlamaHfVocab left <|im_start|> (105) and <eos> (1) NORMAL although tokenizer.json marks them special. llama.cpp only matches special-token text for CONTROL and USER_DEFINED tokens, so every chat turn's <|im_start|> reached the model as seven text tokens and ZAYA1-8B answered off-template (12 + 30 -> 22). Tokens tokenizer.json lists as special are now CONTROL. Converting ZAYA1-8B now gives the published GGUF's 1283 tensors unchanged and token types with 1 and 105 CONTROL.
convert: ZAYA special tokens are CONTROL, so chat templates tokenize
…stores (engine#123, ggml-org#140) Every decode-split reduce_fused variant (direct, cooperative, multipass, in both the ggml.* and qwen3_moe corpora) writes the reduced attention output to global memory and then packs it into next_q8 with pack_completed_q8, which reads that output back with a different wave/lane mapping. The barrier between them was kernel.barrier<workgroup>, which orders workgroup (LDS) memory only; it compiled to s_waitcnt lgkmcnt(0); s_barrier, with no store-completion wait before it and no buffer_gl0_inv after it. A pack wave could therefore read output before another wave's stores landed, or from a stale L0 line (WGP mode, or after a CWSR restore onto another CU), and quantize the arena slot's previous contents. The f32 output stays correct but next_q8, which the following matmul consumes, is wrong, and NaN when the stale bytes decode as NaN. That is the source of the NaN router logits behind engine#123 and of the run-to-run divergence in engine#140. Queue preemption from any process's KFD eviction opens the window, which is why both issues tracked page migration and GPU contention. Use kernel.barrier<global> scope(workgroup) ordering(acq_rel) at the five sites. Measured on gfx1151 with a separate process that only forces KFD queue evictions: - Qwen3-0.6B, 24-token greedy: 0 of 354 divergent (before: 13-34 per ~177) - Qwen3-0.6B, forced compaction + mlockall: 0 of 118 (before: 42-62 %) - Qwen3-Coder-30B, 2113-token prompt, GGML_HRX_FA_PARTIAL_ALIGN=256: 16 of 16 correct, 0 NaN (before: 1 of 8 correct, NaN guard fired 7 times) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…arrier hrx: order the decode-split q8 pack after the reduce's global output stores (engine#123, ggml-org#140)
…23 follow-up) A Loom kernel.barrier fences only the memory space it names. #28 changed the barrier before pack_completed_q8 from <workgroup> to <global>, which orders the reduce's global output stores but drops the LDS fence the cooperative and multipass reducers relied on for their staging buffers. Keep both: a <global> barrier followed by the original <workgroup> barrier (HIP __syncthreads semantics) at the five pack sites. The standalone two-dispatch reducer (ggml_flash_attention_decode_split_reduce_f32 and its qwen3_moe copy) had the same bug as #28: wave 0 rewrites partial_max in global memory with the exp scales and every wave reads it back after a <workgroup>-only barrier, so the reads could hit stale GL1/L0 lines holding the block maxima. It is not dispatched today. Add the <global> barrier there too. Compiled ISA, gfx1151: every pack handoff (direct, cooperative, multipass) is now vmcnt(0), vscnt(0), s_barrier, buffer_gl1_inv, buffer_gl0_inv, s_barrier; the standalone reducer gains buffer_gl1_inv/buffer_gl0_inv (it had none). Measured with a separate process forcing KFD queue evictions: Qwen3-0.6B 0 of 118 divergent; Qwen3-Coder-30B, 2113-token prompt, partial alignment 256: 8 of 8 correct, 0 NaN. Decode speed vs #28 within run-to-run noise (0.6B d0 276-287 vs 296-300, d1000 204-215 vs 199-207; 30B d2100 57-61 vs 59-62 tok/s). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rs-both-spaces hrx: fence both global and LDS at the decode-split barriers (engine#123 follow-up)
The multipass reducer's final pass gave each workitem one output channel and walked every KV block serially, recomputing expf(partial_max - maximum) per (block, element). Only 128 of 256 workitems are live at value_head_size=128, so that pass was the whole residual >2048 decode gap. Compute the per-block scale once in the lane-strided sum pass into a per-row LDS stage (also publish each row's sum), then run one vectorised output pass over all query rows at once: each workitem owns a 4-channel quad and accumulates vector<4xf32> over blocks. The per-channel block order is unchanged, so the reduce output is bit-identical. Measured (gfx1151, Qwen3-Coder-30B-A3B-Instruct-Q4_K_M, -dev HRX0, median of 6 interleaved llama-bench -p 0 -n 8 -r 5 runs): d2100 +15.4%, d3000 +19.8%, d4800 +22.0%; d2100/d2000 0.881 -> 0.976. Untouched <=2048 controls +1.1%/+4.2%. 1.24 GB decode-path logits byte-identical to the previous build; code word ZX-4718-QQ exact at 4700 tokens; 0 GPU faults. benchmarks/NOTE-hrx-124-multipass-output-2026-09-27.md has the full table, method and repro.
HRX: vectorise the decode-split multipass output pass (engine#124)
…indices ggml-hrx: preserve IQ4 table indices after Loom lookup fix
…em/llama.cpp # Conflicts: # tests/test-hrx-ops.cpp
Laya's typed-decisions checkpoint (ggmlc's GGUF) now runs entirely on HRX: 95.5% on the engine's 200 routing cases, as on Vulkan and the CPU scorer. New LOOM kernels in our small_rows_f32.loom, matchers in our dispatch-small-rows.cpp, all below the existing kernels (priority -10) so they only take what those refuse: - NORM (LayerNorm without affine), one workgroup per row - ADD/SUB/MUL/DIV with strided or broadcast inputs - CLAMP, standalone and in place (ggml_clamp's output views its input) - CPY F32 (strided) -> F16 - FLASH_ATTN_EXT for one-block-per-head layouts (online softmax per lane) - MUL_MAT of small F16/F32 weights (any K) and Q8_0 weights (K % 32) AMD files: NORM in eager_capability_declared, its epsilon in op-params, and the mul_mat_postops matcher no longer takes K % 256 != 0 (its kernel's config constraint rejected those at prepare time). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ggml_attention_rows_f32_f16: one 64-lane workgroup per (query, head) stages the query row and a score row in workgroup memory (each score computed once, not once per value lane), workgroup max/sum reductions, then lanes over the value dims. Up to 2048 keys, head sizes <= 256; the per-lane kernel stays for the rest (ONEBIT_HRX_ATTN_PER_LANE=1 forces it). - ggml_mul_mat_rows_q8_0_f32: Q8_0 with K >= 2048, one workgroup per output, lanes over the K blocks (ONEBIT_HRX_Q8_PER_OUTPUT=1: the old kernel). Laya on HRX0, steady state at 64 tokens: 40 -> 32 ms a decision; 95.5% on the 200 routing cases unchanged. test-backend-ops HRX0: FLASH_ATTN_EXT 204/204 (with plain K/V tensors), MUL_MAT 209/209, NORM 10/10, ADD 22/22. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ggml_rope_rotate_half_f32: the eight-node RoPE graph compilers emit (x*cos + CONT(CONCAT(NEG(CONT(x[half:])), CONT(x[:half])))*sin) in one kernel, matched from its first node (x*cos) so the fused match covers the rest; the two half views must sit at x's offset and half a row further. - ggml_geglu_strided_f32: CONT(gate view) -> GELU -> MUL(., up view) in one kernel, GELU in ggml's tanh form. With ggmlc giving its graphs a uid (1bit-MONSTER/ggmlc 1bit/main), so HRX's graph-program cache and graph replay hit instead of re-importing and re-recording every call: Laya on HRX0 at 64 tokens 32 -> 15-16 ms a decision (Vulkan 12 ms), 128 tokens 25 ms (Vulkan 24-31). 95.5% on the 200 routing cases, probabilities within ~1e-3 of Vulkan's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ggml-hrx: kernels to run Laya (ModernBERT encoder) on HRX0
…code) At decode the WMMA mul_mat_id kernel reads a single token's expert at ~45 GB/s (ZAYA1-8B: 80 calls x 85.6 us = 40% of a token's device time). The new ggml_mul_mat_id_decode_f32_wave64 (ops/mul_mat_id_decode_f32.loom, ours) runs the decode GEMV's per-block dequant/dot loop from ops/mul_mat_f32_f32_decode.loom (linked as a library) over rows expert*N + row, one wave64 per (row pair, token*slot); the router id is clamped to the expert count. Registered at priority 150 for <= 4 tokens and <= 64 token*slot rows, below the fused routed-FFN and SwiGLU paths. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CMAKE_CXX_STANDARD 17 -> 26 in ggml/CMakeLists.txt and the top level
(CMAKE_CXX_STANDARD_REQUIRED kept); CMAKE_CXX_SCAN_FOR_MODULES OFF, so no
clang-scan-deps is needed (TheRock amdclang has none).
Source fixes for what C++26 rejects or changes:
- src/models/models.h: declare the explicit specializations of the
eagle3/dflash/t5 graph members before first use. make_unique is constexpr
since C++23, so clang instantiates it where build_arch_graph calls it,
before the specialization was declared ("explicit specialization after
instantiation"; ill-formed NDR in every standard).
- common/chat.cpp: build std::optional<json> in place (json -> optional<json>
is ambiguous with libstdc++ 16 at C++26).
- tools/mtmd/clip-graph.h, tools/mtmd/models/qwen3vl.cpp,
tests/test-backend-ops.cpp: cast ggml_scale_flag to an integer before
combining it with ggml_scale_mode (enum|enum of different types is
ill-formed since C++26, P2864).
- std::to_string(float/double) prints std::format("{}") since C++26 (P2587):
test-jinja failed 5 cases ({{ -1.0 }} -> "-1", 42|float -> "42"; 100.0
would print "1" and 0.0 an empty string through the trailing-zero trim).
Every floating-point to_string call site (found by a deprecated-overload
scan of all TUs with both compilers) now prints "%f" as before:
common/jinja/{string,value}.h, src/llama-impl.cpp, tools/mtmd/clip-impl.h,
tools/llama-bench/llama-bench.cpp, tools/server/server-schema.cpp,
tests/test-backend-ops.cpp.
- ggml-hrx kernel-executable-cache.cpp: std::atomic<std::shared_ptr> instead
of the deprecated atomic_load/store_explicit(shared_ptr*).
- ggml-backend-reg.cpp: path from a UTF-8 string via std::u8string instead of
the deprecated fs::u8path (C++17 keeps u8path).
- vendor/nlohmann/json.hpp: std::is_trivial (deprecated in C++26) spelled as
is_trivially_default_constructible && is_trivially_copyable.
Builds (strixhalo): amdclang 23.0.0git (TheRock) HRX build, with and without
the HIP add-on, and g++ 15.2.0 CPU build: 0 errors, 0 warnings from fork
sources (the add-on's own warnings and its two glm harness targets, which
hit clang 23 cuda_wrappers/new vs libstdc++ 16 constexpr placement delete at
C++26, are left to the add-on). ctest (g++): same results as C++17
(test-tokenizers-ggml-vocabs fails in both: LFS pointer vocab files).
Gates (performance mode, HRX0, C++17 tip f95f2db = A vs this = B):
- test-backend-ops -b HRX0 MUL_MAT 325/325, MUL_MAT_ID 108/108,
MUL_MAT_VEC_FUSION 57/57, ROPE 42/42, RMS_NORM 6/6, SOFT_MAX 8/8,
CONCAT 6/6, SSM_CONV 47/47; x2, both builds.
- Qwen3-0.6B Q4_K_M c512 16 chunks: saved logits byte-identical A vs B,
PPL 26.1766 both, KLD 0.
- 8 identical greedy requests: 1/8 distinct on Qwen3-0.6B and ZAYA1-8B,
output identical to the C++17 build.
- llama-bench medians of 9 (3 interleaved runs x -r 3), B vs A:
Qwen3-0.6B pp512 -0.31% tg128 -0.08%; Qwen3-8B +0.42% / +0.11%;
GLM-4.7-Flash -0.12% / +0.05%; ZAYA1-8B -0.05% / +0.15%.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
build: C++26 by default (amdclang 23 / g++ 15), no module scanning
ggml-hrx: graph inputs are never written back to host; non-replay programs wait for their input upload
… C++17 CI's GNU toolchain has no CXX26 dialect (FindOpenMP try_compile failed), so the standard follows CMAKE_CXX_COMPILE_FEATURES: 26 with amdclang 23 / g++ 15, 17 elsewhere. The source fixes are standard-neutral. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
build: C++26 only where the compiler offers it; older toolchains keep C++17
…nuation engine#315 measured 1 fault in 5 runs on GLM-4.7-Flash and killed a full ZAYA1-8B evaluation with 0 of 30 problems recorded. The guard below refuses to sample a silently-wrong token, which is correct, but it aborts on the FIRST NaN -- so the run and the evidence die together, and every recorded occurrence so far reports only 'vocab index 0'. Measure the corruption before acting on it: count, index range, and whether the run is contiguous. That is the discriminator between a migrated/torn page and a single bad element, and it is the fact an earlier probe could not obtain. With GGML_HRX_NAN_CONTINUE set, substitute -inf for the NaN entries and continue. -inf makes those entries unsampleable -- exactly what masking already uses -- while leaving the rest of the distribution intact, so a long evaluation completes instead of dying at the first fault. Off by default, so the guard's existing behaviour is unchanged unless an operator asks for it.
common: record NaN logits instead of only aborting, with opt-in continuation (engine#315)
First successful sync since 2026-10-01: bump-hrx had failed on every scheduled
run because it rebased, and replaying our 159 commits died at commit 1 of 159.
Merged instead, so one bounded resolution replaced the replay.
Resolution policy, per owner direction ("we use all the pinned sources" /
"people want bleeding edge"): keep both sides, newest wins on overlap.
- 11 content-conflicted files (CMakeLists, dispatch-mul-mat*, dispatch-flash-
attention, kernel-executable-cache, dequant.loom, matmul_id motifs,
mul_mat_q5_k_q8_plane_wmma.loom, router_top8_f32.loom): conflict blocks taken
from AMD's side, the auto-merged remainder preserved so our non-conflicting
work stays.
- loom-libs/manifest.json: semantic union, AMD's entries authoritative on the
90 shared symbols. files 87+216 -> 277 (61 of ours kept that AMD lacks);
exports 145+160 -> 215 (55 of ours kept). AMD's new sources/kernels/
link_modules/plan_cases keys retained.
- ops/flash_attention_decode_split_f32_f16_wmma.loom and
ops/flash_attention_f32_f16_wmma.loom (and the qwen_moe copy): superseded by
AMD's refactor into ops/flash_attention/* + motifs/flash_attention/*
(online_softmax, qk_block_wmma, pv_block_wmma, decode_partials, decode_reduce,
publish_q8, publish_f32, completion_counter_reduce). The symbols still exist at
the new paths, so the dispatch names resolve; keeping the monoliths would have
double-defined them. The pre-merge tip is tagged so nothing is lost.
NOT YET VALIDATED: this tree has not been built or run. GitHub-hosted CI builds
without HRX. Before any engine pin bump that consumes this commit, build it on
Strix Halo and run tests/serve_e2e.sh for hrx and cpu, plus
test-backend-ops -b HRX0.
common: record NaN logits instead of only aborting, with opt-in continuation (engine#315) -> pinned line
…ged corpus has
The AMD sync left this inconsistent. AMD's pin deleted
kDecodeSplitMaxKeyValueTokenCapacity and their kernel refactor removed the
multipass reduce provider, but the merge kept our 32768 constant and our own
comment warning that "nothing covers above 2048" and to "match this bound".
With only direct_f32 (64-256) and cooperative_f32 (257-2048) in the corpus, a
capacity of 2049-32768 offers a dispatch with no provider, so the kernel
selector rejects every candidate ("all_rejected") and the whole decode fails --
the exact failure that comment describes. Long-context decode is affected.
Capped at 2048 so the scheduler falls through to the general
flash_attention_f32_f16_wmma dispatch: still correct, just not split-parallelized
for very long decode contexts (AMD's stated trade-off).
NOT YET VALIDATED: no build or run against this.
The AMD sync auto-merged our kernels into AMD's refactored corpus. Shared motifs ended up with duplicate SSA names, 47 manifest paths had no file, 19 of our exports had no source, and dispatch-mul-mat.cpp kept our body with AMD's include block. None of it compiled. Base the corpus on AMD's pin and keep our kernels under ours/ with their own recipes; rebuild manifest.json as AMD's exports plus our 55, with no shared primary carrying two library lists; re-add the includes the merge dropped, the transient-reuse guard the merge left out of the source list, and the definition of common_q8_prefill_relaxed the merge kept only its callers of. Verified on Strix Halo with -DONEBIT_HRX=ON: cmake --build build --target onebit succeeds with 0 errors, and test-backend-ops -b HRX0 runs to completion. Assisted-by: DeepSeek
runtime/loom-jit-disk-cache.cpp serialises magic, version, launch_config, workload_argument_count, the hsaco and the manifest - and never the host-side launch program. A cache hit therefore restores a result whose launch_program is null, and kernel-executable-cache.cpp rejects it through the `launch_program == nullptr` disjunct while printing "compiled ABI does not match manifest" - an error that names the wrong subsystem and hides the real one. Skip storing any result that carries a launch program (the format cannot hold it), bump kVersion 1 -> 3 so entries written by the broken code are ignored, and point the two kquant exports whose primary is ours/ops/kquant_decode_f32.loom at our own motif files, plus repoint ggml_flash_attention_f32_f16_wmma at our pre-merge kernel (AMD's caps qk_head_size at 512; GLM-4.7-Flash needs 576). Measured on Strix Halo, GLM-4.7-Flash-Q4_K_M: before: "compiled ABI does not match manifest" 57x, 15 distinct kernels, no run after : 3 -> 0 for AMD's kernels; pp512 896.84 t/s, tg32 28.05 t/s serve_e2e.sh ... hrx on Qwen3-0.6B still passes (32 chunks, answers "Paris"). test-backend-ops -b HRX0 is unchanged at 989 OK / 84 FAIL - a separate cause.
The AMD sync pointed the two matcher constants at AMD's tiled/skinny
mul_mat_id kernels, which mishandle several token counts. Our pre-merge
dispatcher had no such split - one ggml_mul_mat_id_f32_f32_wmma for every
token count - and our exports are still in the corpus, so this is a dispatcher
change only, with no manifest edit and no kernel port.
Measured on Strix Halo:
test-backend-ops -b HRX0 -o MUL_MAT_ID
before 78/108 passed, 32 failures at n = 5, 17, 32, 129 with ERR ~ 1.0
after 108/108 passed
GLM-4.7-Flash-Q4_K_M, default -ub 512
before test_gen: failed to decode generation batch (all-NaN logits)
after pp512 902.21 t/s, tg32 27.40 t/s
The second row is engine#315: its all-NaN logits fired for every ubatch of
>= 256 tokens, which is why a -ub 128 workaround existed. GLM now generates at
the default ubatch, so the workaround is no longer needed.
Pointing only the tiled constant at our kernel gave 90/108 and fixed n = 17, 32
and 129, but left n = 5 at ERR ~ 85 - garbage rather than an unwritten output -
because the skinny constant still named a different interface. Both must name
the same kernel, as they did before the sync.
AMD's core changes (everything in llama.cpp/ggml outside ggml/src/ggml-hrx) plus our HRX backend (ggml-hrx as of 522dab4) together are a working configuration, measured in fresh build directories with the engine's own ExternalProject arguments (Release, amdclang, HRX_SOURCE_DIR=hrx-system 98d05d94): test-backend-ops -b HRX0 -o MUL_MAT AMD core + our ggml-hrx 322/322 libggml-hrx.so 5,388,432 AMD core + AMD ggml-hrx 282/322 libggml-hrx.so 6,518,584 GLM-4.7-Flash, 269-token prompt, -ub 512, three requests AMD core + our ggml-hrx nan = 0, reply " Paris." The 40 MUL_MAT failures and the engine#315 all-NaN logits are properties of AMD's ggml-hrx kernel/dispatch set, not of the core sync: with our backend in place on the same core both disappear. hrx-system stays at 98d05d94. ---
Simpler alternative to 1bit/fa-shared-visibility: replace the two subgroup votes with constant predicates (all_visible=false, any_visible=true), so the kernel always takes the existing staging path and never issues a collective. Same emission failure addressed, 6 lines instead of 102, but it pays the staging loop and two subgroup barriers on every key tile instead of only the tiles that need masking - a correctness fallback, not a free fix. Verified: compiles (0 collectives remain); bit-identical output to the pristine kernel and to the shared-visibility variant on a masked-KV check where only the masked rows' V changes (out0=0.0100021362, sum=61.453125). Not verified against the authoritative masked-KV tool; the emission failure does not currently reproduce on demand.
Author
|
Superseded — this change landed on the pin line as So this draft no longer needs review. Leaving the branch in place for reference; close whenever convenient. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
engine#314 candidate 2: always-stage — the small alternative to #95
The same emission failure as #95, addressed with +6/-2 instead of +102/-2. Replaces the two subgroup votes with constant predicates (
all_visible = false,any_visible = true) so the kernel always takes the existing staging path and never issues a collective.1bit-MONSTER/engine#314has the full comparison. Short version:fa-shared-visibilityfa-always-stageout0=0.0100021362 sum=61.453125This is the fallback: simpler and obviously correct, at the cost of the staging loop and two subgroup barriers on every key tile. Pick one, not both — they are alternatives, not a stack.
Not benchmarked (no machine headroom), and not verified against the authoritative masked-KV tool.