#361 (c4d9a9b) is titled "Bump HRX: llama.cpp 02f2880c9f42 + hrx-system
4ba76c18eafe (AMD's tested pair)" but its diff changes only
third_party/hrx-system. The llama.cpp gitlink stayed at e44c9d01a4d5, whose
ggml/src/ggml-hrx/loom-jit.cpp still uses loomc_amdgpu_runtime_global_flags_t,
loomc_amdgpu_emit_options_t and LOOMC_AMDGPU_RUNTIME_GLOBAL_*, which
4ba76c18eafe removed. The pinned pair therefore does not compile:
loom-jit.cpp:70:5: error: unknown type name 'loomc_amdgpu_runtime_global_flags_t'
loom-jit.cpp:553:1: error: unknown type name 'loomc_amdgpu_runtime_global_flags_t'
... 10 errors, all in loom-jit.cpp
GitHub CI builds the pin without HRX (ONEBIT_HRX defaults OFF), so it cannot
catch this class of break; a from-scratch HRX build on strixhalo is what
surfaced it.
Restore third_party/hrx-system to 98d05d94a9f9, the revision #359 pinned and
which exports the API our loom-jit path needs. llama.cpp is unchanged.
Co-authored-by: bong-water-water-bong <bong-water-water-bong@users.noreply.github.com>
Moves BOTH submodules onto the pair ROCm/ggml-staging-automation now pins, each merged with our own commits (the workflow performs the merge; a conflict fails the run and names the files).
e44c9d01a4d5e44c9d01a4d5(on AMD02f2880c9f42)98d05d94a9f94ba76c18eafe(on AMDc0b135a778cc)Our commits, rebased onto AMD's pin:
e44c9d01Merge commit '02f2880c9f422dd1df9db11e3a0bb442bc5f11f5' into HEAD26330cc2merge: take shared-memory tile visibility over always-stage for decode-split FA (engine#314)cf4dfd80ggml-hrx: decode-split FA always-stage fallback (engine#314)f9ff0667ggml-hrx: derive decode-split FA tile visibility from shared memory (engine#314)f2099e9bhrx: keep our ggml-hrx backend on AMD's core56c3c8a3ggml-hrx: route mul_mat_id through our kernel again39588ef9ggml-hrx: do not cache a compiled kernel whose launch program exists420b4e6fggml-hrx: take AMD's corpus wholesale and keep our kernels beside it51e538fahrx: cap the decode-split dispatch at 2048, the last provider the merged corpus hasb04c4e95Merge pull request Model registry and the daily HF census (PORTING step 5) #94 from 1bit-MONSTER/1bit/cxx26-cib3981ce9Merge commit '59771f2a63c30c0edb9eaa79bef8b1b937788733' into HEAD59771f2aMerge AMD's pinned pair b802a507 into 1bit/hrx-vulkan-patchedb05d329fMerge pull request Laya scorer: 8.75 s -> 0.38 s per decision, byte-identical output #93 from 1bit-MONSTER/1bit/hrx-nan-continue4d13516fcommon: record NaN logits instead of only aborting, with opt-in continuationebe16621Merge pull request README, PORTING: the Laya router landed as an opt-in (#90, #91) #92 from 1bit-MONSTER/1bit/cxx26-ci1d4333c3build: C++26 only where the compiler offers it; older toolchains keep C++17522dab47Merge pull request Laya router: the step-4 scorer in C++, routing each request to a device #90 from 1bit-MONSTER/1bit/kv-nan-rootcause6150ff09Merge pull request 1bit route: print the scorer's load and decision times #91 from 1bit-MONSTER/1bit/cxx2669788ac5build: C++26 by default (amdclang 23 / g++ 15), no module scanningf862daedggml-hrx: wait for a non-replay command program that reads registered host memoryb4e75159ggml-hrx: never write a device-rewritten graph input back to host memoryf95f2dbdMerge pull request GPU: public again — re-pin our HRX and ROCmFPX forks (revert #87) #89 from 1bit-MONSTER/1bit/hrx-runtime-overheadefb93ceaggml-hrx: record graph kernels with a bounded lookahead to share barriersd9594130ggml-hrx: transient arena reuse without false dependenciese57beb97Merge pull request NPU private routes: optional routes decline to the fast lane #88 from 1bit-MONSTER/1bit/qwen4exp-hrx001832c3Merge pull request README, CONTRIBUTING: 1bit-MONSTER is private now #86 from 1bit-MONSTER/1bit/hrx-hip-plumbing3527b569ggml-hrx: HIP matchers can declare the ops they claim (hip/hip-capabilities)7740a1adqwen4exp MTP: load and run the Flash-Next draft head in this treef89d5037qwen4exp: run Qwen3.8-Flash-Next on HRX (graph adjustments, dense QSA under top-k width)bfda5f58qwen4exp MTP: refuse a draft vocabulary whose d2t ids are out of range or repeatedf294a5b7qwen4exp MTP: reduced draft vocabulary (d2t + K-row head)1458fd80qwen4exp: NextN/MTP draft head for Qwen3.8-Flash-Next (ggml-org/llama.cpp PR #28243)864deaf6qwen4exp: add hc ops (#28901)34ff68b8qwen4exp: enable rms_norm + mul fusion (#28896)7d028677models : fix GDN normalization frommaxtorsqrt(#28068)4bb1e268kv-cells: look up the n-gram history in the sequence position index (#28040)ffca4b4fqwen4exp: fix seq_cp, block position keying, mtmd input, cuda abort, add tests (#27941)2c684b57qwen4exp: support recurrent state rollback (#28123)6354fe94qwen4exp: sum the indexer heads by slices (#28023)6378935bcommon: rename --tensor-read-lazy to --lazy-mode, add -lzm shorthand (#27969)9175627ellama: improve TENSOR_READ_LAZY handling (#27837)d3516bf1bench: add --tensor-read-lazy (#27881)08a12334model: qwen4exp: reduce number of graph splits (#27880)d5a3a1b3model: add Qwen3.8-Flash-Next (qwen4exp) (#27742)254e0142llama: model_loader: add TENSOR_READ_LAZY (#27794)35207fe9model : allow reshape of tensors during load (#26531)4dade385llama: add token ID tracking to KV cell (#27762)f3c771fdMerge pull request GPU: pin AMD's HRX pair and upstream ROCmFPX; our GPU kernel work behind -DONEBIT_GPU_PRIVATE #87 from 1bit-MONSTER/1bit/hrx-replay-keep2806e9a4ggml-hrx: stop the sleeping wait from oversleeping decode after a promptd1691f40Merge pull request npu-lax: state the parity protocol (NPU routing adopted) #81 from 1bit-MONSTER/1bit/hrx-f32-wmma-accumulated5a9932dMerge 1bit/hrx-vulkan-patched into 1bit/hrx-f32-wmma-accumulate86eee898Merge pull request README: drop the Powered by FastFlowLM line #83 from 1bit-MONSTER/1bit/mtp-pending-h-continuitya679f780Merge pull request bump-xdna: correct the XRT path in the PR's test recipe #82 from 1bit-MONSTER/1bit/hrx-host-writeback-flushd5de4b1ftests: MUL_MAT_VEC_FUSION cases for one-token Q4_K/Q5_K gate/up + SwiGLU8719cf4eggml-hrx: HIP kernel add-on hook (GGML_HRX_HIP_ADDON_DIR)616034b8ggml-hrx: run hipcc code objects through HRX (HIP kernel plumbing)2e5fcf56Merge pull request NPU: move the Qwen3.6-35B-A3B route to a private add-on, keep a generic hook #85 from 1bit-MONSTER/1bit/hrx284-f32acc973e5ce4Merge pull request Bump llama.cpp to c075cc1: HRX0 runs MoE models above 128 experts (Qwen3.6-35B-A3B) #84 from 1bit-MONSTER/1bit/hrx123-decline-guardd336ba65ggml-hrx: fold the MUL_MAT_ID WMMA f16 accumulator into f32 every 16-wide dot1c0241bbhrx: decline the decode-split under a non-production partial-transient layout (engine#123 step 6)b5669e03ggml-hrx: keep the transposed result-tile staging with the f32 accumulatore0f637bfggml-hrx: publish deferred host writebacks before a host-staging upload34e545f5spec: MTP bridges the previous h-row into a fresh sequence (engine MTP: context checkpoints make repeated identical requests give results that alternate every other request #290)85f59832ggml-hrx: the f32_f32 WMMA cores accumulate in f32f5b7f4adMerge pull request README, NOTICE: acknowledge FastFlowLM #74 from 1bit-MONSTER/1bit/hrx-prism-tile-bytesd3dc0f78ggml-hrx: PrismML tile bytes for the low-token SwiGLU (PQ2_0 1.7B decode)44859162Merge pull request docs/lemonade.md: the embedding checklist (measured) #68 from 1bit-MONSTER/1bit/hrx-attention-sinks54054846ggml-hrx: attention-sink dispatch refuses outputs that overlap an input's storagedcbc1d4fggml-hrx: attention sinks for FLASH_ATTN_EXT (gpt-oss), applied after FlashAttention2bd7f585Merge pull request NPU: the build's own lax kernels are the default; Qwen3.6-35B-A3B through Lemonade #72 from 1bit-MONSTER/1bit/hrx-fa-masked-ve6187096ggml-hrx: masked KV rows reach flash-attention P*V as V = +0 (identical requests gave different logits)35027f7bMerge pull request NPU route: reasoning_content, as llama-server returns it #73 from 1bit-MONSTER/1bit/hrx-moe-split-guardba0be9a1ggml-hrx: keep ADD_ID / SWIGLU_OAI with their MUL_MAT_ID (placement guard)388290b7Merge pull request Bump Lemonade to d7c644a9; embedding checklist status #71 from 1bit-MONSTER/1bit/hadamard-onebit-guard63390665Merge pull request 1bit serve --embedding / --reranking (Lemonade checklist #4) #70 from 1bit-MONSTER/1bit/hrx-fa-mla-transpose53c8b179llama-hadamard: fold qwen3next's ssm_ba; refuse stamped files with unfoldable Q4_0 weights56b44534ggml-hrx: no V transpose for MLA's strided V (GLM-4.7-Flash prompts >= 512 tokens)0b56f47bMerge pull request 1bit serve: forward llama-server's other routes (Lemonade checklist #3) #69 from 1bit-MONSTER/1bit/hrx-tq-prf2c62989ggml-hrx: MXFP4 check cases fill bytes 121..127 (the iota must stay a positive i32)74b8dea6Merge pull request Pin Lemonade: third_party/lemonade at the fork's 7650b4f #67 from 1bit-MONSTER/1bit/hrx-gptoss-opsb221e1f4ggml-hrx: scalar reference and check cases for TQ1_0, TQ2_0 and MXFP4 decodeb9c39393ggml-hrx: TQ1_0 decoders never subtract past zero8bd4f890ggml-hrx: TQ1_0 / TQ2_0 weights, and single-token decode lanes for TQ1_0, TQ2_0 and MXFP41bd9ee66Merge pull request 1bit comfy: run the checked binary; fetch the submodule when missing #66 from 1bit-MONSTER/1bit/hrx-mmid-g32-v2d17f05faggml-hrx: SWIGLU_OAI on F32 (gpt-oss clamped SwiGLU)50a63b6cggml-hrx: ADD_ID on F32 (gpt-oss expert biases)40740e99tests: test-hrx-mul-mat-id-k32, MUL_MAT_ID at input sizes that are a multiple of 32 on HRXa8792c33tests: test-hrx-decode-stride, decode-kernel input rows at input sizes that are not a multiple of 2567215efdeggml-hrx: decode kernels read input row t at t * input_size7370b6e7ggml-hrx: MUL_MAT_ID WMMA kernels take input sizes that are a multiple of 32cebcd708Merge pull request ComfyUI.cpp inside the engine: 1bit comfy #65 from 1bit-MONSTER/1bit/hrx-mxfp4-v2adfd541eggml-hrx: MUL_MAT_ID admits only input sizes that are a multiple of 25652b79b96tests: test-hrx-mxfp4 accepts zeros only for rows whose E8M0 scale is subnormalb0d3d46etests: test-hrx-mxfp4, a bit-exact known-answer check of MXFP4 GET_ROWS on HRX68920038ggml-hrx: exact E8M0 half scale for MXFP4 in the shared dequantizer89ddd94fggml-hrx: MXFP4 weights in the shared dequantizere2b946adMerge pull request site: Docs7 analytics #64 from 1bit-MONSTER/1bit/hrx-hadamard-fwht827b486fhrx: Walsh-Hadamard kernel for MUL_MAT hinted GGML_HINT_SRC0_IS_HADAMARDd60cc4f3Merge pull request pr-agent: a bot comment no longer cancels a review #63 from 1bit-MONSTER/1bit/hrx-program-cache-cap6b1061a9Merge pull request NPU lane: fresh hw_context per generate(), fixes the idle timeout #62 from 1bit-MONSTER/1bit/hrx-prism-types93f1381chrx: bound the graph program cache (GGML_HRX_GRAPH_PROGRAM_CACHE, default 64)20751520ggml-cpu: clamp lists PQ2_0 / PTQ1_0 (-Werror=switch on CI)ce709871ggml-hrx: kquant check cases for PTQ1_0 paired with a codebook format (1bit OS: the boot fixes the VM test found (after #56) #60 two-buffer staging)5f2424a6ggml-hrx: PQ2_0 / PTQ1_0 native on HRX0 (one resident copy), Q1_0 on the kquant decode kernelsd0f3cf78ggml: PrismML PQ2_0 / PTQ1_0 types (ids 142 / 143) with a CPU referencebd5b2970Merge pull request 1bit OS: the boot fixes the VM test found (after #56) #60 from 1bit-MONSTER/1bit/hrx-mixgrid326cb7cdhrx: K-quant SwiGLU fills one codebook when gate and up share it (review)68805cc0ggml-hrx: q8_1 x4 prefill takes the 256-aligned head of a remainder chunk (PR-Agent: local coding model on the strixhalo runner #61)2029ecb8ggml-hrx: keep V row-major past copy_transpose_f16's 32768-row range (Build as C++26 #59)8d880f6ellama: load the engine's Hadamard-rotated Q4_0 files (onebit.hadamard_q4_0) through llama-hadamard (Lean: ROCmI4 on Vulkan (our ROCmFPX fork); the bump rebases our commit #58)b00ddff0hrx: K-quant SwiGLU decode keeps separate gate and up codebooks (mixed-format pairs no longer declined)6e42b511ggml-hrx: route Q4_K/Q5_K/IQ4_XS prompt matmuls to the q8_1 x4 kernel (pp512 99 -> 341 on Qwen3.8-27B) (q4nx: document the two Q8 code layouts behind one chunk size #55)dd74f6bbMerge pull request Site SEO: canonical URLs, preview card, JSON-LD, robots.txt, sitemap.xml, IndexNow #54 from 1bit-MONSTER/1bit/hrx-ternary4fbae7aehrx: packed ternary upload accepts all-zero Q4_0 blocks with any scale (review)bf5ad7c9Merge pull request Discord: permanent invite on the site and in the README #53 from 1bit-MONSTER/1bit/hrx-iq1-v230d0701chrx: packed ternary decode for exact-ternary Q4_0 weights (opt-in, GGML_HRX_TERNARY_Q4_0)7de84122hrx: IQ1_S and IQ1_M weights on HRX0 (shared dequantizer, K-quant decode kernels)cde002ddMerge pull request NPU: Qwen3.6-35B-A3B on full ELFs, one runlist submit per token, served by 1bit serve #52 from 1bit-MONSTER/1bit/hrx-gdn-snapshot-ratio7912b2d9hrx: GDN snapshot publish forms decay ratios in log space (fixes NaN after MTP rollback)764a2562Merge pull request blog: correct the MiniMax-H3 note #51 from 1bit-MONSTER/1bit/hrx-iq3xxs-kq5825fdebhrx: IQ3_XXS and IQ2_S weights on HRX0 (shared dequantizer, K-quant decode kernels)d5048ad2Merge pull request context7.json: description under 200 characters (fixes the claim) #50 from 1bit-MONSTER/1bit/hrx-iq2a05c1b65hrx: IQ2 generators: flake8 and pyright clean (output unchanged)0174b32bhrx: kquant decode: block-bytes comment back above its function (review)6cbab7efhrx: IQ2_XXS and IQ2_XS weights on HRX0 (shared dequantizer, K-quant decode kernels)d1747cb7Merge pull request context7.json: claim the Context7 library #49 from 1bit-MONSTER/1bit/hrx-q2kef779b8chrx: kquant header: format 12 is a common weight format now (review)738460eehrx: Q2_K weights on HRX0 (shared dequantizer, K-quant decode kernels)b8d587e0Merge pull request site: GoatCounter visitor counts #34 from 1bit-MONSTER/1bit/hrx-bonsai804a0f75Merge branch '1bit/hrx-vulkan-patched' into 1bit/hrx-bonsaiacf9c747Merge pull request Docs chat on 1bit.gg: Context7 widget + context7.json #48 from 1bit-MONSTER/1bit/hrx-kquant-ready616348a8llama-hadamard: credit PrismML and carry the MIT notice of the code followed8981d2b3llama-hadamard: read the prism.hadamard keys through gguf8bf4fcf3llama: load Hadamard-folded weights (prism.hadamard, Ternary Bonsai)3b954f07hrx: kquant matchers skip q8-only inputs; softplus kernel0d7e3984Merge pull request blog: HRX answers long prompts correctly, and where Unsloth quants pay #47 from 1bit-MONSTER/1bit/hrx-kquant-tokens-vaccef9e3c23hrx: K-quant token kernels accumulate vectors, reduce once per rowdad19addMerge pull request HRX: decode past 256 KV tokens fixed (llama.cpp 79788e9) #46 from 1bit-MONSTER/1bit/hrx-kquant-tokensb57bd75ahrx: K-quant decode kernels for 2-8 tokens (MTP / speculative verify)9e485ebbMerge pull request docs/hrx.md: known issue — --device hrx answers wrongly on mid-length prompts #45 from 1bit-MONSTER/1bit/hrx-kquant-swiglu-priob8061ffahrx: K-quant SwiGLU decode above the int4 lowrow path6617e78aMerge pull request site: the address is 1bit.gg #44 from 1bit-MONSTER/1bit/hrx-sleeping-wait5c6d4765hrx: sleep through long stream waits instead of spinning a CPU coref81814a6Merge pull request blog: speculative decoding, smaller files, one GPU many users #43 from 1bit-MONSTER/1bit/hrx-kquant-q3k316dc223hrx: Q3_K and IQ3_S in the K-quant decode kernelsade79d4cMerge pull request 1bit serve: --parallel, --adaptive, RAG (--embed/--rerank), --mtp-p-min; MMQ-only ROCm build #42 from 1bit-MONSTER/1bit/hrx-kquant-decodefd20a3dahrx: K-quant / IQ4 / Q8_0 decode projections in GGUF block layout8c713849Merge pull request 1bit serve: --mtp (up to 3.4x), --parallel (up to 5.8x), --adaptive (Vulkan + ROCm overflow), --device rocm #41 from 1bit-MONSTER/1bit/hrx-iq4-on-hrx3b053a43hrx: fix the other IQ4_XS remap and the q8-plane SwiGLU placeholder690e9e8cggml-hrx: IQ4_NL / IQ4_XS weights decode correctly (no XOR 12 table remap)c9283cffMerge pull request Bump Laya: source 23a17522, model 55cf4c4e #40 from 1bit-MONSTER/1bit/zaya-cca-qk-norma2a5e15bMerge pull request Bump ZINC: zolotukhin/zinc b9123c649f81 #39 from 1bit-MONSTER/1bit/zaya-cca-conv-decodebc3aa1bahrx: ZAYA CCA query/key mixing and norms at decode in one dispatchb8bc863fhrx: ZAYA CCA convolution at decode in one dispatch1174e7b4Merge pull request Bump XDNA: amd/xdna-driver efddccd30568 #38 from 1bit-MONSTER/1bit/mmid-decode-no-runtime-div1162dda8Merge pull request Lean option: ROCmFP4 (Vulkan) and ROCmI4 (ROCm, W4A4) behind 1bit serve --lean #37 from 1bit-MONSTER/1bit/res-scale-loom-cases3d75c089Merge pull request site: the 1bit.MONSTER template, a cleaned-up blog, three new posts #36 from 1bit-MONSTER/1bit/zaya-decode-noop-copies8a9c7268hrx: MoE decode GEMV without runtime divisors0e5fba50hrx: res_scale_pair without a runtime divisor, with Loom check cases45c17708Merge pull request site: 1bit.monster, with a 404 page for old 1bit.MONSTER links #35 from 1bit-MONSTER/1bit/hrx-res-scale-fuseacd1b433zaya: skip CONTs of fresh contiguous results18be3726hrx: fuse ZAYA residual scale into one dispatch per sideb408831dMerge pull request NPU: read all three Q4NX chunk kinds (q4_1, Q4_K, Q8), repack Q4_K into q4_1 #33 from 1bit-MONSTER/1bit/hrx-jit-disk-cache250178cdhrx: cache Loom JIT results on disk across processes8de144eeMerge pull request Docs site: black and white theme #32 from 1bit-MONSTER/1bit/hrx-moe-decode-gemv5d5c2bc3hrx: one-token MUL_MAT_ID as a GEMV over the selected experts (MoE decode)22ceda3aMerge pull request Documentation site on GitHub Pages (1bit-monster.github.io/engine) #31 from 1bit-MONSTER/1bit/hrx-laya597cb57dggml-hrx: fused rotate-half RoPE and GEGLU for encoder graphs3be7161eggml-hrx: attention scores once per (query, head); Q8_0 rows kernel43855c71ggml-hrx: kernels to run Laya (ModernBERT encoder) on HRX0fcd83eb2Merge pull request 1bit serve --prefill-device hrx: prompt prefill on HRX, decode on Vulkan, one shared KV cache (zero copy) #30 from 1bit-MONSTER/1bit/hrx-124-output-passa34a6b75hrx: vectorise the decode-split multipass output pass (engine#124)00adc2b9Merge pull request HRX: pin our patched branch (IQ3_XXS, honest op claims), rebased on every bump #29 from 1bit-MONSTER/1bit/hrx-decode-split-barriers-both-spacesa14199a3hrx: fence both global and LDS at the decode-split barriers (engine#123 follow-up)9ca96038Merge pull request Vulkan from upstream llama.cpp: its own pin, kept on the latest release #28 from 1bit-MONSTER/1bit/hrx-decode-split-pack-barrier1c15f38ahrx: order the decode-split q8 pack after the reduce's global output stores (engine#123, HRX: decode_split_next_q8 gives wrong, nondeterministic attention at short context (up to 3.66 nats between identical requests) #140)fa226f93Merge pull request sync-lemonade-fork self-test: enable then disable a workflow #27 from 1bit-MONSTER/1bit/zaya-special-tokens2093fa90convert: ZAYA special tokens are CONTROL, so chat templates tokenizefdd8f1c6Merge pull request sync-lemonade-fork: selftest input to prove the token's write access #26 from 1bit-MONSTER/1bit/hrx-moe-router-nan-loudc16424f5hrx: make the all-NaN router-logit case loud instead of a silent wrong expert895d63f0Merge pull request Sync our Lemonade fork with upstream daily #25 from 1bit-MONSTER/1bit/hrx-moe-router-expert-id53c07152hrx: never publish the MoE router no-winner sentinel as an expert id358cafc2Merge pull request Docs: the onebit recipe lives in our Lemonade fork #24 from 1bit-MONSTER/1bit/hrx-zaya-kernelsf56203e1zaya: decode graph that runs entirely on HRX1f09354cggml-hrx: grouped F16 matmul and short-row kernels (softmax, sum_rows, argsort, narrow get_rows, strided copy, broadcast repeat)81c2740fMerge pull request Embed the engine into Lemonade, not Lemonade into the engine: drop the vendored Lemonade #22 from 1bit-MONSTER/1bit/zaya-conv-seqfold2eaff436zaya: fold sequences into tokens in the grouped conv, so Vulkan keeps it on the GPU8dd75eb4Merge pull request Bump HRX: llama.cpp f1a0aca141de, hrx-system 51b1739ae5fd #13 from 1bit-MONSTER/feat/hrx-decode-split-multipassf30cc439Merge pull request Laya: pin the router's source and checkpoints, hash-verified, kept current #18 from 1bit-MONSTER/1bit/zaya1-vlf191edd6zaya: ZAYA1-VL (vision LoRA on image tokens, bidirectional image attention, converter)914ed044ggml-hrx: eliminate global scale write-back in multipass decode-split reducer (engine#115)83e1c41aMerge pull request Step 3a: NPU full ELFs generated in C++ — no xclbin, no per-context files #8 from 1bit-MONSTER/census/register-sibling-hf-namesa61aa513convert: drop unused imports in the new converters01bea650opt, gptneo, codegen, gptj: fixes from checking each against transformers05cdb4a6Merge remote-tracking branch 'origin/1bit/hrx-vulkan-patched' into sib43483ce7convert: drop the BERT decoder aliases; Gemma4UnifiedForCausalLM onto the unified classb6c35357loom: drop leftover force-coop experiment lines - the multipass wrapper must apply reduce_completed.multipass (the cooperative reducer is invalid above 32 blocks and produced garbage)706ad1a7Merge pull request Lemonade zinc backend: GGUF through the engine's own ZINC build #15 from 1bit-MONSTER/1bit/zaya-multiseq7c797951zaya: correct CCA conv input with several sequences in a ubatch567d2dc7Merge pull request ZINC: pin upstream, serve it through Lemonade, NVIDIA verified on an RTX 5090 #14 from 1bit-MONSTER/1bit/zaya-legacyedcc1c7fconvert: fetch chat_template.jinja with --remoteaea4eea7convert: ZAYA legacy (Megatron-style) checkpointsdabb95aebenchmarks: fault decisively localized to reduce_completed.multipass; uniform-loop fixes A and B both fail (engine#115)1e4317a5benchmarks: DECISIVE - the fault is in reduce_completed.multipass (forced-cooperative is 0/5 at d2100 and d3000); divergent %lane loop bound is the suspect (engine#115)4380dafaMerge pull request HRX: follow AMD's live ggml-hrx, pinned to AMD's tested pair and kept current #11 from 1bit-MONSTER/1bit/zaya-swaee330cb0revert: drop the wave-padding experiment (2x compute, did not fix the fault)8493527adocs+tests: wave-padding falsified; residual fault depends on transient size, token count and arena layout; aliasing is the best fit (engine#115)b9214cdfbenchmarks: auditor rework notes - measured speed, d3000 still faults, OOB search narrowed (engine#115)f59f6d37HRX: multi-pass KV-block reduction for flash-attention decode-split (engine#115)d8c84b2fzaya: read sliding_window_pattern as a period, an array, or not at all1e775cdfMerge pull request XDNA: pin upstream amd/xdna-driver (and its XRT), build it privately, keep it current #12 from 1bit-MONSTER/fix/moe-router-experts-per-wave-launch-geometry6f469985zaya: sliding-window attention and a second rope base (ZAYA1-74B)290186aaggml-hrx: fix Qwen MoE fused router gfx1151 launch geometry (engine#108)460e5b7ecambrian: map CambrianQwenForCausalLM onto qwen284ae855bpico: map PicoDecoderHF onto the llama architectureaf55e780codegen: implement CodeGenForCausalLM (runtime graph + converter)bdd7f043gptneo: implement GPTNeoForCausalLM (runtime graph + converter)fa1a4563Merge pull request Steps 3b-3c: NPU fast lane runtime on full ELFs, served through Lemonade #9 from 1bit-MONSTER/fix/hrx-decode-split-fa-capacity-bound-clean53e85479opt: implement OPTForCausalLM (runtime graph + converter)5f708712gptj: implement the architecture (runtime graph + converter)6690b55cggml-hrx: cap the decode-split flash-attention dispatch at the kernel's real capacity895b526bconvert: register sibling HF names for architectures the runtime already has0f2ca133Merge pull request Step 2: HRX + Vulkan in one llama.cpp build, wired into Lemonade #7 from 1bit-MONSTER/fix/moe-router-route-ids-binding-length83ea7be3ggml-hrx: fix the Qwen MoE router's route-ids binding length when experts run on another device5556bf2cMerge pull request Steps 1-2: embedded Lemonade, and HRX + Vulkan in one llama.cpp build #6 from 1bit-MONSTER/1bit/hrx-setrows-95b1b87a99hrx: add scalar-scatter SET_ROWS for the non-FA KV-cache write (HRX: chat on Qwen3-Coder-30B-A3B and GLM-4.7-Flash fails with 500 "Compute error." #95)e2e6ca29Merge pull request Remove the CPU reference and tokenizer; port the existing engine instead #5 from 1bit-MONSTER/1bit/zaya-gpubd063e6fzaya: run on HRX and ROCm as well as Vulkanec07f7ebMerge pull request Tokenizer: byte-level BPE from GGUF, exact against HF tokenizers #2 from 1bit-MONSTER/1bit/zaya96f6b89bMerge pull request Server: 1bit-server speaks Lemonade's backend protocol #4 from 1bit-MONSTER/1bit/hrx-95d925875eggml-hrx: flash attention value_stride covers a whole KV row (HRX: chat on Qwen3-Coder-30B-A3B and GLM-4.7-Flash fails with 500 "Compute error." #95)19465d25ggml-hrx: GGML_HRX_DISABLE_DISPATCH and GGML_HRX_LOG_DISPATCH debug switchesee7ed7dfggml-vulkan: dma-buf import on Linux only2dc65867kv-share: build on Windowsff07f1c7hrx: fix MoE routing tables for strided route-id views (HRX: chat on Qwen3-Coder-30B-A3B and GLM-4.7-Flash fails with 500 "Compute error." #95)80ade7d8model: ZAYA1 (Zyphra), with a converter for the transformers checkpoint99d9e9e2kv-share: braced extern "C" so GCC builds it07a2ee4ehrx: claim flash attention for MLA and full-context MoE shapes (HRX: chat on Qwen3-Coder-30B-A3B and GLM-4.7-Flash fails with 500 "Compute error." #95)d7ff3afbhrx: fix non-flash-attention MoE graph compile failures (HRX: chat on Qwen3-Coder-30B-A3B and GLM-4.7-Flash fails with 500 "Compute error." #95)c075cc1cMerge pull request CPU reference: GGUF reader, dequant, Qwen3 fp32 forward, golden test #1 from 1bit-MONSTER/1bit/hrx-35b-prefill9a7aad66ggml-hrx: claim nodes that only a fused dispatch runs, inside model graphs96049a20ggml-hrx: MoE router partition table in the common layout above 128 experts79788e90ggml-hrx: decode flash attention past 256 KV tokens reduces only its own query rows267d8644Shared KV: 1 GiB chunks, so Vulkan never maps one buffer past its limit8fc19f0cZero-copy KV split: prompt prefill on HRX, decode on Vulkan, one shared KV cached2a9239fggml-hrx: IQ3_XXS dispatch in its own files, under our copyright notice61a53087ggml-hrx: claim only nodes the dispatcher can executed7b4d78fggml-hrx: IQ3_XXS matvec kernel for MUL_MATCI here builds without HRX. Before merging, on Strix Halo (and
test-backend-ops -b HRX0):cmake -B build -G Ninja -DONEBIT_HRX=ON && cmake --build build --target onebit && for d in hrx cpu; do tests/serve_e2e.sh build/1bit <Qwen3-0.6B Q4_K_M .gguf> $d; done