Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
381 commits
Select commit Hold shift + click to select a range
a61aa51
convert: drop unused imports in the new converters
Sep 26, 2026
83e1c41
Merge pull request #8 from 1bit-MONSTER/census/register-sibling-hf-names
bong-water-water-bong Sep 26, 2026
914ed04
ggml-hrx: eliminate global scale write-back in multipass decode-split…
Sep 26, 2026
f191edd
zaya: ZAYA1-VL (vision LoRA on image tokens, bidirectional image atte…
Sep 26, 2026
f30cc43
Merge pull request #18 from 1bit-MONSTER/1bit/zaya1-vl
bong-water-water-bong Sep 26, 2026
8dd75eb
Merge pull request #13 from 1bit-MONSTER/feat/hrx-decode-split-multipass
bong-water-water-bong Sep 26, 2026
2eaff43
zaya: fold sequences into tokens in the grouped conv, so Vulkan keeps…
Sep 26, 2026
81c2740
Merge pull request #22 from 1bit-MONSTER/1bit/zaya-conv-seqfold
bong-water-water-bong Sep 26, 2026
1f09354
ggml-hrx: grouped F16 matmul and short-row kernels (softmax, sum_rows…
Sep 26, 2026
f56203e
zaya: decode graph that runs entirely on HRX
Sep 26, 2026
358cafc
Merge pull request #24 from 1bit-MONSTER/1bit/hrx-zaya-kernels
bong-water-water-bong Sep 26, 2026
53c0715
hrx: never publish the MoE router no-winner sentinel as an expert id
bong-water-water-bong Sep 27, 2026
895d63f
Merge pull request #25 from 1bit-MONSTER/1bit/hrx-moe-router-expert-id
bong-water-water-bong Sep 27, 2026
c16424f
hrx: make the all-NaN router-logit case loud instead of a silent wron…
Sep 27, 2026
fdd8f1c
Merge pull request #26 from 1bit-MONSTER/1bit/hrx-moe-router-nan-loud
bong-water-water-bong Sep 27, 2026
2093fa9
convert: ZAYA special tokens are CONTROL, so chat templates tokenize
Sep 27, 2026
fa226f9
Merge pull request #27 from 1bit-MONSTER/1bit/zaya-special-tokens
bong-water-water-bong Sep 27, 2026
1c15f38
hrx: order the decode-split q8 pack after the reduce's global output …
Sep 27, 2026
9ca9603
Merge pull request #28 from 1bit-MONSTER/1bit/hrx-decode-split-pack-b…
bong-water-water-bong Sep 27, 2026
a14199a
hrx: fence both global and LDS at the decode-split barriers (engine#1…
Sep 27, 2026
00adc2b
Merge pull request #29 from 1bit-MONSTER/1bit/hrx-decode-split-barrie…
bong-water-water-bong Sep 27, 2026
a34a6b7
hrx: vectorise the decode-split multipass output pass (engine#124)
Sep 27, 2026
fcd83eb
Merge pull request #30 from 1bit-MONSTER/1bit/hrx-124-output-pass
bong-water-water-bong Sep 27, 2026
4e14145
Merge pull request #119 from AMD-Ecosystem/fix/hrx-iq4-table-indices
rsuderman Sep 28, 2026
ff318b8
Merge branch 'hrx-graph-develop-v2' of https://github.com/AMD-Ecosyst…
rsuderman Sep 28, 2026
43855c7
ggml-hrx: kernels to run Laya (ModernBERT encoder) on HRX0
Sep 28, 2026
3be7161
ggml-hrx: attention scores once per (query, head); Q8_0 rows kernel
Sep 28, 2026
597cb57
ggml-hrx: fused rotate-half RoPE and GEGLU for encoder graphs
Sep 28, 2026
22ceda3
Merge pull request #31 from 1bit-MONSTER/1bit/hrx-laya
bong-water-water-bong Sep 28, 2026
5d5c2bc
hrx: one-token MUL_MAT_ID as a GEMV over the selected experts (MoE de…
Sep 29, 2026
8de144e
Merge pull request #32 from 1bit-MONSTER/1bit/hrx-moe-decode-gemv
bong-water-water-bong Sep 29, 2026
59e3098
ggml-hrx: organize gated DeltaNet kernels into motifs
rsuderman Sep 29, 2026
5dd2747
Revert "ggml-hrx: preserve IQ4 table indices after Loom lookup fix"
rsuderman Sep 29, 2026
887810b
ggml-hrx: order vector RoPE dependencies after postprocessing
AaronStGeorge Sep 28, 2026
37b7d50
ggml-hrx: organize elementwise kernels into motifs
rsuderman Sep 29, 2026
250178c
hrx: cache Loom JIT results on disk across processes
Sep 29, 2026
b408831
Merge pull request #33 from 1bit-MONSTER/1bit/hrx-jit-disk-cache
bong-water-water-bong Sep 29, 2026
18be372
hrx: fuse ZAYA residual scale into one dispatch per side
Sep 29, 2026
acd1b43
zaya: skip CONTs of fresh contiguous results
Sep 29, 2026
45c1770
Merge pull request #35 from 1bit-MONSTER/1bit/hrx-res-scale-fuse
bong-water-water-bong Sep 29, 2026
0e5fba5
hrx: res_scale_pair without a runtime divisor, with Loom check cases
Sep 29, 2026
8a9c726
hrx: MoE decode GEMV without runtime divisors
Sep 29, 2026
3d75c08
Merge pull request #36 from 1bit-MONSTER/1bit/zaya-decode-noop-copies
bong-water-water-bong Sep 29, 2026
1162dda
Merge pull request #37 from 1bit-MONSTER/1bit/res-scale-loom-cases
bong-water-water-bong Sep 29, 2026
30f1cd3
Merge pull request #121 from rsuderman/main
rsuderman Sep 29, 2026
1174e7b
Merge pull request #38 from 1bit-MONSTER/1bit/mmid-decode-no-runtime-div
bong-water-water-bong Sep 29, 2026
b8bc863
hrx: ZAYA CCA convolution at decode in one dispatch
Sep 29, 2026
bc3aa1b
hrx: ZAYA CCA query/key mixing and norms at decode in one dispatch
Sep 29, 2026
a2a5e15
Merge pull request #39 from 1bit-MONSTER/1bit/zaya-cca-conv-decode
bong-water-water-bong Sep 29, 2026
c9283cf
Merge pull request #40 from 1bit-MONSTER/1bit/zaya-cca-qk-norm
bong-water-water-bong Sep 29, 2026
690e9e8
ggml-hrx: IQ4_NL / IQ4_XS weights decode correctly (no XOR 12 table r…
Sep 29, 2026
e892927
ggml-hrx: restore IQ4 table indices after Loom lookup fix
rsuderman Sep 29, 2026
d97e291
ggml-hrx: organize copy, layout, and row kernels into motifs
rsuderman Sep 29, 2026
3b053a4
hrx: fix the other IQ4_XS remap and the q8-plane SwiGLU placeholder
Sep 29, 2026
80c6b0b
ggml-hrx: organize normalization and RoPE kernels into motifs
rsuderman Sep 29, 2026
e096d96
ggml-hrx: organize SSM convolution kernels into motifs
rsuderman Sep 29, 2026
2ed008e
ggml-hrx: organize routed matmul kernels into motifs
rsuderman Sep 29, 2026
8c71384
Merge pull request #41 from 1bit-MONSTER/1bit/hrx-iq4-on-hrx
bong-water-water-bong Sep 29, 2026
b627e7b
ggml-hrx: organize Flash Attention kernels into motifs
rsuderman Sep 29, 2026
31629fc
ggml-hrx: restore logical IQ4 indices after lookup correction
AaronStGeorge Sep 29, 2026
fd20a3d
hrx: K-quant / IQ4 / Q8_0 decode projections in GGUF block layout
Sep 29, 2026
ade79d4
Merge pull request #42 from 1bit-MONSTER/1bit/hrx-kquant-decode
bong-water-water-bong Sep 29, 2026
2ba2ddc
Merge pull request #122 from AaronStGeorge/fix-ppl-36606527331-3d6968…
rsuderman Sep 29, 2026
316dc22
hrx: Q3_K and IQ3_S in the K-quant decode kernels
Sep 29, 2026
432ce55
Merge amdeco/hrx-graph-develop-v2
rsuderman Sep 29, 2026
f81814a
Merge pull request #43 from 1bit-MONSTER/1bit/hrx-kquant-q3k
bong-water-water-bong Sep 29, 2026
5c6d476
hrx: sleep through long stream waits instead of spinning a CPU core
Sep 29, 2026
fcc0785
ggml-hrx: reduce warmup staging and scale JIT
rsuderman Sep 29, 2026
6617e78
Merge pull request #44 from 1bit-MONSTER/1bit/hrx-sleeping-wait
bong-water-water-bong Sep 29, 2026
b8061ff
hrx: K-quant SwiGLU decode above the int4 lowrow path
Sep 29, 2026
9e485eb
Merge pull request #45 from 1bit-MONSTER/1bit/hrx-kquant-swiglu-prio
bong-water-water-bong Sep 29, 2026
b57bd75
hrx: K-quant decode kernels for 2-8 tokens (MTP / speculative verify)
Sep 29, 2026
dad19ad
Merge pull request #46 from 1bit-MONSTER/1bit/hrx-kquant-tokens
bong-water-water-bong Sep 30, 2026
ef9e3c2
hrx: K-quant token kernels accumulate vectors, reduce once per row
Sep 30, 2026
0d7e398
Merge pull request #47 from 1bit-MONSTER/1bit/hrx-kquant-tokens-vacc
bong-water-water-bong Sep 30, 2026
3b954f0
hrx: kquant matchers skip q8-only inputs; softplus kernel
Sep 30, 2026
8bf4fcf
llama: load Hadamard-folded weights (prism.hadamard, Ternary Bonsai)
Sep 29, 2026
8981d2b
llama-hadamard: read the prism.hadamard keys through gguf
Sep 29, 2026
616348a
llama-hadamard: credit PrismML and carry the MIT notice of the code f…
Sep 29, 2026
acf9c74
Merge pull request #48 from 1bit-MONSTER/1bit/hrx-kquant-ready
bong-water-water-bong Sep 30, 2026
804a0f7
Merge branch '1bit/hrx-vulkan-patched' into 1bit/hrx-bonsai
bong-water-water-bong Sep 30, 2026
b8d587e
Merge pull request #34 from 1bit-MONSTER/1bit/hrx-bonsai
bong-water-water-bong Sep 30, 2026
738460e
hrx: Q2_K weights on HRX0 (shared dequantizer, K-quant decode kernels)
Sep 30, 2026
ef779b8
hrx: kquant header: format 12 is a common weight format now (review)
Sep 30, 2026
d1747cb
Merge pull request #49 from 1bit-MONSTER/1bit/hrx-q2k
bong-water-water-bong Sep 30, 2026
6cbab7e
hrx: IQ2_XXS and IQ2_XS weights on HRX0 (shared dequantizer, K-quant …
Sep 30, 2026
0174b32
hrx: kquant decode: block-bytes comment back above its function (review)
Sep 30, 2026
a05c1b6
hrx: IQ2 generators: flake8 and pyright clean (output unchanged)
Sep 30, 2026
042b391
hrx: publish prefill alternate layouts
rsuderman Sep 30, 2026
d0b1544
ggml-hrx: consolidate vector routed matmul kernels
rsuderman Sep 30, 2026
d5048ad
Merge pull request #50 from 1bit-MONSTER/1bit/hrx-iq2
bong-water-water-bong Sep 30, 2026
5922224
ggml-hrx: retire redundant Qwen routed vector kernels
rsuderman Sep 30, 2026
3ec1477
hrx: retire unused Qwen Q4_K token embedding
rsuderman Sep 30, 2026
3cdd69c
ggml-hrx: enforce a minimum HRX source revision
AaronStGeorge Sep 30, 2026
142cbe8
ggml-hrx: share the revision check through Python
AaronStGeorge Sep 30, 2026
f0be33c
ggml-hrx: simplify revision checking
AaronStGeorge Sep 30, 2026
f5ae1a5
ggml-hrx: retire obsolete Qwen prefill flash attention
rsuderman Sep 30, 2026
3ba8786
ggml-hrx: retire duplicate Qwen expert table builders
rsuderman Sep 30, 2026
2010368
hrx: retire Qwen split-decode FlashAttention packet
rsuderman Sep 30, 2026
6fbb135
hrx: retire unreachable Qwen same-format QKV experiment
rsuderman Sep 30, 2026
d1674fd
hrx: retire dead Qwen dense controls
rsuderman Sep 30, 2026
58e5a7e
hrx: retire redundant Qwen kernel paths
rsuderman Sep 30, 2026
c2f9610
ggml-hrx: deduplicate dequant grid constants
rsuderman Sep 30, 2026
14330fb
Merge pull request #127 from rsuderman/main
rsuderman Oct 1, 2026
12f086c
ggml-hrx: fuse Lemonade model operations
rsuderman Oct 1, 2026
4709e9c
ggml-hrx: fix MoE routing JIT catalog coverage
rsuderman Oct 1, 2026
0c822e7
Merge remote-tracking branch 'amdeco/hrx-graph-develop-v2'
rsuderman Oct 1, 2026
c7baa2b
ggml-hrx: consolidate Qwen MoE kernels
rsuderman Oct 1, 2026
5825fde
hrx: IQ3_XXS and IQ2_S weights on HRX0 (shared dequantizer, K-quant d…
Oct 1, 2026
764a256
Merge pull request #51 from 1bit-MONSTER/1bit/hrx-iq3xxs-kq
bong-water-water-bong Oct 1, 2026
7912b2d
hrx: GDN snapshot publish forms decay ratios in log space (fixes NaN …
Oct 1, 2026
cde002d
Merge pull request #52 from 1bit-MONSTER/1bit/hrx-gdn-snapshot-ratio
bong-water-water-bong Oct 1, 2026
7de8412
hrx: IQ1_S and IQ1_M weights on HRX0 (shared dequantizer, K-quant dec…
Oct 1, 2026
cec43f6
Merge pull request #126 from AMD-Ecosystem/users/astgeorg/hrx-pin
rsuderman Oct 1, 2026
30d0701
hrx: packed ternary decode for exact-ternary Q4_0 weights (opt-in, GG…
Oct 1, 2026
a653d0c
ggml-hrx: add Lemonade publication fusions
rsuderman Oct 1, 2026
f602a42
ggml-hrx: add Lemonade compound fusions
rsuderman Oct 1, 2026
bf5ad7c
Merge pull request #53 from 1bit-MONSTER/1bit/hrx-iq1-v2
bong-water-water-bong Oct 1, 2026
4fbae7a
hrx: packed ternary upload accepts all-zero Q4_0 blocks with any scal…
Oct 1, 2026
dd74f6b
Merge pull request #54 from 1bit-MONSTER/1bit/hrx-ternary
bong-water-water-bong Oct 1, 2026
1b61a7d
hrx: consolidate configurable binary fusion kernels
rsuderman Oct 1, 2026
58bc9fa
Merge pull request #128 from rsuderman/main
rsuderman Oct 1, 2026
a72748b
test: update HRX RMSNorm JIT corpus coverage
rsuderman Oct 1, 2026
a5a5a62
hrx: unify alternate activation publication
rsuderman Oct 1, 2026
6e42b51
ggml-hrx: route Q4_K/Q5_K/IQ4_XS prompt matmuls to the q8_1 x4 kernel…
bong-water-water-bong Oct 1, 2026
269eef7
hrx: complete consumer-driven activation publication
rsuderman Oct 1, 2026
04b6b12
ggml-hrx: let Loom select AMDGPU runtime globals
AaronStGeorge Oct 1, 2026
b00ddff
hrx: K-quant SwiGLU decode keeps separate gate and up codebooks (mixe…
Oct 1, 2026
69e721c
hrx: consolidate RMSNorm gate publication kernels
rsuderman Oct 1, 2026
8d880f6
llama: load the engine's Hadamard-rotated Q4_0 files (onebit.hadamard…
bong-water-water-bong Oct 2, 2026
2029ecb
ggml-hrx: keep V row-major past copy_transpose_f16's 32768-row range …
bong-water-water-bong Oct 2, 2026
68805cc
ggml-hrx: q8_1 x4 prefill takes the 256-aligned head of a remainder c…
bong-water-water-bong Oct 2, 2026
326cb7c
hrx: K-quant SwiGLU fills one codebook when gate and up share it (rev…
Oct 2, 2026
bd5b297
Merge pull request #60 from 1bit-MONSTER/1bit/hrx-mixgrid
bong-water-water-bong Oct 2, 2026
0be30b6
ggml-hrx: add reusable IQ codebook matmuls
rsuderman Oct 2, 2026
bda40b5
hrx: consolidate RMSNorm binary publication roots
rsuderman Oct 2, 2026
c37033e
hrx: reuse kernels across dynamic workloads
rsuderman Oct 2, 2026
d0f3cf7
ggml: PrismML PQ2_0 / PTQ1_0 types (ids 142 / 143) with a CPU reference
Oct 2, 2026
5f2424a
ggml-hrx: PQ2_0 / PTQ1_0 native on HRX0 (one resident copy), Q1_0 on …
Oct 2, 2026
ce70987
ggml-hrx: kquant check cases for PTQ1_0 paired with a codebook format…
Oct 2, 2026
2075152
ggml-cpu: clamp lists PQ2_0 / PTQ1_0 (-Werror=switch on CI)
Oct 2, 2026
93f1381
hrx: bound the graph program cache (GGML_HRX_GRAPH_PROGRAM_CACHE, def…
Oct 2, 2026
6b1061a
Merge pull request #62 from 1bit-MONSTER/1bit/hrx-prism-types
bong-water-water-bong Oct 2, 2026
d60cc4f
Merge pull request #63 from 1bit-MONSTER/1bit/hrx-program-cache-cap
bong-water-water-bong Oct 2, 2026
827b486
hrx: Walsh-Hadamard kernel for MUL_MAT hinted GGML_HINT_SRC0_IS_HADAMARD
Oct 2, 2026
e2b946a
Merge pull request #64 from 1bit-MONSTER/1bit/hrx-hadamard-fwht
bong-water-water-bong Oct 2, 2026
89ddd94
ggml-hrx: MXFP4 weights in the shared dequantizer
Oct 2, 2026
6892003
ggml-hrx: exact E8M0 half scale for MXFP4 in the shared dequantizer
Oct 2, 2026
b0d3d46
tests: test-hrx-mxfp4, a bit-exact known-answer check of MXFP4 GET_RO…
Oct 2, 2026
52b79b9
tests: test-hrx-mxfp4 accepts zeros only for rows whose E8M0 scale is…
Oct 2, 2026
adfd541
ggml-hrx: MUL_MAT_ID admits only input sizes that are a multiple of 256
Oct 2, 2026
cebcd70
Merge pull request #65 from 1bit-MONSTER/1bit/hrx-mxfp4-v2
bong-water-water-bong Oct 2, 2026
7370b6e
ggml-hrx: MUL_MAT_ID WMMA kernels take input sizes that are a multipl…
Oct 2, 2026
7215efd
ggml-hrx: decode kernels read input row t at t * input_size
Oct 2, 2026
a8792c3
tests: test-hrx-decode-stride, decode-kernel input rows at input size…
Oct 2, 2026
40740e9
tests: test-hrx-mul-mat-id-k32, MUL_MAT_ID at input sizes that are a …
Oct 2, 2026
50a63b6
ggml-hrx: ADD_ID on F32 (gpt-oss expert biases)
Oct 2, 2026
d17f05f
ggml-hrx: SWIGLU_OAI on F32 (gpt-oss clamped SwiGLU)
Oct 2, 2026
1bd9ee6
Merge pull request #66 from 1bit-MONSTER/1bit/hrx-mmid-g32-v2
bong-water-water-bong Oct 2, 2026
8bd4f89
ggml-hrx: TQ1_0 / TQ2_0 weights, and single-token decode lanes for TQ…
Oct 2, 2026
b9c3939
ggml-hrx: TQ1_0 decoders never subtract past zero
Oct 2, 2026
b221e1f
ggml-hrx: scalar reference and check cases for TQ1_0, TQ2_0 and MXFP4…
Oct 2, 2026
74b8dea
Merge pull request #67 from 1bit-MONSTER/1bit/hrx-gptoss-ops
bong-water-water-bong Oct 2, 2026
f2c6298
ggml-hrx: MXFP4 check cases fill bytes 121..127 (the iota must stay a…
Oct 2, 2026
0b56f47
Merge pull request #69 from 1bit-MONSTER/1bit/hrx-tq-pr
bong-water-water-bong Oct 2, 2026
56b4453
ggml-hrx: no V transpose for MLA's strided V (GLM-4.7-Flash prompts >…
Oct 2, 2026
53c8b17
llama-hadamard: fold qwen3next's ssm_ba; refuse stamped files with un…
Oct 2, 2026
6339066
Merge pull request #70 from 1bit-MONSTER/1bit/hrx-fa-mla-transpose
bong-water-water-bong Oct 2, 2026
388290b
Merge pull request #71 from 1bit-MONSTER/1bit/hadamard-onebit-guard
bong-water-water-bong Oct 2, 2026
e5afc24
ggml-hrx: reuse SwiGLU block metadata on integration base
AaronStGeorge Sep 30, 2026
3cbd76a
Merge pull request #130 from AaronStGeorge/fix-bump-pr-87-1
rsuderman Oct 2, 2026
80b9259
Merge pull request #125 from AMD-Ecosystem/experiment/hrx-metadata-st…
rsuderman Oct 2, 2026
ba0be9a
ggml-hrx: keep ADD_ID / SWIGLU_OAI with their MUL_MAT_ID (placement g…
Oct 2, 2026
35027f7
Merge pull request #73 from 1bit-MONSTER/1bit/hrx-moe-split-guard
bong-water-water-bong Oct 2, 2026
e618709
ggml-hrx: masked KV rows reach flash-attention P*V as V = +0 (identic…
Oct 2, 2026
2bd7f58
Merge pull request #72 from 1bit-MONSTER/1bit/hrx-fa-masked-v
bong-water-water-bong Oct 2, 2026
dcbc1d4
ggml-hrx: attention sinks for FLASH_ATTN_EXT (gpt-oss), applied after…
Oct 2, 2026
5405484
ggml-hrx: attention-sink dispatch refuses outputs that overlap an inp…
Oct 2, 2026
4485916
Merge pull request #68 from 1bit-MONSTER/1bit/hrx-attention-sinks
bong-water-water-bong Oct 2, 2026
d3dc0f7
ggml-hrx: PrismML tile bytes for the low-token SwiGLU (PQ2_0 1.7B dec…
Oct 2, 2026
f5b7f4a
Merge pull request #74 from 1bit-MONSTER/1bit/hrx-prism-tile-bytes
bong-water-water-bong Oct 2, 2026
c4c14f4
hrx: add packed Q1_0 Q8_1 decode matmul
rsuderman Oct 2, 2026
f104f8d
hrx: optimize Q1_0 token-64 tiled matmul
rsuderman Oct 2, 2026
2ae98de
hrx: decouple skinny matmul token capacity schema
rsuderman Oct 2, 2026
973c934
ggml-hrx: extend reusable IQ codebook matmuls
rsuderman Oct 2, 2026
02122cf
ggml-hrx: add packed Q1_0 paired SwiGLU route
rsuderman Oct 2, 2026
f681f85
Merge amdeco/hrx-graph-develop-v2 into main
rsuderman Oct 2, 2026
437e1cb
ggml-hrx: speed up IQ4_NL dequantization
rsuderman Oct 2, 2026
b802a50
Merge pull request #132 from rsuderman/main
rsuderman Oct 2, 2026
85f5983
ggml-hrx: the f32_f32 WMMA cores accumulate in f32
Oct 3, 2026
34e545f
spec: MTP bridges the previous h-row into a fresh sequence (engine #290)
Oct 3, 2026
e0f637b
ggml-hrx: publish deferred host writebacks before a host-staging upload
Oct 3, 2026
b5669e0
ggml-hrx: keep the transposed result-tile staging with the f32 accumu…
Oct 3, 2026
1c0241b
hrx: decline the decode-split under a non-production partial-transien…
Oct 3, 2026
d336ba6
ggml-hrx: fold the MUL_MAT_ID WMMA f16 accumulator into f32 every 16-…
Oct 3, 2026
973e5ce
Merge pull request #84 from 1bit-MONSTER/1bit/hrx123-decline-guard
bong-water-water-bong Oct 3, 2026
2e5fcf5
Merge pull request #85 from 1bit-MONSTER/1bit/hrx284-f32acc
bong-water-water-bong Oct 3, 2026
616034b
ggml-hrx: run hipcc code objects through HRX (HIP kernel plumbing)
Oct 4, 2026
8719cf4
ggml-hrx: HIP kernel add-on hook (GGML_HRX_HIP_ADDON_DIR)
Oct 5, 2026
d5de4b1
tests: MUL_MAT_VEC_FUSION cases for one-token Q4_K/Q5_K gate/up + SwiGLU
Oct 5, 2026
a679f78
Merge pull request #82 from 1bit-MONSTER/1bit/hrx-host-writeback-flush
bong-water-water-bong Oct 5, 2026
86eee89
Merge pull request #83 from 1bit-MONSTER/1bit/mtp-pending-h-continuity
bong-water-water-bong Oct 5, 2026
d5a9932
Merge 1bit/hrx-vulkan-patched into 1bit/hrx-f32-wmma-accumulate
Oct 5, 2026
d1691f4
Merge pull request #81 from 1bit-MONSTER/1bit/hrx-f32-wmma-accumulate
bong-water-water-bong Oct 5, 2026
2806e9a
ggml-hrx: stop the sleeping wait from oversleeping decode after a prompt
Oct 5, 2026
f3c771f
Merge pull request #87 from 1bit-MONSTER/1bit/hrx-replay-keep
bong-water-water-bong Oct 5, 2026
4dade38
llama: add token ID tracking to KV cell (#27762)
ngxson Aug 26, 2026
35207fe
model : allow reshape of tensors during load (#26531)
ggerganov Aug 4, 2026
254e014
llama: model_loader: add TENSOR_READ_LAZY (#27794)
ngxson Aug 27, 2026
d5a3a1b
model: add Qwen3.8-Flash-Next (qwen4exp) (#27742)
danielhanchen Aug 27, 2026
08a1233
model: qwen4exp: reduce number of graph splits (#27880)
ngxson Aug 28, 2026
d3516bf
bench: add --tensor-read-lazy (#27881)
ngxson Aug 28, 2026
9175627
llama: improve TENSOR_READ_LAZY handling (#27837)
ngxson Aug 30, 2026
6378935
common: rename --tensor-read-lazy to --lazy-mode, add -lzm shorthand …
ggerganov Aug 30, 2026
6354fe9
qwen4exp: sum the indexer heads by slices (#28023)
ServeurpersoCom Sep 1, 2026
2c684b5
qwen4exp: support recurrent state rollback (#28123)
ServeurpersoCom Sep 1, 2026
ffca4b4
qwen4exp: fix seq_cp, block position keying, mtmd input, cuda abort, …
danielhanchen Sep 1, 2026
4bb1e26
kv-cells: look up the n-gram history in the sequence position index (…
ServeurpersoCom Sep 1, 2026
7d02867
models : fix GDN normalization from `max` to `rsqrt` (#28068)
danielhanchen Sep 6, 2026
34ff68b
qwen4exp: enable rms_norm + mul fusion (#28896)
am17an Sep 14, 2026
864deaf
qwen4exp: add hc ops (#28901)
am17an Sep 16, 2026
1458fd8
qwen4exp: NextN/MTP draft head for Qwen3.8-Flash-Next (ggml-org/llama…
danielhanchen Sep 26, 2026
f294a5b
qwen4exp MTP: reduced draft vocabulary (d2t + K-row head)
Oct 1, 2026
bfda5f5
qwen4exp MTP: refuse a draft vocabulary whose d2t ids are out of rang…
Oct 1, 2026
f89d503
qwen4exp: run Qwen3.8-Flash-Next on HRX (graph adjustments, dense QSA…
Oct 4, 2026
7740a1a
qwen4exp MTP: load and run the Flash-Next draft head in this tree
Oct 5, 2026
3527b56
ggml-hrx: HIP matchers can declare the ops they claim (hip/hip-capabi…
Oct 5, 2026
001832c
Merge pull request #86 from 1bit-MONSTER/1bit/hrx-hip-plumbing
bong-water-water-bong Oct 5, 2026
e57beb9
Merge pull request #88 from 1bit-MONSTER/1bit/qwen4exp-hrx
bong-water-water-bong Oct 5, 2026
d959413
ggml-hrx: transient arena reuse without false dependencies
Oct 5, 2026
efb93ce
ggml-hrx: record graph kernels with a bounded lookahead to share barr…
Oct 5, 2026
f95f2db
Merge pull request #89 from 1bit-MONSTER/1bit/hrx-runtime-overhead
bong-water-water-bong Oct 5, 2026
b4e7515
ggml-hrx: never write a device-rewritten graph input back to host memory
Oct 5, 2026
f862dae
ggml-hrx: wait for a non-replay command program that reads registered…
Oct 5, 2026
69788ac
build: C++26 by default (amdclang 23 / g++ 15), no module scanning
Oct 5, 2026
6150ff0
Merge pull request #91 from 1bit-MONSTER/1bit/cxx26
bong-water-water-bong Oct 5, 2026
522dab4
Merge pull request #90 from 1bit-MONSTER/1bit/kv-nan-rootcause
bong-water-water-bong Oct 5, 2026
1d4333c
build: C++26 only where the compiler offers it; older toolchains keep…
Oct 5, 2026
ebe1662
Merge pull request #92 from 1bit-MONSTER/1bit/cxx26-ci
bong-water-water-bong Oct 5, 2026
4d13516
common: record NaN logits instead of only aborting, with opt-in conti…
Oct 6, 2026
b05d329
Merge pull request #93 from 1bit-MONSTER/1bit/hrx-nan-continue
bong-water-water-bong Oct 6, 2026
59771f2
Merge AMD's pinned pair b802a507 into 1bit/hrx-vulkan-patched
Oct 6, 2026
b3981ce
Merge commit '59771f2a63c30c0edb9eaa79bef8b1b937788733' into HEAD
Oct 6, 2026
b04c4e9
Merge pull request #94 from 1bit-MONSTER/1bit/cxx26-ci
bong-water-water-bong Oct 6, 2026
51e538f
hrx: cap the decode-split dispatch at 2048, the last provider the mer…
Oct 6, 2026
420b4e6
ggml-hrx: take AMD's corpus wholesale and keep our kernels beside it
bong-water-water-bong Oct 6, 2026
39588ef
ggml-hrx: do not cache a compiled kernel whose launch program exists
bong-water-water-bong Oct 6, 2026
56c3c8a
ggml-hrx: route mul_mat_id through our kernel again
bong-water-water-bong Oct 6, 2026
f2099e9
hrx: keep our ggml-hrx backend on AMD's core
Oct 6, 2026
cf4dfd8
ggml-hrx: decode-split FA always-stage fallback (engine#314)
bong-water-water-bong Oct 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
10 changes: 10 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,16 @@ cmake_minimum_required(VERSION 3.14...3.28) # for add_link_options and implicit
project("llama.cpp" C CXX)
include(CheckIncludeFileCXX)

# C++26 where the compiler has it (amdclang 23, g++ 15); older toolchains keep C++17
if("cxx_std_26" IN_LIST CMAKE_CXX_COMPILE_FEATURES)
set(CMAKE_CXX_STANDARD 26)
else()
set(CMAKE_CXX_STANDARD 17)
endif()
set(CMAKE_CXX_STANDARD_REQUIRED true)
# no C++20 module scanning: no module sources, and some toolchains lack clang-scan-deps
set(CMAKE_CXX_SCAN_FOR_MODULES OFF)

#set(CMAKE_WARN_DEPRECATED YES)
set(CMAKE_WARN_UNUSED_CLI YES)

Expand Down
91 changes: 91 additions & 0 deletions benchmarks/NOTE-hrx-124-multipass-output-2026-09-27.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
# HRX decode-split multipass output pass: drop the 128x redundant expf and the scalar block loop (engine#124)

Worktree: `~/1bit-engine-176/third_party/llama.cpp`, branch `1bit/hrx-124-output-pass` (base `00adc2b`).
Box: gfx1151 (Strix Halo), `Qwen3-Coder-30B-A3B-Instruct-Q4_K_M`, `-dev HRX0`.

## What was wrong

`@ggml.flash_attention.decode_split.reduce_completed.multipass`
(`ggml/src/ggml-hrx/kernel-corpus/kernels/loom-libs/ops/flash_attention_decode_split_f32_f16_wmma.loom`)
finished with a scalar output pass: `workitem -> one output channel`, then a serial
`scf.for %block = [%c0 to %active_block_count]` over **every** KV block, run by all 256
workitems. Only 128 are live at `value_head_size = 128`, and each one recomputed
`scale = expf(partial_max[block] - maximum)` for every block, so `expf` ran once per
`(block, output element)` instead of once per `(block)`, and every element walked the
block dimension alone. That is the residual ~16% gap above capacity 2048 (issue #124).

## The change (bit-exact)

1. **LDS scale stage.** The lane-strided sum pass already computes
`scale = expf(partial_max[block] - maximum)` for every block; store it into a
per-row workgroup stage (`scale_stage_view[row, block]`) and also publish the row
`sum` (`sum_stage_view[row]`). The output pass then just reloads the exact f32
value. Removes ~128x redundant `expf` and the redundant `partial_max` reloads.
2. **Vectorised all-rows output pass.** After the per-row max/sum phase, one pass
covers all query rows at once: workitem `w` owns a 4-channel *quad*
(`quad = tile*256 + w`, `row = quad / quads_per_row`, `channel = (quad %
quads_per_row)*4`), loads `vector<4xf16>` from `partial_output`, and accumulates
`vector<4xf32>` over blocks with `unroll(%c4) schedule(interleaved)`. For the
production GQA shape (8 query heads per KV head, `value_head_size = 128`) there are
exactly 256 quads, so the whole workgroup (all four subgroups) is busy and each
channel still accumulates its blocks in the same order as before.

Because the per-channel accumulation order is unchanged and the vector ops are
elementwise, the reduce output is **bit-identical** to the previous multipass reducer
(verified end-to-end below).

## Result

`llama-bench -p 0 -n 8 -r 5 -d 1900,2000,2100,3000,4800`, 6 interleaved A/B rounds,
median of the 6 run means:

| depth | capacity | blocks | base t/s | patched t/s | delta | path |
|---|---:|---:|---:|---:|---:|---|
| 1900 | 1920 | 30 | 70.72 | 71.48 | +1.1% | cooperative (untouched control) |
| 2000 | 2048 | 32 | 66.06 | 68.82 | +4.2% | cooperative (untouched control) |
| 2100 | 2112 | 33 | 58.19 | 67.15 | **+15.4%** | multipass |
| 3000 | 3008 | 47 | 50.84 | 60.89 | **+19.8%** | multipass |
| 4800 | 4864 | 76 | 41.83 | 51.02 | **+22.0%** | multipass |

The boundary cliff is gone: `d2100/d2000` moves **0.881 -> 0.976** (and
`d4800/d2000` 0.633 -> 0.741). The two depths the reducer does not touch moved by
+1.1% / +4.2%, i.e. run-to-run noise.

A throwaway diagnostic that deleted the output block loop entirely (keeping the
max/sum passes) measured **+21.6% / +50.0% / +30.5%** at d2100 / d3000 / d4800, which
is how the output pass was identified as the whole gap rather than the reduction math
(issue #124's diagnosis).

## Evidence

- **Bit-exact.** `llama-perplexity -c 2049 -b 1 -f <6600-token text> --save-all-logits`
(token-by-token decode, so the decode-split kernel is dispatched; `GGML_HRX_LOG_DISPATCH=1`
confirms `flash_attention_decode_split`). The 1,244,725,284-byte logits dumps for base
and patched are **byte-identical** (`cmp`), PPL 21.9760 +/- 0.81813 on both.
- **Correctness.** Buried code word `ZX-4718-QQ` at 4700 prompt tokens (capacity 4864,
multipass), `llama-server`, temperature 0, seed 42, `cache_prompt:false` -> answer
`ZX-4718-QQ`. PASS.
- **Faults.** 6 A/B rounds x 5-rep sweeps at all five depths on the patched build:
0 `HSA_STATUS_ERROR_MEMORY_FAULT`, 0 `res = -3`. The cooperative/direct paths are
untouched, so <= 2048 (d1900/d2000) shows no tok/s or fault regression.

## Reproduce

```
# build (engine worktree ~/1bit-engine-176, HRX llama.cpp build dir)
cmake --build build/hrx/llama --target llama-bench llama-perplexity -j 16

# dispatch needs libhsa from the HRX toolchain
export IREE_HAL_AMDGPU_LIBHSA_PATH=/opt/rocm-therock/lib/python3.14/site-packages/_rocm_sdk_core/lib/libhsa-runtime64.so.1

# performance
build/hrx/llama/bin/llama-bench -m models/Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf \
-dev HRX0 -ngl 99 -p 0 -n 8 -r 5 -d 1900,2000,2100,3000,4800

# bit-exact reference vs the previous build
build/hrx/llama/bin/llama-perplexity -m models/Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf \
-f ref-text.txt -c 2049 -b 1 -dev HRX0 -ngl 99 --save-all-logits logits.bin
```

Raw runs and the A/B driver used here live under the session scratch dir
(`ab5-*.json`, `ab5-summary.txt`, `diag-noloop.json`).
Loading