Skip to content

perf: Qwen3.5 levers3; GLM-5.3 decode and MTP-verify levers (routed-expert ring, top-k at 2-4 rows, L2 issue stop, hyper-connection front) - #84

Merged
marcospaulo merged 67 commits into
mainfrom
train/engine-10
Oct 2, 2026
Merged

marcospaulo merged 67 commits into
mainfrom
train/engine-10

Conversation

@marcospaulo

@marcospaulo marcospaulo commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Commits on top of #83 (82a23c9). Every kernel change is behind a *_LEGACY switch (the graph-order and harness changes have none), and every change has its row in TORAD.md with its numbers, its tests and what stays open.

Qwen3.5 and glm5next KDA

331d727: Qwen3.5 levers between the qkv group and the recurrence, each behind its own switch.

  • The alpha/beta pair is folded into the conv-state update.
  • A PQ2_0 launch prefetches into L2 the weights of the kernels before the next launch.
  • The conv's and the recurrence's weights and states are requested before their PDL waits.
  • The recurrence and mul_mat_vec_f trigger the next launch at their start.
  • Register caps.

Served bits (Ternary Bonsai 2 27B, one slot and four, 128 greedy tokens with top-5 log-probabilities) are identical to all switches set. Open, in its row: the fold does not engage at this commit; the caps leave 16–32 B stacks; the Qwen speed sweep that sets the defaults is still to run.

e74ed87: glm5next's KDA output gate goes into the graph before the recurrence. g_a runs beside f_a as one mat-vec launch, g_b before the recurrence: one small mat-vec launch fewer a KDA layer, the chain −0.36 % (RTX 5070 Ti).

7536ed7, then b850d70: the KDA gated norm written by the recurrence's kernel, then taken out. Bit-identical, and neutral against e74ed87 over 4 rotated rounds (tg128 +0.24 %, CI −0.95 to +1.42); the norm's time is the wait for the state stores to drain, which stays on the chain wherever the norm runs. b850d70's tree is e74ed87's.

GLM-5.3 decode (44-layer proxy, RTX 5080 + 5070 Ti)

  • bbe5370: the routed down ring lists its experts and streams its first tiles before its PDL wait (its ids are whole: the gate/up before it read them). Down 17.76 → 16.19 us (5080), 18.94 → 17.25 (5070 Ti) under -sm tensor. New MOE_FFN_CHAIN test.
  • 3a68863: the routed experts' weighted sum in one launch, bit for bit the MUL and ADDs it replaces. New MOE_WEIGHTED_SUM test, compared bit for bit.
  • 90d1b98: the hyper-connection front writes its normed mix's q8_1 copy, two quantize launches fewer a layer. New DSV4_HC_PRE_Q8_1 test.
  • 4bce940: the front's comb (Sinkhorn) runs beside the stream, off the chain to the sublayer's projections; −1.03 % / −0.69 % kernel time a token. New DSV4_HC_PRE_POST test. Opt-in since 98a6656 (below).
  • 6588e1d: test-backend-ops runs the reference first, so a kernel a backend leaves running past its return shows as a failure.
  • c6b4ecf: the comb beside the stream only where the front's weights outlive the front, so the allocator cannot put another output over them.

GLM-5.3 MTP verify (3 tokens)

  • 5c03ed9: the fused MoE top-k at 2–4 rows, where the allocator put its weights and ids over its logits and the fusion was refused. Routing chain 8.32 → 3.39–3.90 us a MoE layer; kernel time −2.35 % / −1.78 % under -sm tensor. PPL at -ub 3 moves 385507.0946 → 385634.0270: the fused kernel's own normalization rounding, as a decode's always had (KLD vs -ub 1 0.006763 against 0.006756).
  • b3fc826: the paced L2 issue stops when a routed-expert ring starts. Before, it ran 27 us into the gate/up and set its end. One card: gate/up 125.15 → 108.26 us a KDA layer, pp3 wall +2.67 % (CI +0.34 to +5.01). Declared and missed: the stop within 3 us (it takes 9.3). With it, 5c03ed9 shows on one card too: wall +1.48 % (CI +0.50 to +2.47).
  • 3dbefa8: at several tokens the ring decodes each IQ3_XXS fragment once for its expert's pairs. It is a kernel instance of its own, since one kernel with both paths spilled.
    • One card: gate/up 107.81 → 75.65 us and down 61.41 → 49.66 us a KDA layer, kernel time −12.12 %.
    • -sm tensor: kernel time −7.01 % (5080) and −6.52 % (5070 Ti).
    • pp3 at -ub 3: +9.68 % (CI +8.58 to +10.79) on one card, +6.63 % (CI +4.33 to +8.93) under -sm tensor.
    • All 8,302 of b3fc826's kernels keep the same SASS encodings, so decode is unchanged by construction.
    • Declared and missed: the down at ≤ 45 us. Its tokens' vectors, 55 KB, don't fit beside the ring and are read from global memory.
    • The proxy routes every token to all 8 of its experts, the best case for this change. GLM-5.3-Flash has 288 experts; its gain depends on how many experts an MTP verify's tokens share. 2c5fb99 (below) keeps it only where they share.
  • 04dde5d: the ring takes each pair's vector and output offsets from host tables and its tile and slot by fastdiv. ncu showed the down at 48 % of DRAM with 99.5 % L1 hits and ~15 % of its instructions in integer divisions, not starved for its vectors as 3dbefa8's row supposed.
    • One card: down 49.92 → 40.51 us, kernel time −3.42 %.
    • -sm tensor: both rings faster in every round on both cards.
    • pp3 +2.57 % (CI +0.36 to +4.78). Same bits.
    • On the 288-expert proxy both rings run at 89–92 % of DRAM, and the several-token instance is within 0.2 % of the one-pair one under ncu.
  • c33735a: the ring's note on what it touches before its dependency wait now names the stop word.
  • 18fa139: at 2–8 columns and rows × cols ≤ 1024, a float mat-mul that mul_mat_f would run on fewer blocks than the GPU has SMs takes mul_mat_vec_f, and pairs on one input fuse. That covers glm5next's KDA f_a/g_a (4096 → 128 bf16, 8.6 us each on 4 blocks) and beta (6.4 us on 2 blocks).
    • The limit is measured: the vector kernel lost 1.1–4.4× past rows × cols 2048 in a first build without it. With it, all 74 shapes it takes run at 0.18–0.68 of mul_mat_f's time on the 5070 Ti and 0.23–0.63 on the 5080.
    • A KDA layer's gate span goes 33.63 → 18.18 us. Kernel time −2.76 % on one card, −5.71 % / −4.46 % under -sm tensor. pp3 +4.26 % (CI +1.51 to +7.02).
    • At -d 2048, where the DSA indexer scores: −3.1 % kernel time.
    • Not the same bits at several tokens: the vector kernel keeps the activations in f32 where mul_mat_f rounds them to bf16, and sums in another order. PPL at -ub 3 385634.0270 → 388534.0622. The KL divergence against the switch, 0.006795, matches -ub 1 against -ub 3 (order alone), 0.006763. Bit for bit at -ub 1 and under -sm tensor.
  • 2c5fb99: the ring decodes once for an expert's pairs only where ntokens × n_used ≥ n_experts. On GLM-5.3-Flash (8 of 288 experts) nearly every expert meets one pair, and that instance was 2.4 % / 1.4 % slower at -p 3 / -p 4 on 2 RTX PRO 6000. Measured locally:
    • The 288-expert proxy now takes the one-pair instance, rings −0.6 to −1.5 %.
    • A 16-expert proxy (~1.7 pairs an expert) keeps the several-pair instance, which beats one-pair there by 1.3–2.9 % (gate/up) and 7.9–9.6 % (down).
    • Same bits.
  • 98a6656: the comb beside the stream becomes opt-in (GGML_CUDA_HC_COMB_SIDE=1). GLM-5.3-Flash's tg64 was 3.7 % faster without it on 2 RTX PRO 6000 (4 of 4 pairs). Locally it is only +0.70 % (CI −0.15 to +1.55). Same bits.
  • 4135a7e, ff9a73c, cb7d859, c38e7a8, e815919, 09c216b, 2e699cc, 2606ae5, c47f13e, a993de2, 5b9b8e5: TORAD rows.

PPL is bit for bit wherever a row says so: 386825.5220 at -ub 1, 387892.2627 under -sm tensor, and at -ub 3 385507.0946 before 5c03ed9, 385634.0270 after it, and 388534.0622 from 18fa139.

Gates

  • 5b9b8e5 (tip), built alone in a scratch tree: test-backend-ops MUL_MAT 1475/1475, MUL_MAT_ID 1076/1076, MUL_MAT_VEC_FUSION 1032/1032, GATED_DELTA_NET 71/71, GATED_DELTA_NET_CACHE_FUSION 77/77 (and with the L2 persistence, PDL on and off), DSV4_HC_POST 4/4, DSV4_HC_PRE_FUSED 37/37, LORA_RANK1 3/3, SSM_CONV 45/45, SSM_CONV_STATE_UPDATE 48/48, MOE_FFN_CHAIN 10/10, MUL_MAT_ID_FUSION 13/13, MOE_WEIGHTED_SUM 18/18, DSV4_HC_PRE_Q8_1 16/16, DSV4_HC_PRE_POST 5/5, TOPK_MOE 320/320, ARGSORT 98/98, MUL_MAT_PAIR 120/120; DSV4_HC_PRE_FUSED, DSV4_HC_PRE_POST and DSV4_HC_PRE_Q8_1 again with GGML_CUDA_HC_COMB_SIDE=1; test-llama-archs; test-backend-meta split, views and capture on both cards.
  • 2606ae5 and 2e699cc, each built alone in a scratch tree, the same counts at both: test-backend-ops MUL_MAT 1433/1433, MUL_MAT_ID 1074/1074, MUL_MAT_VEC_FUSION 1032/1032, GATED_DELTA_NET 71/71, GATED_DELTA_NET_CACHE_FUSION 77/77 (and with the L2 persistence, PDL on and off), DSV4_HC_POST 4/4, DSV4_HC_PRE_FUSED 37/37, LORA_RANK1 3/3, SSM_CONV 45/45, SSM_CONV_STATE_UPDATE 48/48, MOE_FFN_CHAIN 10/10, MUL_MAT_ID_FUSION 13/13, MOE_WEIGHTED_SUM 18/18, DSV4_HC_PRE_Q8_1 16/16, DSV4_HC_PRE_POST 5/5, TOPK_MOE 320/320, ARGSORT 98/98; test-llama-archs; test-backend-meta split, views and capture on both cards.
  • Earlier tips: b850d70, bbe5370, 3a68863, 90d1b98, 4bce940 and c6b4ecf each passed the same scratch gate: every test-backend-ops group they touch, test-llama-archs, and the meta split/views/capture tests on both cards.

GLM-5.3-Flash: the sparse indexer, the mask scan and the hyper-connection front (Sep 30)

Eight commits on a 45-block proxy (33 KDA layers, 11 DSA, 1 dense lead), RTX 5080 + RTX 5070 Ti,
-sm tensor -ts 1/1 -fa 1. Every lever carries a *_LEGACY switch, and the three that claim bit
equality carry a *_CHECK switch that runs both paths and traps on the first differing bit.

  • 4a88ac9: the rows of a scatter are unique — the sparse mask's filler slots and the pooled-key
    write's unused slots. Its own evidence is still being settled by its author and its row says so; it is
    not claimed bit-identical here.
  • 69b9332: a node with no sources is mirrored. handle_generic returned UNKNOWN for such a node
    and allocation aborted on GGML_ASSERT(ret.axis != GGML_BACKEND_SPLIT_AXIS_UNKNOWN); ARANGE routes
    there and the sparse mask's dump columns are an arange, so every GLM-5.3 decode under -sm tensor
    aborted in ggml_gallocr_alloc_graph. New test-backend-meta-sourceless: red on the parent (exit
    134), green here. Upstream master carries the same UNKNOWN.
  • 9c506a4: every warp of the MMA flash-attention combine reaches one barrier, ported by hand from
    ggml-org/llama.cpp b74f590ea (ggml-cuda: fix divergent barrier in f16 flash attention ggml-org/llama.cpp#27870). compute-sanitizer synccheck: upstream's repro 3680 errors →
    0, GLM-5.3's cases 512 → 0. PPL bit for bit; flash_attn_ext_f16 within 1.79 % of base by ncu at base
    clocks with L2 flushed, against the instrument's own 1.44 % spread.
  • f8352b6: the sparse mask scan reads 32768 columns a round (1024 threads, four 16-byte loads
    each, a block-wide scan placing each thread's cells) where upstream's read 2048 with scalar loads.
    Index lists identical, so attention is bit for bit. Scan median a launch: 2.56 / 3.26 us at -d 8192
    and 4.54 / 4.42 at -d 32768, against 3.6 and 11.2 before. The FLASH_ATTN_EXT node at -d 8192 is
    0.62x / 0.58x of the dense path's, all three kernels counted. Wall tg32 @ d32768: +6.8 % and
    +4.1 %
    over two interleaved runs. Two new FLASH_ATTN_EXT cases at 40960 mask cells cover a second
    scan round; three mutants fail the sparse cases and only those.
  • 252f29c: two indexer changes landing together, since the second rewrites the kernel the first
    teaches to read rows.
    • ggml_lightning_indexer_rows: GGML_OP_LIGHTNING_INDEXER gains an optional src[4], I32 rows, so
      key i of stream s is k's row rows[i, s]. glm5next passes the f16 index cache's pooled head and
      pool_reps instead of materialising a get_rows f32 copy of every pool's key — at 32K cached
      tokens an ~20 us gather a layer, 8258 blocks and 4.2 MB written, 11 times a token. CPU reads the
      rows; Metal and SYCL refuse the variant; the meta backend mirrors it. Wall +2.86 % and +1.24 %.
    • lightning_indexer_kernel_quad: a key scored on a quad of lanes, two shuffles a key and head where
      the vector kernel took five and served one key with them; the vector kernel's scores bit for bit.
      By ncu at base clocks with all caches flushed, 20 launches a side: 0.516x / 0.424x / 0.536x at
      8258 pooled keys batch 1 and 3 and at 65536 f32 keys, 0.0–0.9 % drift on a repeated arm; the same
      shapes at boost clocks with a warm cache read 0.449x / 0.445x / 0.350x.
    • 26 new LIGHTNING_INDEXER rows cases; PPL identical to the last digit at ub 512 and ub 3, two runs
      each; two mutants, one of which (products summed xor-4-first) is invisible to a CPU comparison and
      is caught only by the check kernel.
  • c40b9d5: the routed-expert ring quantizes its own down-projection vectors. A quantize_q8_1
    launch ended 3.14 us after the ring it follows, 43 a token, 167 us of serial chain, and the ring then
    copied that q8_1 into shared memory anyway; its 512 consumer threads now produce it there past the
    dependency wait, through the same reduction tree and rounding, bit for bit.
  • 6fbde0b: the hyper-connection front runs in one launch instead of two a sublayer (88 a token).
    Each block's writers fence, thread 0 takes the token's ticket, and the block with the last ticket does
    the second kernel's work for the whole token through the same device functions, so the bits are the
    same. The two kernels were 2.39 + 3.90 us a front with 93 % / 91 % of cycles holding no eligible warp
    — latency, not work. Two mutants fail band 1 in distinct ways (320 and 544 check lines).
  • d769cb1: TORAD rows for the above.

Three wall bands missed, and why the misses are the instrument's

c40b9d5ce (declared ≥ 0.6 %), 6fbde0b8e (≥ 0.8 %) and the quad kernel in 252f29c0a (≥ 0.8 %) all
miss, and two of the three produced runs disagreeing in sign (−2.76 % / +1.52 %, and −1.22 % / +1.20 %).
That is the instrument, not the code:

lever measured node cost share of a 21.7 ms token ceiling at the wall
mask scan scan 11.2 → 4.5 us, 11 a token ~1.5 % several % (earned)
keys in place ~20 us gather a layer, 4.2 MB written ~1 % 1–3 % (earned)
ring quantize 43 × 3.14 us of serial chain = 167 us 0.8 % ~1 %
HC front 553 us cut to 0.75x 2.5 % ~0.6 %
quad indexer indexer_pool_score 196.8 us, halved 0.91 % ~0.5 %

Two runs of the same arm span 0.5 % at best on this host and 8.0 % when another build shares its CPU
(90.97 and 84.10 tok/s on one arm). tg32 -r 6 therefore cannot resolve a sub-2 % lever, and the quad
kernel's ≥ 0.8 % band was unreachable before a single run — the whole node is 0.91 % of the wall.
Below ~2 %, node attribution is the deciding instrument and the wall is a sanity check. The bands were
declared from each change's shape rather than from the node's measured share, which was already in hand;
each miss is recorded beside its declaration rather than dropped.

Where the two cards actually wait

From an NVTX capture at -d 32768 with graphs off (so host time is an upper bound), over the last 31
decode tokens, every >100 us idle gap charged to the nodes on either side of it:

  • A token is 21.7 ms on both cards. Card 0 is busy 13.84 ms (36.3 % idle); card 1 is busy 9.75 ms
    (55.1 % idle).
  • The idle is not launch overhead: sub-10 us gaps are 1.1–1.3 ms a token, while the >100 us class alone
    is 4.3 ms (card 0) and 8.7 ms (card 1).
  • Card 1 waits 28.2 times a token before a DSV4_HC_POST node, 6.58 ms in total; card 0 waits after
    its dsa_out / kda_out / ffn_shexp matmuls, 3.48 ms. The same per-layer tensor-parallel dependency
    seen from both ends.
  • The token boundary is 1.0 a token and costs 629 us (card 0) and 749 us (card 1) — symmetric, and
    3.5 % of a token, so not where the time goes.
  • Card 0 does 42 % more GPU work, so removing all idle still lands at its 13.84 ms: the ceiling from
    perfect overlap is 21.7 → 13.8 ms a token, 1.57x. Beating it needs a -ts other than 1/1, measured.

This is why single-kernel levers of the sizes above cannot move the wall much however good the kernel is,
and it names the next lever: fewer, later sync points a layer, and a weighted split.

Gates (continued)

  • 9c506a4, built alone in a detached scratch worktree and run from its own bin: test-backend-ops
    MUL_MAT 1475/1475, MUL_MAT_ID 1076/1076, MUL_MAT_VEC_FUSION 1032/1032, GATED_DELTA_NET 71/71,
    GATED_DELTA_NET_CACHE_FUSION 77/77, DSV4_HC_POST 4/4, DSV4_HC_PRE_FUSED 37/37, LORA_RANK1 3/3,
    DSV4_HC_PRE_POST 5/5, DSV4_HC_PRE_Q8_1 16/16, MUL_MAT_PAIR 120/120, TOPK_MOE 320/320,
    FLASH_ATTN_EXT 3229/3229; the cache-fusion persist leg at PDL on and off (77/77 each);
    test-llama-archs; test-backend-meta split, views, capture and sourceless; and the proxy under
    -sm tensor with CUDA graphs on and off, 4 of 4 rows and 0 asserts.
  • Each lever above also ran its own op suites on both cards by default, with its *_LEGACY switch, and
    under its *_CHECK switch where it has one.

Final node attributions, and one lever turned off by them

The node attribution the three missed wall bands pointed at has now run (NVTX, graphs off, the last 31
decode tokens at -d 32768, both arms from the same bin). It confirms two levers, bounds a third, and
reverses the fourth:

lever node, default → switch (card 1) declared verdict
f8352b6c2 mask scan FLASH_ATTN_EXT 268.4 → 666.4 us, 0.403x (card 0 0.452x) ≤ 0.5x earned, and 0.393x at d65536
252f29c0a keys in place indexer total 234.1 → 421.2 us, 0.56x; the 11 ~20 us pool-key GET_ROWS a token are gone ≤ 250 us (switch ~420) earned, every band
252f29c0a quad kernel indexer_pool_score 164.7 → 230.0 us, 0.716x (card 0 0.598x) ≤ 0.4x miss — the node is partly a wait, so a 2x kernel cannot halve it
c40b9d5ce ring quantize MUL_MAT_ID ffn_moe_down 862.5 → 1085.9 us, −223.4 us; the 43 quantize_q8_1-after-mmvq_moe a token are gone ≥ 100 us earned, 2.2x the bar
6fbde0b8e HC front one launch 759.4 us a token vs the two kernels' 535.4 — 1.62x, a regression ≤ 0.75x miss; made opt-in in e1c1ca7d8

Two notes on method, both of which changed a conclusion:

  • The ffn_moe_gate node is the ring lever's control: it is unchanged (1545.1 vs 1548.4 us), so the node
    that should not have moved did not.
  • For the HC front, a sum of two kernel times omits the gap between them, which is exactly what fusing
    removes — the wrong denominator for a latency lever. The NVTX issue span, which includes that gap, was
    measured too: 1475.4 us a token with one launch against 1395.0 with two. Both agree, so the regression
    stands. Cause: dsv4_hc_pre_gram_f32 summed every partial on 16 blocks in parallel, and the fused
    kernel's last-ticket block does all of it alone after fencing on 64 blocks. The 93 % / 91 % of cycles
    with no eligible warp that motivated the change were a small grid, not the launch boundary.

e1c1ca7d8 therefore makes the two kernels the default and the one launch opt-in behind
GGML_CUDA_HC_FRONT_ONE=1 — the same move 98a6656ae made for the comb beside the stream. Every property
6fbde0b8e verified is kept and was re-run against the flip: DSV4_HC_PRE_FUSED 37/37, DSV4_HC_PRE_POST
5/5, DSV4_HC_PRE_Q8_1 16/16, DSV4_HC_POST 4/4 by default, with the one launch, and under the check with 0
check lines. GGML_CUDA_HC_FRONT_CHECK=1 now implies the one launch so the check still has something to
compare.

Also earned here: 252f29c0a's KLD band, with its own control — default against the switch reads mean KLD
0.000000, max 0.000003, same-top 99.992 %, and the switch against itself reads exactly the same, so the
3e-6 maximum is the run-to-run floor rather than the lever.

…tate update, and the chain's weights and states requested before the PDL waits

The kernels between the qkv group and the Gated DeltaNet recurrence are a chain of short launches, each one's DRAM misses
paid on the critical path (rig's roofline research, 2026-09-27). Each lever below has its own switch; set to 1, the
switch restores the code before this commit.

The alpha/beta fold (GGML_CUDA_SSM_CONV_AB_LEGACY=1):
- Qwen3.5's two bf16 matvecs on the layer's normed input run inside the conv-state update (ggml_cuda_try_ssm_conv_ab,
  ssm_conv_ab_rows): one launch and one boundary fewer on the chain, 48 times a token. The blocks compute the rows
  before their dependency wait, mul_mat_vec_f's 256-thread arithmetic, so the values are the pair launch's bit for bit,
  and write them after it.
- The normed input is read before the wait, so it arrives by a release/acquire handoff, not the PDL chain.
  griddepcontrol.wait makes a prerequisite grid's writes visible to the waiting grid only. rms_norm_fwht_cuda releases
  to a handoff slot (ggml_cuda_ssm_conv_ab_slots) once its blocks have stored the input, and every conv block acquires
  it. The pair folds in only when the writer released this evaluation; otherwise it keeps its own launch.
- The PQ2_0 group writing the conv's inputs takes at most 152 registers and triggers its dependents after its own wait.
  At 152, a sub-partition keeps room for the conv's 56-register blocks beside it; at 168 it did not. Every other group
  keeps 168 and its early trigger; GGML_CUDA_PQ2_MMA_GROUP_REGS_LEGACY=1 puts that group at 168 too.

Requested before the waits:
- GGML_CUDA_PQ2_PREFETCH_BETWEEN_LEGACY=1: a PQ2_0 launch's L2 prefetch covers, before the next launch's heads, the
  weights of the kernels between the two (up to the budget). GET_ROWS tables, expert stacks and Hadamard rotation
  tables are skipped, since no kernel reads them whole.
- GGML_CUDA_SSM_CONV_PREWAIT_LEGACY=1: the conv's weights are loaded before its dependency wait.
- GGML_CUDA_SSM_CONV_STATE_PREFETCH_LEGACY=1: the conv's cache row goes into L2 before its wait.
- GGML_CUDA_GDN_STATE_PREFETCH_LEGACY=1: with the fused gather, the recurrence requests its head's state into L2 before
  its wait.
- GGML_CUDA_GDN_TRIGGER_LEGACY=1 and GGML_CUDA_MMVF_TRIGGER_LEGACY=1: the recurrence and mul_mat_vec_f let the next
  launch start at their start (mul_mat_vec_f_vec already did).

Also:
- The conv-state update is capped at 56 registers and has a 4-token instance (decode, an MTP verify of up to 3 drafts).
- llama-bench puts an NVTX range "prompt" around the measured prompt pass.

An adversarial review of the earlier version (29 agents, three votes a finding) confirmed three defects, fixed here:
- the fold never registered;
- every group took the 152 cap;
- the between-prefetch spent about 1.9 MB a segment on the rotation table.

Measured on the GLM-5.3 proxy, which has no bf16 alpha/beta pair and so no fold: tg64 59.72 / 59.72 / 59.45 on the
RTX 5070 Ti, same session. On Ternary Bonsai 2 27B, the Qwen3.5 model the fold is for, it is still to be measured, and
the bit-identity check is still to be run.
…ph before the recurrence, g_a beside f_a

The output gate g_b(g_a(x)) reads the layer input, not the recurrence's output, but the graph built it after the
recurrence: two small mat-vec launches on the chain between the recurrence and the output projection. It is now expanded
into the graph right after the convolution: g_a beside f_a, two mat-vecs of one input, which ggml-cuda's MUL_MAT pair runs
as one launch (5d58cc2), and g_b before the recurrence.

The 44-layer GLM-5.3 proxy on an RTX 5070 Ti alone, -cts f16, tg16 node trace: one small mat-vec launch fewer a KDA layer,
the token's chain -0.36 % (about 60 us of 16.6 ms); PPL at -ub 1 386825.5220 before and after, to the last digit. It is
also the order the next step needs: the recurrence's kernel applying the gated norm itself reads the gate.
…rnel, two launches fewer a KDA layer

After a KDA layer's GATED_DELTA_NET, the output gate normalizes the attention rows and multiplies them by
sigmoid(gate): RMS_NORM -> MUL(ssm_o_norm), then SIGMOID(gate) -> MUL. Before this, ggml-cuda ran them as two launches, a
fused rms_norm + weight and a fused sigmoid + mul. The first waited on the recurrence's just-written rows, the second on
the first, both on the layer's critical path.

ggml_cuda_try_gdn_gated_norm now registers the four nodes for the GDN node, which skips them. The recurrence's kernel then
writes the gated norm itself:
- Each (sequence, head) has a ticket. The last of its blocks to finish reads the head's rows back and normalizes them, a
  warp a row, and sets the ticket back to 0 (ggml_cuda_gdn_norm_tickets, zeroed once a context).
- The arithmetic is the unfused kernels'. rms_norm_f32 sums squares by a butterfly over each 32 columns and combines the
  32-column sums as its block_reduce does. Then its scale, the weight, and unary_gated_op_kernel's sigmoid product. So
  the output is the four nodes' output bit for bit.
- It fuses only up to 16 tokens a sequence (decode and the MTP verify), for 32- to 128-wide heads that the unfused norm
  reduces in one pass. Not on the chunked prefill path, not while the graph forks streams, and not when the output shares
  memory with what the GDN reads or writes, except exactly its attention rows or the gate.

glm5next builds the chain in place over the attention rows (ggml_rms_norm_inplace, ggml_mul_inplace), which nothing
reads after the norm. Built as separate tensors, the allocator put the output where it pleased, often on the GDN's own
inputs, which die at the GDN while other heads' blocks still read them. The matcher declined those (4 of 33 layers at
decode, all 33 in the 3-token verify graph). In place, the output is the attention rows in every graph, and all 33 fuse.
In place, the norm's RMS_NORM -> MUL no longer fuses on the unfused path (GGML_CUDA_GDN_GATED_NORM_LEGACY=1, or another
backend): one launch more there, the same values.

On the 44-layer GLM-5.3 proxy, RTX 5070 Ti, the norm on against GGML_CUDA_GDN_GATED_NORM_LEGACY=1:
- test-backend-ops: GATED_DELTA_NET_GATED_NORM 13/13 on and legacy (in place and as separate tensors: decode and verify
  over 1-2 seqs, f16/q8_0/f32 caches, the gathered state, 64- and 32-wide scalar-gate heads, and 20 tokens, which decline).
  GATED_DELTA_NET_CACHE_FUSION 77/77 and GATED_DELTA_NET 71/71. A kernel without the sigmoid fails the 8 fused cases of the
  out-of-place build and passes the declining one.
- PPL at -ub 1 (386825.5220) and at -c 128 -ub 8 over two sequences (389464.8783): every chunk the same.
- llama-server at one slot, greedy, 128 tokens with top-5 log-probabilities: bit-identical, plain and with the MTP draft,
  and the same over repeated runs of each arm. A kernel that flips the last bit of every fused output differs from
  position 0 in both.
- The node trace of tg16: rms_norm 23 a token against 56, unary_gated 0 against 33, all 33 KDA layers; llama-server's
  trace, every one of its 594 GDN launches followed by the wo projection's quantize.
- test-llama-archs: 0 FAIL (glm5next OK on the GPU, the CPU and both Meta configs).

GGML_CUDA_GDN_GATED_NORM_LEGACY=1 keeps the four nodes.
…536ed7), which measured neutral

7536ed7 had the recurrence's kernel write glm5next's KDA output gate: two launches fewer a KDA layer, bit-identical.
Measured against e74ed87 (the same graph built out of place, the norm on its own two launches), the 44-layer GLM-5.3
proxy, llama-bench rotated over 4 rounds:
- RTX 5070 Ti, tg128: +0.24 % (95 % CI -0.95 to +1.42); pp3 at ubatch 3, the MTP verify's shape: -1.34 % (-3.33 to +0.65).
- RTX 5080 + 5070 Ti under -sm tensor, tg128: -0.71 % (-1.75 to +0.33) for a restructured epilogue (the ticket before the
  state stores, loads hoisted), which on the 5070 Ti measured +0.83 % (-0.67 to +2.33) tg128 and -0.43 % pp3.

The node trace says why. The norm's first launch took about 4 us a layer, but that was the wait for the recurrence's
state stores to drain, which stays on the chain wherever it lands. Under PDL and CUDA graphs the two small launches cost
about 1 us each, and the epilogue's tail cost as much: the ticket's round trip, then the rows reloaded from L2, then the
stores. A change measured neutral doesn't earn its place on the default path. It also built the chain in place, which cost
the unfused path (other backends, the legacy switch) its RMS_NORM -> MUL fusion.

The tree is e74ed87's again. The restructured epilogue is kept as a patch in rig's research notes
(glm-matvec-l2-2026-09-29/gatednorm-v3-restructure.diff), with the served-bits check and the A/B that measured it.
…s order (e74ed87) and the gated norm taken out (7536ed7, b850d70)

The levers' row records what the served-bits check found beside the bits: the alpha/beta fold does not engage at
331d727 (the node trace's launch count is the same with its switch and without), and the register caps leave 16-32 B
stacks. The gated norm's row records why it measured neutral: the rms_norm's time after the recurrence is the wait for
its state stores to drain, which stays on the chain wherever the norm runs.
…its first tiles before its PDL wait

The routed-expert ring (mmvq-moe.cu) triggered the next launch at its start and read ids only past its dependency wait,
so a down projection's first rows were asked for only after its gate/up had ended and the q8_1 of the GLU's output had
run. The 44-layer GLM-5.3 proxy under -sm tensor, graph node trace: the down charged 18.94 us a layer on the RTX 5070 Ti,
against 14.5 at its DRAM peak.

- The ring triggers the next launch only past its own wait, so whatever the kernels before it wrote (the layer's top-k)
  is whole before any later kernel on the stream starts.
- A ring launch whose ids an earlier ring launch on the same stream read in this evaluation (the down after its gate/up)
  lists its experts and issues its block's first own tiles before its wait: a ring's worth at most, and never a ticket
  from the stream's counter, which the launch before may still be taking. ids are read past L1.
  GGML_CUDA_MMVQ_MOE_IDS_EARLY_LEGACY=1 keeps the old order.

Measured on the proxy, -sm tensor, -cts f16, RTX 5080 + 5070 Ti; graph node traces, 2 rounds alternated, the median of
every MoE layer:
- The down: 18.94 -> 17.25 us on the 5070 Ti, 17.76 -> 16.19 on the 5080.
- The gate/up is level (34.24 against 34.30); the quantize between is +0.12.
- A token's charged kernel time: 5080 7,779.5 -> 7,724.9 us (-0.70 %), 5070 Ti 8,885.7 -> 8,830.1 (-0.63 %).
- With the switch set, both cards are level with the base, so moving the trigger alone costs nothing.
- llama-bench tg128, rotated over 4 rounds: +0.75 % (95 % CI -1.43 to +2.94 over the rounds), too noisy to resolve a
  change this size.

Bit-identical: PPL at -ub 1 and -ub 3, and under -sm tensor, the same to the last digit in all three arms.

Tests:
- MOE_FFN_CHAIN, new: GLM's routed FFN, gate/up with and without the SwiGLU limit, then down on one ids; 1 and 3 tokens;
  a 2,048 and a 1,024 FFN. 8/8.
- A build whose tiles before the wait read the next expert's rows fails 8 of 8. It passes under the switch, and passes
  MUL_MAT_ID 1066/1066, whose single launches never take the path.
- MUL_MAT_ID 1066/1066, MUL_MAT_ID_FUSION 13/13, MUL_MAT_VEC_FUSION 1032/1032.
- The IQ3_XXS instances stay at 96 registers, with no stack.
…he activation fold that lost

The row records the mechanism measure (median charged time per MoE layer from node traces), each card's kernel time a
token, the round-paired throughput interval (with the finding that ab-arms.sh's old interval treated a run's samples as
independent, about 3x too narrow), bit-identical PPL, and the new MOE_FFN_CHAIN test with its mutants. It also records
the down quantizing its own activations, measured after it and not landed: +3.1 and +2.4 us a MoE layer, because every
block quantizes every pair's vector after its wait while its landed tiles wait in the ring.
…it the MUL and the ADDs it replaces

build_moe_ffn writes the weighted sum as MUL(experts, weights), a view of each slot, and the views added in slot order.
The MUL ran as one k_bin_bcast and the ADDs as a second, fused. ggml_cuda_op_moe_weighted_sum does both in one kernel.
Each thread takes one element of a token: e0*w0 + e1*w1 + ..., in slot order, with each product and each sum rounded on
its own (__fmul_rn, __fadd_rn), as the nodes round them. The matcher checks the MUL's shape, each view's offset and row
stride, the ADD chain's order, the subgraph's use counts, and that the output lies over neither input.
GGML_CUDA_MOE_WSUM_LEGACY=1 runs the two launches.

test-backend-ops MOE_WEIGHTED_SUM: n_embd 4096 and 2880, 8/4/2 slots, 1/3/8 tokens, compared bit for bit (the error is
the count of elements whose bits differ, the tolerance 0). 18/18 pass fused and 18/18 under the switch. An nsys trace of
the run shows k_moe_weighted_sum 18 times and no k_bin_bcast. Summing the slots in reverse order fails the 12 cases with
4 or 8 slots (1,304 to 20,226 elements differ); the 2 slot cases pass, as swapping two addends is the same add.

44-layer GLM-5.3 proxy under -sm tensor on a 5080 + 5070 Ti, -cts f16, tg32 node traces over two rotated rounds:
- the weighted sum, 1.83 and 1.76 us a MoE layer in two launches, is 1.25 and 1.22 in one;
- the chain from topk through the shared expert is -0.56 and -0.58 us a layer;
- each card's kernel time a token, without the all-reduce, is -0.33 % (7,730.9 -> 7,705.1 us) and -0.34 %
  (8,309.1 -> 8,280.7).
PPL is the same to the last digit: 386825.5220 at -ub 1 and 385507.0946 at -ub 3 on one card, 387892.2627 under
-sm tensor.
…a68863) and the three levers that lost

The row records the bit-for-bit MOE_WEIGHTED_SUM test with its reverse-order mutant, the sum's charged time a MoE
layer and each card's kernel time a token under -sm tensor (-0.33 % and -0.34 %), and bit-identical PPL. It also
records three levers measured the same way and not landed: mul_mat_vec_q triggering the next launch past its PDL wait
(it only moves charged time between launches), the shared expert's gate/up moved ahead of the routed experts with
topk's ids marked whole (+2.3 to +3.7 % a token, through the L2 issuer's plan losing the next layer's q, k and v), and
the issuer counting a MUL_MAT_ID as heavy (+4.82 / +3.95 % on the landed order).
…opy, in place of a quantize launch at its reader

The front's second kernel (dsv4_hc_pre_gram_f32) makes the sublayer's normed mix a warp to a q8_1 block, so it
quantizes each block as it writes it, with quantize_q8_1's arithmetic on the values it stores: the same bits. The
evaluation makes the copy for every such mix with a quantized reader on mul_mat_vec_q before any node runs (one copy a
key, which each layer's front writes in turn, as ggml-alloc gives every layer's mix the same bytes), and a reader of
one or more reads it. A copy remembers the index of the node that last wrote it, so the writes of the nodes in the
front's fused group, or before it, leave it whole. Where the front does not run fused, or n_embd is not a multiple of
MATRIX_ROW_PADDING, the first reader quantizes it as before. GGML_CUDA_MMVQ_Q8_1_PRODUCER_LEGACY=1 turns it off.

test-backend-ops DSV4_HC_PRE_Q8_1, new: the front's mix read by one or two Q8_0 or Q4_K MUL_MATs, 16/16. A doubled
block scale in the front's copy fails the 12 cases whose readers read it; the 4 that pass are the two where the reader
quantizes it (n_embd 256, 12 tokens) and Q4_K at 8 tokens, which Blackwell runs on MMQ.

The 44-layer GLM-5.3 proxy under -sm tensor on a 5080 + 5070 Ti, -cts f16, tg32 node traces: two quantize launches a
layer fewer (2,640 a run of 30 evaluations); each card's kernel time a token past the all-reduce -0.76 to -0.81 % and
-0.30 to -0.34 % over three runs, -0.68 / -0.56 % with the L2 issuer off. PPL bit for bit: 386825.5220 at -ub 1,
385507.0946 at -ub 3, 387892.2627 under -sm tensor.
)

The row records the new DSV4_HC_PRE_Q8_1 test and its doubled-scale mutant, two quantize launches fewer a layer,
each card's kernel time a token over three runs with the L2 issuer and one without, why the 5070 Ti (no all-reduce
slack, its fronts beside an L2 issue) keeps about half of the gain, the trigger variant that did not help, bit-identical
PPL, and the review's open item: no test fails if the copy silently stops being read.
…the chain to the sublayer's projections

The front's second kernel ran the comb's Sinkhorn (20 row and column normalizations in one warp) before it could end,
and the sublayer's first projections wait for it to end, though only the sublayer's DSV4_HC_POST reads the comb. It now
leaves the comb's inputs in their slots of the weights, and a warp a token on a stream of its own (forked after it)
makes the comb from them with the same function on the same values, so the same bits, and writes it over them. The
evaluation's stream waits for it before any node that reads those weights, in every DSV4_HC_POST, and at the
evaluation's end; not while the graph runs concurrent streams. GGML_CUDA_HC_COMB_SIDE_LEGACY=1 keeps it on the chain.

test-backend-ops DSV4_HC_PRE_POST, new: the front, then a DSV4_HC_POST reading its weights at once, 5/5. Without the
waits the post races the comb: the 1,000-iteration case fails 5 runs of 5, the 20-iteration ones by chance. Without
the side kernel, DSV4_HC_PRE_FUSED fails 31 of 36 (the 5 that pass run unfused), DSV4_HC_PRE_POST 5 of 5 and
DSV4_HC_PRE_Q8_1 16 of 16.

The 44-layer GLM-5.3 proxy under -sm tensor on a 5080 + 5070 Ti, tg32 node traces over two rotated rounds: the front's
second kernel ends 1.00 and 0.96 us sooner past its first, and the next launch still starts before it ends (PDL kept
across the fork); each card's kernel time a token past the all-reduce -1.03 % and -0.69 % against the switch set. PPL
bit for bit: 386825.5220 at -ub 1, 385507.0946 at -ub 3, 387892.2627 under -sm tensor.
… so work it leaves running shows

ggml_backend_compare_graph_backend computed the backend under test, then the reference, then read both. The
reference's evaluation sat between the backend's return and the read of its outputs, so a kernel the backend left
running past the return (a side stream's, not waited for at the evaluation's end) had finished by the read. The
reference now runs first. The results of a backend that finishes its work by the return are the same.

test-backend-ops DSV4_HC_PRE_FUSED, one case more: a comb of 1,000 iterations with no DSV4_HC_POST in the graph, so
only the evaluation's end waits for the comb CUDA makes beside the stream. The front's weights are marked an output,
as the tests read them back. With the evaluation's end not waiting (ggml-cuda.cu's ggml_cuda_dsv4_hc_comb_join after
the node loop removed), the new case fails 5 runs of 5 and 7 to 17 of the 20-iteration cases fail each run; with the
old order all 37 passed 5 runs of 5.
…nto weights read past the front

The side kernel writes the comb into the front's weights after the front's launch, and only a reader of the weights
or the evaluation's end waits for it. Weights read by the front's own DSV4_HC_PRE alone are free to the allocator
once that node is placed, so a later node's output may lie over them and the comb would land in it. The front now
makes the comb beside the stream only where the weights outlive it: a view of them besides pre's is in the graph (the
post and comb views, read by a DSV4_HC_POST that waits, or past the evaluation's end, which waits), or they are an
output; otherwise its second kernel makes the comb as before. The use count is the whole graph's, which the meta
backend's and the scheduler's subgraphs keep, so under -sm tensor, where the post is past the all-reduce in the next
subgraph, the comb stays beside the stream.

No model builds a front without its post, so nothing changes where one runs: the 44-layer GLM-5.3 proxy under -sm
tensor launches 100 dsv4_hc_comb_side a token as before, and PPL is bit for bit (386825.5220 at -ub 1, 385507.0946 at
-ub 3, 387892.2627 under -sm tensor). test-backend-ops cannot fail for the case this closes: it gives every tensor
its own memory, so nothing lies over the weights.
… to lie over its logits

ggml_cuda_check_fusion_memory_ranges let the top-k fusion's outputs overlap its logits at one row only, and ggml-alloc
places a layer's weights and ids over the logits whenever it can, so at an MTP verify's 3 tokens every MoE layer of
the GLM-5.3 proxy ran the unfused chain: 8 kernels (sigmoid, bias add, argsort, get_rows, sum, clamp, div, scale) where
one runs at a decode. topk_moe_cuda now takes 4 rows a block, a warp each, and its warps meet at a barrier once every
row's logits are read, before a weight or an id is written; ggml_cuda_topk_moe_reads_before_writes says so for the
rows that fit one block, and the check lets those overlap. GGML_CUDA_TOPK_MOE_ALIAS_LEGACY=1: at one row only.

The proxy at -p 3 -ub 3 (nsys, CUDA graphs): the routing chain from the hyper-connection front to the routed gate/up
8.32 -> 3.90 us a MoE layer on an RTX 5070 Ti, 8.32 -> 3.39 on an RTX 5080 and 8.19 -> 3.58 on the 5070 Ti under -sm
tensor; an evaluation's kernel time under -sm tensor -1.73 % (5080) and -1.55 % (5070 Ti), and its wall time -1.93 +-
0.49 % over 4 rotated rounds of 1,500 (the switch on: +0.01 +- 0.52). On the 5070 Ti alone the wall time moved -0.24
+- 0.39 %: the paced L2 issuer's request after the attention output runs ~25 us into the routed gate/up, which then
ends where the issue lets it, whenever it starts (with GGML_CUDA_L2_ISSUE_LEGACY=1 the evaluation's kernel time is
-0.96 %). A decode is unchanged (+0.10 % and -0.10 % a token).

PPL is bit for bit at -ub 1 (386825.5220) and under -sm tensor (387892.2627). At -ub 3 it is 385634.0270 against
385507.0946: the fused kernel normalizes the weights with its own rounding, as a decode's always did. KLD against the
-ub 1 logits: 0.006763 +- 0.000061 against the unfused chain's 0.006756 +- 0.000061. test-backend-ops TOPK_MOE (320),
MOE_WEIGHTED_SUM, MOE_FFN_CHAIN, MUL_MAT_ID_FUSION, MUL_MAT_ID and ARGSORT pass on the 5080, test-llama-archs too.
test-backend-ops gives every tensor its own memory, so no case overlaps the outputs and the logits; the proxy's layers
at 3 rows put the weights over the logits at the same rows, and without the barrier PPL was the same.
The plan sizes an issue by the graph nodes of the chain beside it (GGML_CUDA_L2_ISSUE_NODE_US each), but where the
evaluation fuses (a hyper-connection front, a top-k) many nodes are one launch, so the issue after the attention output
ran on into the routed gate/up. The ring then shared DRAM with it and ended where the issue let it: on the 44-layer
GLM-5.3 proxy at -p 3 -ub 3 on an RTX 5070 Ti, 97 % of the gate/up launches started with an issue running, which ended
27.2 us into them, and the gate/up took 125.2 us a KDA layer against 107.9 with no issuer, where the shared expert it
had brought into L2 gained back 9.2 us.

A ring launch (mmvq_moe) now bumps a word of the context's (ggml_cuda_l2_issue_stop, made before any capture as the
tile counters are) as its reads start, and the issuer reads the word when it starts and every 4th piece after, and
stops requesting once it has changed. GGML_CUDA_L2_ISSUE_STOP_LEGACY=1: the issuer ignores it.

The proxy at -p 3 -ub 3, nsys over 4 rotated rounds, the switch set against not: on the 5070 Ti alone the issue ends
9.3 us into the gate/up (p90 10.6) against 27.2, the gate/up 125.15 -> 108.26 us a KDA layer (107.94 with no issuer),
the shared expert's gate +3.2 and up +1.95, the layer -11.5 us, and an evaluation's kernel time -2.30 % (-0.58 % with
no issuer at all). Under -sm tensor no issue is running when a ring starts, with or without the switch, and every
issue lasts as long: the stop never fires there. PPL is bit for bit: 386825.5220 at -ub 1, 385634.0270 at -ub 3,
387892.2627 under -sm tensor. test-backend-ops MUL_MAT_ID, MUL_MAT_ID_FUSION, MOE_FFN_CHAIN, TOPK_MOE,
MOE_WEIGHTED_SUM and ARGSORT pass on the RTX 5080 and the 5070 Ti, test-llama-archs too.
… fragment once for its expert's pairs

At 3 tokens (an MTP verify's shape) the ring was bound by its math, not by DRAM: each pair routed to an expert
decoded every IQ3_XXS fragment of the expert's rows again (8 grid gathers from shared memory, 4 sign lookups and
the sign ops) before its 8 dot products. On the 44-layer GLM-5.3 proxy, where every token routes to each of its
8 experts, the gate/up took 113-116 us at 51-53 % of DRAM with its memory pipes 72-74 % busy (ncu, RTX 5070 Ti).

A launch of several tokens now runs its own instance (mmvq_moe<type, nmat, rpw, pairs_once>, IQ3_XXS at rpw*nmat<=2):
the expert's pairs 4 at a time (MMVQ_MOE_PB), each fragment decoded once into signed bytes (iq3_xxs_frag_decode)
and met with each pair's vector in turn (vec_dot_iq3_xxs_frag_q8), each pair's sums added in the same order,
so the same bits. Its own instance, as both paths in one kernel spilled (96 registers a thread at 17 warps an
SM): the one-pair path's gate/up went 108 -> 139 us. The one-token instances are unchanged: all 8,302 kernels of
b3fc826's libggml-cuda have the same SASS encodings (control words included), and the 3 kernels added are the
pairs_once instances. iq3_xxs_frag_pair holds the sign logic both use; vec_dot_iq3_xxs_frag meets each pair of
ints as it is decoded, as before, where decoding the whole fragment first put the one-pair down's 18 predicated
q8_1 loads behind branches and took it 61.4 -> 66.9 us a layer. GGML_CUDA_MMVQ_MOE_PAIRS_LEGACY=1: a pair at a time.

The proxy at -p 3 -ub 3, -cts f16, nsys over 4 rotated rounds: on the 5070 Ti alone the gate/up 107.81 -> 75.65 us
a KDA layer, the down 61.41 -> 49.66, an evaluation's kernel time -12.12 %; under -sm tensor -7.01 % on the RTX 5080
and -6.52 % on the 5070 Ti. llama-bench pp3 at -ub 3, 1,500 repetitions over 4 rotated rounds: 140.41 -> 154.00 t/s
on one card (+9.68 %, 95 % CI +8.58 to +10.79), 229.27 -> 244.48 under -sm tensor (+6.63 %, +4.33 to +8.93). tg128
reads +1.30 % (-1.39 to +3.99) on one card and -0.31 % (-2.43 to +1.82) under -sm tensor, measured with a draft
whose one-token instances were slower; these are b3fc826's. The down stays 21 us over its DRAM floor (28.7 us):
its tokens' vectors, 55 KB at 3 tokens of 8 slots, do not fit beside the ring and are read from global memory.

PPL is bit for bit, with the switch and without: 386825.5220 at -ub 1, 385634.0270 at -ub 3, 387892.2627 under
-sm tensor. test-backend-ops gains MUL_MAT_ID at IQ3_XXS with 8 experts all used at 2, 3, 5 and 8 tokens (past 4
pairs a pass at 5 and 8), gate/up and down shapes, and MOE_FFN_CHAIN on those experts at 3 and 5 tokens; a mutant
that stores pair q's sums from pair q+1's fails all 16 IQ3_XXS cases of several tokens on the ring (10 MUL_MAT_ID,
6 MOE_FFN_CHAIN) and passes the one-token ones. MUL_MAT_ID, MUL_MAT_ID_FUSION, MOE_FFN_CHAIN, TOPK_MOE and MUL_MAT
pass on the RTX 5080 and the 5070 Ti, test-llama-archs too.
…wait

With ids whole before the launch, a ring launch bumps the L2 issue's stop word (b3fc826) before its PDL wait; the
note on what the ring touches before that wait now says so. Only the issuer on its own stream reads the word.
@marcospaulo marcospaulo changed the title perf: Qwen3.5 levers3, glm5next's KDA output gate before the recurrence, the gated-norm fusion measured neutral and taken out perf: Qwen3.5 levers3; GLM-5.3 decode and MTP-verify levers (routed-expert ring, top-k at 2-4 rows, L2 issue stop, hyper-connection front) Sep 30, 2026
… tables and its tile and slot by fastdiv

ncu of the ring at 3 tokens (the 44-layer GLM-5.3 proxy, RTX 5070 Ti) put the down at 48 % of DRAM, its loads hitting
L1 99.5 % of the time and no warp eligible 55 % of the cycles, and ~15 % of its instructions in integer divisions:
each pair of each tile found its vector and its dst row from p / n_used, p % n_used and slot % nchannels_y, and each
tile its expert, row tile, slot and phase from tile / ntr and i / nslots, 11 division sequences a warp a tile.

The host now makes pair p's vector offset and dst offset (mmvq_moe_dev_args::pair_y, pair_dst: 64 pairs at most, a
constant load at an index the warp shares), and the kernel no longer takes the strides they came from; tile / ntr and
i / nslots are fast_div_modulo on values made on the host. Every address is the same, so the same bits. The down's
instance for several tokens goes 3,296 -> 2,792 instructions, the gate/up's 3,920 -> 3,408, the one-token down's 2,072
-> 1,872; 96 registers or fewer, no stack.

The proxy at -p 3 -ub 3, -cts f16, nsys over 4 rotated rounds, against 3dbefa8: on the 5070 Ti alone the down 49.92
-> 40.51 us a KDA layer, the gate/up 75.01 -> 74.66, an evaluation's kernel time -3.42 %; with
GGML_CUDA_MMVQ_MOE_PAIRS_LEGACY=1 the gate/up 101.86 (108 before). Under -sm tensor the gate/up 36.35 -> 34.82 us on
the RTX 5080 and 39.2 -> 37.3 on the 5070 Ti, the down 30.18 -> 29.5 and 33.7 -> 33.15, the same in each of the 4
rounds; the 5080's kernel time -0.68 %, the 5070 Ti's +2.21 % pooled from rounds that jump in both arms. pp3 at
-ub 3 on one card +2.57 % (95 % CI +0.36 to +4.78, host load 12-22). A decode's ring launches read the same or less:
the down -2.2 % and -2.8 % under -sm tensor, the gate/up -0.5 % on one card.

On the 4-layer proxy with GLM-5.3-Flash's 288 experts (the 5080, random routing: nearly every expert meets one pair)
the rings move -0.2 % (gate/up) and -1.2 % (down), and the several-token instance stays 1.2 % and 2.1 % over the
one-pair one there, as it was.

PPL is bit for bit, with the switch and without: 386825.5220 at -ub 1, 385634.0270 at -ub 3, 387892.2627 under -sm
tensor. test-backend-ops MUL_MAT_ID, MUL_MAT_ID_FUSION, MOE_FFN_CHAIN, TOPK_MOE and MUL_MAT pass on the RTX 5080 and
the 5070 Ti, test-llama-archs too; adjacent pairs' dst offsets swapped fail 37 MUL_MAT_ID cases and MOE_FFN_CHAIN 10
of 10.
…Ms takes the vector kernel at 2-8 columns

At an MTP verify's 3 tokens glm5next's KDA gate projections ran on mul_mat_f, which launches a block for each 32 rows
of each dst channel and reads a row's whole K in it: ssm_f_a and ssm_g_a (4096 -> 128 bf16) 8.6 us each on 4 blocks
and ssm_beta (4096 -> 64) 6.4 us on 2, RTX 5070 Ti, where at one token mul_mat_vec_f ran f_a and g_a as one launch.

On NVIDIA from Ampere on, at 2 to 8 columns and rows x cols <= 1024, where mul_mat_f would launch fewer blocks than
the device has SMs, mul_mat_vec_f now takes the product, and ggml_cuda_mul_mat_runs_mmvf says so, so a pair of them on
one src1 (f_a and g_a) runs as one launch (ggml_cuda_mul_mat_vec_f_pair). GGML_CUDA_MMVF_UNDERFILLED_LEGACY=1: mul_mat_f.

The limit is measured: test-backend-ops perf over f16/bf16 weights of 64-2048 rows, K 1024-8192 and 2-8 columns (RTX
5070 Ti) had the vector kernel at 0.21-0.67 of mul_mat_f's time at rows x cols <= 1024 but 1.1-4.4x slower past 2048
(mul_mat_f's time there follows K, the vector kernel's rows x cols). With it, all 74 such shapes take 0.18-0.68 of
mul_mat_f's time, 3 rounds (128 x 4096 bf16 at 3 columns 2.06 us against 7.62); the 118 other shapes 0.94-1.03.
On the RTX 5080, one round: 0.23-0.63.

The 44-layer GLM-5.3 proxy at -p 3 -ub 3, -cts f16, nsys over 4 rotated rounds on the 5070 Ti alone: a KDA layer's
span from its conv to its recurrence 33.63 -> 18.18 us (its float mat-muls 32.67 -> 16.90), an evaluation's kernel
time -2.76 % against 2606ae5 and -3.03 % against the switch; pp3 +4.26 % (95 % CI +1.51 to +7.02). Under -sm
tensor on the 5080 and the 5070 Ti, -5.71 % and -4.46 % against 2606ae5, the switch within 0.32 % of it. A decode
launches the same 1,247 kernels as before.

Not the same bits at several tokens: the vector kernel keeps the activations in f32 where mul_mat_f rounds them to
bf16, and sums in another order. PPL at -ub 3 385634.0270 -> 388534.0622, and bit for bit at -ub 1 (386825.5220) and
with the switch; against the switch's logits the mean KL divergence is 0.006795, where -ub 1 against -ub 3 (summation
order alone, on this random-weight proxy) is 0.006763. test-backend-ops MUL_MAT (new cases on both sides of each
limit, batched and broadcast), MUL_MAT_PAIR (glm5next's f_a/g_a pair), MUL_MAT_VEC_FUSION, SSM_CONV,
SSM_CONV_STATE_UPDATE, GATED_DELTA_NET_CACHE_FUSION and test-llama-archs pass on the RTX 5080 and the 5070 Ti; dst
zeroed after the new launch fails 31 MUL_MAT cases, all shapes the rule takes, and MUL_MAT_PAIR none (its pairs fuse).
… only where they are as many as the experts

3dbefa8 gave the ring an instance for launches of several tokens that decodes each IQ3_XXS fragment once and meets
it with every pair of its expert. That pays for an expert that meets several pairs and costs the instance's
bookkeeping on one that meets one. Where 3 tokens route to 8 of GLM-5.3-Flash's 288 experts, nearly every expert a
launch reads meets one pair: on a 4-layer proxy with its 288 experts (random routing, RTX 5080) that instance ran
the gate/up 1.2 % and the down 2.1 % slower than the one-pair one, and GLM-5.3-Flash on 2 RTX PRO 6000 under -sm
tensor was 2.4 % and 1.4 % slower end to end at 3 and 4 tokens (GGML_CUDA_MMVQ_MOE_PAIRS_LEGACY=1 against the
default, llama-bench -p 3 and -p 4, two rounds each).

The host now takes that instance only where the launch's pairs (ntokens * n_used) are at least as many as the
experts (ggml_cuda_mmvq_moe_args::n_experts, src0's channels): not for GLM-5.3-Flash at up to 8 tokens, still for
the 44-layer proxy, whose 3 tokens each route to all 8 of its 8 experts. GGML_CUDA_MMVQ_MOE_PAIRS_LEGACY=1 keeps
the one-pair instance whatever the counts.

At -p 3 -ub 3, nsys over 2 rotated rounds on the RTX 5080: on the 288-expert proxy the rings now launch the one-pair
instance (their demangled names in the traces) and a layer's gate/up takes 171.67 and 175.54 us where it took 174.23
and 176.95, its down 83.50 and 84.83 where 84.00 and 85.71 (the two layer kinds; the switch's arm, the same instance,
within 0.9 % of it). On a 4-layer proxy with 16 experts, where an expert that any pair meets meets ~1.7 of them at
random routing, the rule keeps the several-pair instance, which beats the one-pair one there: gate/up 107.17 and
109.71 against 108.58 and 113.03 us, down 51.36 and 51.71 against 55.79 and 57.22. The 44-layer proxy launches the
same kernels and instances as before.

The two instances make the same bits: PPL at -ub 3 388534.0622 on the 44-layer proxy and 337649.7140 on the
288-expert one, before and after.
test-backend-ops MUL_MAT_ID (a new case at 4 tokens on 32 experts, the rule's edge), MUL_MAT_ID_FUSION, MOE_FFN_CHAIN
and test-llama-archs pass on the RTX 5080 and the 5070 Ti.
…n (GGML_CUDA_HC_COMB_SIDE=1)

4bce940 made each fused hyper-connection front's comb on a stream of its own, off the chain to the sublayer's
projections, measured on node traces only: each card's kernel time a token 1.03 % and 0.69 % lower under -sm tensor
(the 44-layer GLM-5.3 proxy, RTX 5080 + 5070 Ti). On GLM-5.3-Flash on 2 RTX PRO 6000 under -sm tensor, tg64 was
3.7 % faster with GGML_CUDA_HC_COMB_SIDE_LEGACY=1 in 4 of 4 interleaved pairs (rig-glm, 2e699cc, llama-bench), and
on the proxy locally the stream is only +0.70 % over it (95 % CI -0.15 to +1.55, tg64, 4 rotated rounds).

So by default the front's second kernel makes the comb on the chain, as under the old switch, which is gone;
GGML_CUDA_HC_COMB_SIDE=1 makes it beside the stream as before. The same function on the same values either way:
PPL bit for bit (387892.2627 under -sm tensor by default and with the switch, 386825.5220 at -ub 1, 388534.0622 at
-ub 3). A tg node trace under -sm tensor launches no dsv4_hc_comb_side by default and 88 an evaluation with the
switch, one for each of the proxy's fronts. test-backend-ops DSV4_HC_POST, DSV4_HC_PRE_FUSED, DSV4_HC_PRE_POST and
DSV4_HC_PRE_Q8_1 pass on the RTX 5080 and the 5070 Ti each way, test-llama-archs too.
…vectors, bit for bit

Measured first, on an nsys capture at -d 32768 on the RTX 5070 Ti's main stream: a quantize_q8_1 launch
ends 3.14 us (median) after the gate/up ring it follows, 43 of them a token, 167 us of serial chain; the
ring's down projection then copies that q8_1 into shared memory anyway.

Where a routed launch's vectors are each its own (ne11 == n_used, which is a down projection), no other
MUL_MAT reads their q8_1, and the plan keeps them in shared memory, the host now passes the f32 vectors
and the ring's 512 consumer threads quantize them into that shared copy past the dependency wait: a
thread a float4 (eight of them a round trip to L2), 8 lanes a q8_1 block, the max and sum taken in
warp_reduce_max/sum<QK8_1>'s own tree (xor 16, 8, 4 across the 8 lanes as xor 4, 2, 1, then xor 2 and 1
within the lane's float4), then d, the rounding and ds exactly as quantize_q8_1 computes them. So the
results are the launch's bit for bit. GGML_CUDA_MMVQ_MOE_QUANTIZE_LEGACY=1 restores the launch, and
GGML_CUDA_MMVQ_MOE_QUANTIZE_CHECK=1 runs the ring from both and traps on the first result whose bits
differ.

Measured on an RTX 5080 + RTX 5070 Ti pair, -sm tensor -ts 1/1 -fa 1, on a 45-block proxy (33 KDA
layers, 11 DSA, 1 dense lead):

- test-backend-ops MUL_MAT_ID 1076/1076 and MOE_FFN_CHAIN 10/10 on both cards, by default and with the
  legacy switch, and under the check with no trap.
- R3: a mutant adding the float4's elements in another order ((0 + 1) + (2 + 3)) traps under the check
  on the down cases and passes test-backend-ops without it.
- Bit for bit: PPL under the check with the legacy switch's value, at -c 1024 -ub 1 --chunks 1 and
  -c 2048 -ub 3 --chunks 1 (the ring taking 1 to 8 tokens), and identical to the last digit at
  -c 8192 -b 8192 ub 512, ub 3 and ub 1 (349190.7774, 348976.2995, 321534.0329).
- Wall, tg32 at -d 32768, -r 6, reps 2-6 median, two interleaved runs: 90.19 and 91.67 tok/s against
  92.75 and 90.30 with the launch. The two runs disagree in sign, -2.76 % and +1.52 %.

The wall band asked for >= 0.6 % in both runs and is a miss, but the number to act on is the
disagreement, not the mean: 167 us of removed serial chain is about 1 % of a token here, and two runs of
one arm span more than that on this host while other work shares its CPU. tg32 at -r 6 does not resolve
a 1 % lever. The node attribution this lever was declared against -- no quantize_q8_1 after an mmvq_moe,
and the ffn_down_exps node at least 100 us lower on card 1 -- is the instrument that can see it, and it
has not run yet; it is recorded as owed beside the declaration rather than papered over with the one
positive run.
The front (dsv4_hc_front's Gram path) was two launches a sublayer, 88 a decode token on the 44-layer
proxy: dsv4_hc_mix_gram (64 blocks, the dot products and Gram partials over a 64-column slice) and
dsv4_hc_pre_gram_f32 (16 blocks, each summing every partial, then the pre weights, the RMS, block 0's
Sinkhorn, and its slice of the normed mix and its q8_1). An earlier capture on the RTX 5070 Ti at
-d 32768 with graphs off put them at 2.39 and 3.90 us a front, 210 and 343 us a token, while ncu read
0.23 and 0.05 waves with 93 % and 91 % of cycles holding no eligible warp. Latency, not work.

dsv4_hc_front_one runs mix_gram's blocks; each block's writers fence, thread 0 takes the token's ticket
(ggml_cuda_hc_front_tickets, one unsigned int a stream and token, zeroed before a graph evaluation the
way the PQ2_0 tile counters are), and the block holding the last ticket sets it back to 0 and does
pre_gram's work for the whole token through the same device functions dsv4_hc_pre_gram_f32 used -- the
partial sums, the pre weights and RMS, the weights, each element and its q8_1 -- so the bits are the
same. It reads the other blocks' partials streaming through L2 (__ldcg). One launch a front.
GGML_CUDA_HC_FRONT_ONE_LEGACY=1 launches the two kernels; GGML_CUDA_HC_FRONT_CHECK=1 runs the two into
scratch before the one launch and traps on any differing bit of the normed mix, the weights or the q8_1
copy.

Measured on an RTX 5080 + RTX 5070 Ti pair, -sm tensor -ts 1/1 -fa 1, on a 45-block proxy (33 KDA
layers, 11 DSA, 1 dense lead):

- test-backend-ops DSV4_HC_PRE_FUSED 37/37, DSV4_HC_PRE_POST 5/5, DSV4_HC_PRE_Q8_1 16/16 and
  DSV4_HC_POST 4/4 on both cards, by default and with the legacy switch, and under the check with 0
  check lines.
- R3, two mutants, each failing band 1 in its own way: the last ticket taken as n_slices - 2 produces
  320 check lines and 5 test failures; the ticket never set back to 0 leaves the next launch on the
  stream with no last block, 544 check lines and 56 failures, collapsing the suites to 6/37, 1/5 and
  1/16.
- Bit for bit: PPL identical to the last digit by default and with the legacy switch at -c 8192 -b 8192
  ub 512, ub 3 --chunks 1 and ub 1 --chunks 1 (349190.7774, 348976.2995, 321534.0329).
- Wall, tg32 at -d 32768, -r 6, reps 2-6 median, two interleaved runs: 86.17 and 86.72 tok/s against
  87.23 and 85.69 with the two launches. The runs disagree in sign, -1.22 % and +1.20 %.

The wall band asked for >= 0.8 % in both and is a miss. It was also the wrong gate for a lever this
size: the two kernels are 553 us of a token on card 1, so cutting them to 0.75x saves about 138 us,
near 0.6 % of the wall, while two runs of one arm on this host span 0.5 % at best and several percent
when another build shares the CPU. The band this lever should be judged on is its node attribution --
the one launch's median against the two medians, and 88 fewer launches a token on each card -- which
has not run yet and is recorded as owed beside the declaration.
…r, the wide mask scan, the indexer reading pooled keys in place with its quad kernel, the ring quantizing its down vectors, and the one-launch hyper-connection front
… measured slower than the two kernels

6fbde0b landed the front as one launch with its wall result unresolved and its node attribution owed.
That attribution has now run, and it is a regression. NVTX with graphs off, the last 31 decode tokens at
-d 32768, 87.8 fronts a token in both arms:

                 one launch (median, a token)    two kernels (medians, a token)      ratio
  RTX 5080        8.48 us,  754.5 us              2.11 + 4.22 us,  563.2 us          1.34x
  RTX 5070 Ti     8.38 us,  759.4 us              2.15 + 3.01 us,  535.4 us          1.62x

against a declared <= 0.75x. The launches a token do fall as designed, 1561.1 to 1473.3, exactly 88 fewer.

A sum of two kernel times omits the gap between them, which is the very thing fusing removes, so the
NVTX issue span was measured too -- it does include that gap. The DSV4_HC_POST ranges total 1475.4 us a
token with one launch against 1395.0 us with two, on card 1. Both denominators agree, so the conclusion
holds under the correction.

The cause is the ticket's shape, not the fusion idea. dsv4_hc_pre_gram_f32 summed every partial on 16
blocks in parallel; in the fused kernel the block holding the last ticket does all of that work alone for
the whole token, after fencing on all 64 blocks. Replacing 16-way parallelism with one block costs more
than the launch it removes. The two kernels' 93 % and 91 % of cycles with no eligible warp, which
motivated the change, were a symptom of a small grid rather than of the launch boundary.

So the two kernels become the default and the one launch is opt-in behind GGML_CUDA_HC_FRONT_ONE=1 --
the same move this fork made for the comb beside the stream at 98a6656 when it measured worse on the
cards. GGML_CUDA_HC_FRONT_CHECK=1 now implies the one launch, so the check always has something to
compare. Everything 6fbde0b verified is kept: DSV4_HC_PRE_FUSED 37/37, DSV4_HC_PRE_POST 5/5,
DSV4_HC_PRE_Q8_1 16/16 and DSV4_HC_POST 4/4 on the 5070 Ti by default, with the one launch, and under the
check with 0 check lines, re-run against this change.

A next attempt should keep the 16-way reduction: the last-ticket block spreading pre_gram's work across
its own warps, or a reduction over more than one block, rather than one block doing all of it.
…anges, the ub-512 logits part on a fusion the layout flips, and its decode cost
…he top-k, not a liveness chain in every layer

4a88ac9 made the rows of the mask's scatter unique with a chain per DSA layer: the exp of the pool bias at the selected pools,
repeated to cells, an arange of dump columns, the index arithmetic in f32, casts, and a concat of zero columns onto a copy of
sel_mask. It cost ~1.1% of decode kernel time and 128 launches, and its per-layer view of pool_bias, an input, was a split input
of its own (24 + 11 > 30), so the meta backend cut a second split and captured two CUDA graphs a token.
The top-k now does it. llm_graph_input_kpool carries select_k dump pools after the real ones: pool_dump, an f32 input of
-FLT_MAX concatenated onto the pool scores, above a dead pool's -inf and under every live score, so a row takes a dump pool
exactly where fewer than select_k pools are live, never a dead one; pool_cells names their cells n_kv + d*kpool + t, and
sel_mask carries those columns, -inf in every row, so the scatter writes them, no two slots of a row name one cell, and the view
back to n_kv drops them. build_attn_sparse scatters top_k as it is into a dup of sel_mask; `live` is gone from build_indexer and
build_attn_sparse. pool_cells' 3-D view is built once in build_inp_kpool, not in each indexer layer: a view of an input made
per layer was a split input, and a host copy each token, of its own.
test-kpool-input builds the maps with 0, a few or n_pools dump pools and checks, both paths, that the dump cells are n_kv + c and
the real ones under n_kv, and that every dump column is -inf in every row; it fails with a dump cell repeated (n_kv + c/r) and
with the dump columns left unwritten. test-kpool-can-reuse refuses a change of the dump count.
…rom the top-k, the parent's logits at ub 3 and 1, 95.7 fewer launches a token and the split inputs halved
requires_imatrix was asked of the target type alone, and raised before the
cur_type != new_type check that copies an already-correct tensor verbatim. So
requantizing a model that holds IQ-family tensors was impossible without an
imatrix even when those tensors were untouched: forcing the experts of an
IQ3_XXS model to IQ3_XXS, to hold them while the non-expert tensors move,
failed on blk.1.ffn_down_exps.weight with "this quantization requires an
importance matrix" for a tensor that would have been copied byte for byte.

--dry-run only records the requirement in will_require_imatrix instead of
raising it, so the dry run passed and the real run failed on identical
arguments -- the divergence that makes a dry run worthless as a check.

Gate the requirement on the type actually changing. The two later guards
(will_require_imatrix at the quantize decision, and the imatrix lookup) read
the same flag and would have fired spuriously for the same tensors.
The 8.9 GB of non-expert weight is 70 % of what a glm5next token reads, so the
byte lever lives there, and a format that cuts bytes 47 % while losing 30 % of
achieved bandwidth is not a 47 % win. Adds the four real shapes -- two KDA
projections, the DSA output and the 4096 x 154880 head -- across Q8_0, Q4_K,
IQ4_XS, MXFP4 and NVFP4, at one row and at an MTP verify's width.

bs=1 takes the vector kernel, where the block-scaled MMA that NVFP4 and MXFP4
exist for cannot apply; bs=3 is where it can, so the ranking may differ between
them and both are measured.

Caveat for whoever reads the output: perf reports us/run and TFLOPS, never
bandwidth, and at these shapes every arm but the head fits in an RTX 5080's
67 MB L2 -- Q8_0 there appears to reach 2816 GB/s against a 960 GB/s DRAM peak.
Only the head exceeds L2 in every format, and there all five formats land
within 0.7 % of each other at 905-912 GB/s with time tracking bytes to 0.4 pp.
The smaller shapes need ncu --cache-control all to say anything about DRAM.
KL divergence was scored only over the second half of each chunk, and the
per-token values were sorted before reporting, so position was discarded. For a
recurrent architecture that hides the thing worth measuring: a KDA layer carries
its state forward, so a quantized tensor written INTO the state (attn_k, attn_v)
has its error compound with distance from the start of the sequence, while a
tensor feeding only the readout (attn_q) stays flat. The two are
indistinguishable in a corpus mean and are not remotely the same risk. At
n_ctx 32768 the default also means positions 0-16384 are never evaluated, so no
amount of context length produces an early-position bin.

- LLAMA_KLD_FIRST sets the first scored position. One helper serves both the
  base-logits writer and the reader, which must agree: it sets how many values a
  chunk occupies in the base file.
- The three buffers were sized from the hardcoded n_ctx/2 before first was
  defined; a lower first overran all of them. They now size from first.
- The header carries n_ctx, n_vocab and n_chunk but not first, so a base file
  written at the default and read at 0 would misalign every chunk and report a
  plausible, meaningless dKLD. The payload length is now checked against what
  first implies and the run refuses with both counts.
- dKLD is reported per position bin (0-512, 512-4k, 4k+) before the sort: count,
  mean, median, p99 and mean |dp|. An empty bin prints why rather than vanishing,
  because an empty late bin means the run cannot see accumulation at all.

Verified on a 4-layer proxy, three claims: an identical model gives dKLD
0.000000 with p99 0.000002 in both populated bins, so the binning indexes
correctly; the bins fill with 2048 and 2044 of 4092 values over 4 chunks and 4k+
reports itself empty; and reading a first=0 base file at the default refuses,
"holds 1267570656 payload bytes but this run expects 633165792" -- the 2.0020
ratio of 1023 to 511 values a chunk.
…profiler

nsys cannot measure this call: it inflates it about 39x. A standalone 400-node
graph that its own process times at 1.3 us with clock_gettime reads 50.6 us by
that same clock once nsys is attached, and nsys itself reports 51.1 us -- the
instrumentation consumes the host time, it does not merely misreport it. So any
host-side attribution drawn from a trace is wrong at this scale.

With this flag and no profiler, a tensor-split decode on two cards reports a mean
of 5.9-13.6 us a launch, against the 246-294 us the same launches measure under
nsys. That matters because a trace had put 66 % of a decode token's 444.6 us
device-idle gap on this call; the real share is about 3 %, and three fixes aimed
at it (fusing nodes to shrink the graph, launching the two devices' graphs from
separate threads, and the scheduler's extra input copies) all measured null for
that reason.

Off by default, costing one test of a static flag. Prints straight to stderr
because llama-bench installs a log callback that drops INFO and WARN.
… goes

An RAII accumulator over build_graph, alloc_graph, set_inputs and graph_compute,
with decode() itself as the denominator so whatever the phases do not account for
is printed as "unattributed" rather than assumed to be zero. Timed with
clock_gettime in-process, because nsys inflates CUPTI RUNTIME durations 30-40x on
this machine and cannot answer a host-attribution question.

What it finds on a tensor split, decode only, no profiler: decode() returns in
1.32 ms while a token takes 10.08 ms, so over 128 tokens the host spends 169 ms
against the device's 1290 ms and is idle 87 % of the time. set_inputs is 2.7 us.
The host runs 7.6x ahead of the device and is not the bottleneck.

That retires the "444.6 us token boundary" this campaign chased: it was mostly the
profiler's own host cost. It also explains why fusing nodes, threading the two
devices' launches, and the scheduler's extra input copies all measured null --
there was no host boundary for them to recover.

Off by default, one static flag test a timer. Prints to stderr because llama-bench
drops INFO and WARN.
The dKLD table bins the divergence by position; a head that picks a
different token than the base is a different question from a divergence
spread over the distribution's tail, so each bin now also reports how
often the top token agrees with the base's, with its binomial error.
The column is appended last: the earlier columns are unchanged.
…norm a warp to a row, the hc norm writing cuBLAS's BF16 operand

Measured first, nsys of a pp4096 prefill on the 44-layer proxy (RTX 5080 + RTX 5070 Ti, -sm tensor -ts 1/1
-fa 1 -b 4096 -ub 1024, graphs traced by node), each kernel billed its own time: its end less the later of
its start and the end of the kernel before it on its stream. Under PDL a kernel starts while its
predecessor runs and nsys's duration bills it that wait.

- KDA's per-head output norm (rms_norm_f32<256> with its weight, rows 128 wide, 32 a token) read 690 us a
  launch in nsys's durations: 622.6 of them waiting on gated_delta_net_cuda, 67.8 its own.
- Each hc mix, a weightless RMS_NORM over the 16384-wide streams into hc_fn (a BF16 MUL_MAT, on cuBLAS at
  prefill), was the norm (143.4 / 161.2 us, 5080 / 5070 Ti), convert_unary's cast of its output to BF16
  (148.2 / 164.7 us) and a 21 us GEMM.

Two changes:

1. rms_norm_f32_warp_rows: a row of at most 256 columns gets a warp, a block eight rows, where the block
   kernel gave it 256 threads (half idle on a 128-wide row) and two barriers. A lane plays the block's
   threads lane, lane+32, ..., lane+224: each virtual warp's squares go through the same butterfly and the
   eight partials through block_reduce's second stage, so every output is the block kernel's bit for bit.
   GGML_CUDA_RMS_NORM_WARP_ROWS_LEGACY=1 keeps the block kernel.

2. A weightless RMS_NORM read only by a MUL_MAT that runs on cuBLAS in BF16 writes the BF16 copy cuBLAS
   would have cast (rms_norm_f32 takes a dst type, rounded with convert_unary's ggml_cuda_cast) to a pool
   buffer cuBLAS reads, and the F32 output is neither written nor read. ggml_cuda_mul_mat_runs_cublas_bf16
   asks ggml_cuda_mul_mat's own predicates, and ggml_cuda_mul_mat_cublas_compute_type (split out of
   ggml_cuda_mul_mat_cublas) the compute type: GGML_PREC_F32 or a GGML_CUDA_CUBLAS_COMPUTE_TYPE that would
   cast the operand to F32 keeps the pair unfused. GGML_CUDA_RMS_NORM_BF16_LEGACY=1 runs the norm and the
   cast on their own.

Results, same proxy and cards:

- Bit for bit: the proxy's KLD base file (-c 2048 --chunks 8 on wikitext-2: each scored token's scale and
  min log-prob as floats, then its 16-bit log-probs over the vocab), 2,535,206,868 bytes, cmp-identical
  to the parent's with both changes on (and the next commit's decay precompute), with the warp rows off,
  and with every switch set. R3: a mutant summing a lane's squares before one butterfly differs in
  2,194,756,864 of those bytes, one rounding the BF16 toward zero in 2,371,020,141. The five f32
  instantiations of rms_norm_f32 compile to the same SASS as before its dst type.
- KDA's output norm: its own time 67.8 -> 24.5 us a launch, 0.41 % of the prefill's device time. pp4096
  in 4 interleaved pairs +1.53 %, 95 % CI +0.04 to +3.03 (t 3.182), which holds the 0.4 % the kernel time
  predicts and does not resolve it.
- The hc mix: norm and cast 291.6 / 325.9 -> 93.9 / 117.0 us a mix (5080 / 5070 Ti), 704 a card; the
  prefill's device time 5590.1 -> 5285.1 ms (-5.5 %). pp4096 in 6 interleaved pairs +4.40 %, 95 % CI
  +3.60 to +5.19 (t 2.571).
- test-backend-ops on both cards by default, and on the 5080 with each switch: RMS_NORM 57/57 (new: rows
  over more than one 8-row block, the last partial), RMS_NORM_MUL 6/6 (new: the mul-only fusion KDA's norm
  takes, which RMS_NORM_MUL_ADD never
  reached), RMS_NORM_MUL_ADD 36/36, RMS_NORM_MUL_MAT 27/27 (new: 64 columns, on cuBLAS; its F16 and F32
  weights stay unfused). An nsys of RMS_NORM_MUL_MAT launches the BF16 norm at both block sizes and no
  f32 -> bf16 cast.
… recurrence from 32 tokens, bit for bit

The recurrent kernel gives each state column a warp, and with KDA's per-channel gate each warp of a head
activated and exponentiated all S_v of a token's decays itself (exp(raw_lb * sigmoid(-(g * raw_a[h]))),
two expf and a division an element): every value S_v times over. From 32 tokens a sequence
gdn_kda_precompute_decay now writes each once, to a pool buffer the kernel reads as its decay
(G_PRECOMPUTED, which with KDA leaves RAW to activate beta alone). The formulas are the kernel's own, so
the values are the same bit for bit. Under 32 tokens (decode, a verify) the kernel keeps computing them,
32 being the GB10 scalar-gate precompute's threshold, not one measured here.
GGML_CUDA_KDA_DECAY_PRECOMPUTE_LEGACY=1 leaves them to the kernel at every size.

Measured on the 44-layer proxy, RTX 5080 + RTX 5070 Ti, -sm tensor -ts 1/1 -fa 1 -b 4096 -ub 1024:

- Bit for bit: the proxy's KLD base file (-c 2048 --chunks 8 on wikitext-2: each scored token's scale and
  min log-prob as floats, then its 16-bit log-probs over the vocab), 2,535,206,868 bytes, cmp-identical
  to the grandparent's with this and the parent's changes on. R3: a mutant shifting every decay one ulp
  toward zero differs in 2,270,055,044 of those bytes. (A mutant taking __expf for expf did not differ:
  ggml-cuda builds with -use_fast_math, where expf is __expf.)
- nsys of a pp4096 prefill, graphs traced by node, each kernel billed its own time: gated_delta_net_cuda
  1351.2 / 1387.3 -> 1216.8 / 1274.3 us a launch (5080 / 5070 Ti), plus 16.4 / 21.2 us of
  gdn_kda_precompute_decay, net -8.7 / -6.6 %; the prefill's device time -1.0 %. Taking the gate's
  instructions out of the token loop took 8-10 % off the kernel: they were not what bounds it.
- pp4096 in 6 interleaved pairs +1.29 %, 95 % CI -0.42 to +3.00 (t 2.571): it holds the 1 % the device
  time predicts and does not resolve it.
- test-backend-ops GATED_DELTA_NET 72/72 (new: KDA at 64 tokens with snapshot slots, two sequences and
  GQA, and at 32 with the rows-indexed state) and GATED_DELTA_NET_CACHE_FUSION 77/77 on both cards, by
  default and with the switch.
…a warp to a row, the hc norm writing cuBLAS's BF16 operand, +4.40 % pp4096; and KDA's decay once per token (58ef198), -1.0 % device time; all three bit for bit with mutants that differ
…ts prefill kernel time halved, pp4096 +2.9 %

KDA (GLM-5.3-Flash's linear attention) ran its prefill on the recurrent kernel, 12.6 % of a pp4096's device time on
the 44-layer proxy: the chunked pipeline took only the scalar gate. It now takes KDA's gate, one per key channel, in
FLA's chunk_kda form, G being each channel's log-decay summed along a 16-token chunk, so every exponent is <= 0:

- stage 1, cgdr_kda_fwdsub_intra_kernel: the activated gate (RAW with the recurrent kernel's formulas), G, and in one
  fp32 pass sharing each exp(G[t] - G[s]) both L[t][s] = beta[t] sum_i k[t][i] k[s][i] exp(G[t][i] - G[s][i]) and the
  masked A[t][s] = sum_i scale q[t][i] k[s][i] exp(...), so stage 2 does not run; then (I + L) x = b for
  beta exp(G) k and beta v as for the scalar gate
- stage 3, cgdr_state_wmma_kernel<..., KDA>: q scaled by exp(G[t]) and k by exp(G_last - G[t]) channel by channel before
  their GEMMs, and each state row k decayed by its own exp(G_last[k]); g_cum holds G per channel ([CS][K] a chunk)
- K > 1: the last K-1 tokens run on the recurrent KDA kernel from the chunked state, g read at the beta strides
  times S_v

GGML_CUDA_KDA_CHUNKED_LEGACY=1 keeps KDA on the recurrent kernel (GGML_CUDA_GDN_CHUNKED=0 still turns the chunked
path off for both gates). Where the chunked path runs, an f16 or q8_0 KDA state cache keeps its cpy and a gathered s0
its GET_ROWS, as for the scalar gate (the pipeline writes and reads f32).

Measured on the 44-layer proxy, RTX 5080 + RTX 5070 Ti, -sm tensor -ts 1/1 -fa 1 -b 4096 -ub 1024, the served env:

- nsys of a pp4096 prefill, graphs traced by node, each kernel billed its own time: KDA 1215.6 + 16.4 (the decay
  precompute) -> 173.5 + 433.7 us a layer (stage 1 + stage 3) on the 5080, 1271.4 + 20.2 -> 208.3 + 441.5 on the
  5070 Ti, -50.7 / -49.7 %; the prefill's device time 5202.9 -> 4819.3 ms (-7.4 %), of which the kept GET_ROWS gives
  back 1.1 / 16.4 ms.
- llama-bench in interleaved pairs: pp4096 +2.93 % over 6 (95 % CI +1.60 to +4.26, t 2.571); pp16384 +3.81 % over 2
  (95 % CI -2.08 to +9.69, t 12.706), unresolved.
- Quality, llama-perplexity KL against a recurrent base on wikitext-2: -c 2048 --chunks 8 (positions 1024-2047) mean
  0.006462, max 0.010349, same top p 79.86 %, ln(PPL ratio) -0.00044 +- 0.00125; -c 16384 --chunks 2 (positions
  8192-16383) mean 0.006547, max 0.011151, same top p 78.98 %, -0.00006 +- 0.00089: no growth along the context. The
  recurrent kernel against its own base: 0.000000, 99.98 %. The proxy is chaotic (PPL 344,555): in the 2026-09-28
  bisect (-c 512) a -ub 512 -> 256 change alone cost it 0.0058 (78.4 / 79.2 % same top p) and a reduction order
  alone (-sm tensor with an f32 wire against -sm layer) 0.0026. What it costs the trained model is the first
  measurement for a box with the real pack, the switch backing it out in the serving env.
- test-backend-ops GATED_DELTA_NET 79/79 (new: KDA from 128 tokens, activated and raw gates, the partial last chunk
  over two sequences, GQA, permuted q/k/v, snapshot tails at K 4 and 9) and GATED_DELTA_NET_CACHE_FUSION 80/80 (new:
  KDA prefill with the f32 cache fused, an f16 one behind its cpy, a gathered s0) on both cards, by default and with
  the switch. The chunked KDA cases' NMSE is at most 1.26e-7 over 6 runs on both cards, under the chunked path's
  2e-7 (the scalar gate's at most 1.02e-7).
- R3: a mutant decaying each state row by its neighbour channel's G_last at the chunk ends fails all 7 chunked KDA
  cases, NMSE 0.75-1.72. With KDA's test gate drawn independently per token, as it was, the same mutant failed only
  the two snapshot-tail cases and passed the 32-head 512-token one at NMSE 9e-8: a channel slow on one token is fast
  on the next, so little state reached a chunk boundary (a float64 simulation of the raw gate's test distributions
  puts the mutant's effect at 1e-13 to 2e-11). init_kda_gate gives each channel and head a rate and each token a
  little jitter, so a slow channel stays slow; both GATED_DELTA_NET tests use it for KDA.
…nel time halved, pp4096 +2.93 %, KL 0.0065 on the proxy at 2k and 16k, the trained model's KL the next box's first measurement
…A's now does

init_kda_gate gave KDA's gate a rate per channel and head with per-token jitter, so a
slow channel stays slow and its state survives a chunk. The scalar gate kept the i.i.d.
per-token draw that motivated that fix, so it had the same hole in a path already
serving: every head is fast on some token of every chunk, no state reaches a boundary,
and nothing tests what carries it there.

The generator already generalises -- for the scalar gate's [1, heads, tokens, seqs],
`i % (ne0*ne1)` reduces to the head index -- so it is renamed init_decay_gate and used
for both, and the ranges are unchanged. Only the correlation across tokens changes, and
that correlation is the whole coverage.

Measured, since the point is a gate that can fail. A row carries state across a
16-token chunk when |g| < -ln(1e-3)/16 = 0.43. Drawn per token that needs all 16 draws
in the slowest 2.2 % of the range: p = 2e-27 analytically, and 0 of 20,000 rows in a
direct draw from the generator. With a per-row rate it is 0.69 analytically and
13,767/20,000 = 68.8 % measured.

Tests on card 0: GATED_DELTA_NET 79/79, GATED_DELTA_NET_CACHE_FUSION 80/80, so the
harder data does not move the real path off its 2e-7 bar.

Not yet proven: that a mutant breaking the scalar gate's chunk-boundary carry now fails.
That needs a temporary edit in gated_delta_net.cu or chunk_gated_delta_net.cu, which
another seat holds, so it is requested rather than done. The KDA half of this argument
was proven that way (NMSE 0.75-1.72 against 9e-8 before), and this change is the same
mechanism on the same generator.
… query rows

Step one of query-tiled sparse attention. The sparse kernel is one query a tile by
construction (fattn-mma-f16.cuh's "a tile's rows share its indices"), so each query
gathers its own ~2080 cells: at p28672 flash-attention is 22.0 % of own device time,
second only to mmq, and the only item that grows with depth on a head served at
-c 524288.

flash_attn_mask_to_sparse_indices is now templated on NROWS, the query rows sharing one
index list, and ORs their finite-bitmasks before the existing popcount, scan and
compaction -- which are all row-agnostic, so the union is a four-line change inside the
scan loop. Correctness does not depend on the overlap: flash_attn_ext_f16_load_mask
already reads each row's own mask value at every gathered cell, so a cell a row did not
select reads -inf for that row.

Default width is 1 and FLASH_ATTN_EXT passes 3231/3231, so the shipped path is
unchanged. GGML_CUDA_FATTN_SPARSE_ROWS=N (2, 4, 8, 16) asks for a wider list, and
because the kernel still reads one list a query a width above 1 would read the wrong
list rather than fail -- so it aborts instead, and the abort is verified reachable
rather than assumed.

What this does NOT establish is whether a union pays. That is the overlap between
adjacent queries' selections, a property of the trained indexer: a tile of 8 loads 1.4x
one query's cells at 95 % overlap and 8x at none. A random-weight proxy reports the
no-overlap case by construction, so the size of this lever is not measurable here by any
instrument, and no predicted gain is recorded. The mechanics and bit-identity are
testable on the proxy; the economics need one leg on a real-weight box.
… 16-byte load a thread a chunk ahead: its time halved, bit for bit

After 7755e85, chunked KDA's stage 3 was 70 % of its time, 433.7 / 441.5 us a layer on the proxy's 5080 / 5070 Ti.
Each of its 128 blocks (32 heads, 4 v-tiles of 32) loaded three 16x128 fp32 operands a chunk with eight scalar loads a
thread, converted them to fp16, and applied two exps an element (q * scale * exp(G), k * exp(G_last - G), and
k_cumdecay as is). That per-block work is what bounds it: halving the tile to BV=16 doubles the blocks and costs
+61 / +78 %, and 128 threads instead of 256 cost more.

Stage 1 (cgdr_kda_fwdsub_intra_kernel) holds q * scale, k, G and k_cumdecay in shared memory already. It now writes
stage 3's operands in their final form, fp16 [CS][K] a chunk, with each state row's decay over the chunk exp(G_last[k]),
in place of fp32 k_cumdecay and G: the same formulas on the same values, rounded to fp16 as stage 3 rounded them.
cgdr_kda_state_wmma_kernel, KDA's own stage 3, loads each operand as one 16-byte load a thread and issues it for the
next chunk as soon as this chunk's copy is in shared memory (nothing but H carries from chunk to chunk), so the loads
land while the chunk's GEMMs run. The scalar gate's cgdr_state_wmma_kernel is back to its text before 7755e85. KDA's
scratch is 6.25 bytes a token, head and channel, from 8.

Measured on the 44-layer proxy, RTX 5080 + RTX 5070 Ti, -sm tensor -ts 1/1 -fa 1 -b 4096 -ub 1024, the served env:

- Bit for bit: the proxy's KLD base file (-c 2048 --chunks 8 on wikitext-2), 2,535,206,868 bytes, cmp-identical to
  7755e85's chunked KDA with the final artifact (libggml-cuda 511f59662369). R3: q's operand one ulp toward zero
  before its rounding to fp16 differs in 2,229,794,626 bytes.
- nsys of a pp4096 prefill, graphs traced by node, each kernel billed its own time: stage 3 433.7 / 441.5 ->
  212.2 / 210.8 us a layer, stage 1 173.5 / 208.3 -> 167.7 / 201.0; KDA 607.2 / 649.8 -> 379.9 / 411.8 us a layer
  (-37.4 / -36.6 %), against the recurrent kernel's 1232.0 / 1291.6 (-69.2 / -68.1 %); the prefill's device time
  4819.3 -> 4709.4 ms (-2.3 %).
- pp4096 against 7755e85's library (its chunk object compiled with the build's flags and relinked, the same
  executables) in 10 interleaved pairs: +1.52 %, 95 % CI +0.55 to +2.50 (t 2.262).
- test-backend-ops GATED_DELTA_NET 79/79 and GATED_DELTA_NET_CACHE_FUSION 80/80 on both cards, by default and with
  GGML_CUDA_KDA_CHUNKED_LEGACY=1.
…33.7 -> 212.2 us a layer, pp4096 +1.52 % against 7755e85, bit for bit
Every batched sparse FLASH_ATTN_EXT case (nb 8, 16, 64, 512) sits in
make_test_cases_perf, which times kernels and never compares them against a
backend. Eval reached the sparse path with at most 3 queries, so batched sparse
attention (the shipped one-query-a-tile path) had no reference check at all. A
probe in the dispatch showed it: 9 sparse dispatches in the whole suite, the
largest Q->ne[1] 3.

Four eval cases reach it: nb 8, 9 (a tile and a 1-row tail) and 16 at
n_kv_max 512 over 4096 cells, where the heuristic admits the gather cheaply,
and GLM-5.3's own bound (n_kv_max 2080, 32 heads on the latent) over 16640
cells at nb 8. FLASH_ATTN_EXT 3235/3235 on the RTX 5080 and the 5070 Ti.
…e sparse gather

The sparse kernel gathered each query's own cells, one query a tile. With
GGML_CUDA_FATTN_SPARSE_TILE=1 and at least 8 query rows, GLM-5.3's DSA shape
(DKQ 512, DV 512, ncols2 8) runs ncols1 = 8 instead. The tile's index list is
the union of its rows' cells, so a tile gathers each cell once instead of once
a query.

- The union is correct whatever the overlap.
  flash_attn_ext_f16_load_mask now loops over the tile's rows. Each row reads
  its OWN mask value at every gathered cell, so a cell a row did not select
  reads -inf for it.
- The builder's width is launch_fattn's own ncols1. The list length is
  min(ncols1 * n_kv_max, n_kv), and the list offset is one list a tile. This
  makes a host/kernel width mismatch impossible, so the abort that guarded
  GGML_CUDA_FATTN_SPARSE_ROWS is gone with the env.

The gate is FLASH_ATTN_EXT, 3235/3235 on the RTX 5080 and the 5070 Ti, with
the tile off and on. The branch is taken 4 times, at Q->ne[1] 8, 9 and 16. Two
mutants fail all four cases that reach it, while the other 3231 stay green:
- the builder emitting row 0's cells instead of the union, ERR 0.82-0.87;
- every row reading row 0's mask, ERR 1.14-1.65;
both against the 0.0005 bar.

It is off by default, and no gain is claimed. The saving is the trained
indexer's overlap between adjacent queries' selections: a tile of 8 loads 1.4x
one query's cells at 95 % overlap and 8x at none. A random-weight proxy shows
the no-overlap case by construction.
…: parked for v2

GLM-5.3 v1 (rig's GLM-V1-PLAN.md) stops the query-tiled sparse flash attention. Its
value is the trained indexer's selection overlap, which the proxy cannot measure,
and it would ship switched off. The kernel is finished and verified at fbd07f7:
FLASH_ATTN_EXT 3235/3235 with the tile off and on, on both cards, and two mutants
that fail every case reaching it. This revert keeps it in history and takes it out
of the release. fattn.cu, fattn-common.cuh and fattn-mma-f16.cuh are byte for byte
as at d5cff0d.

Restoring it for v2 is one `git revert` of this commit. The eval cases that reach
batched sparse attention (a78d07e) stay: they check the shipped one-query path,
which had no reference check above 3 queries.
…orm fused into it where its MUL_MAT runs on cuBLAS in BF16, bit for bit

dsv4_hc_post_f32 gave each output a thread: every thread read x and the four residual streams for the one stream it
makes and unflattened its index in 64 bits. dsv4_hc_post_rows_f32 gives a token a block of 1024 threads, n_embd / 1024
columns a thread (n_embd a multiple of 1024 up to 8192), and makes all four streams at a column from one read of x and
the streams; each output is the old kernel's expression on the same values.

At a prefill ubatch the next front's weightless RMS_NORM of the flat streams feeds hc_fn's MUL_MAT on cuBLAS in BF16,
and the norm read the streams back. Where ggml_cuda_dsv4_hc_post_norm_supported and ggml_cuda_mul_mat_runs_cublas_bf16
hold, the post's kernel sums the squares of the values it holds in rms_norm_f32<1024>'s order (column t + 1024 j,
j = idst * NQ + q), reduces them as that kernel does and writes the norm's BF16 copy, which cuBLAS reads. The F32 norm
is neither written nor read. ggml_cuda_mul_mat_cublas_bf16_src1 is the BF16 view both fusions now hand cuBLAS.

No __restrict__ on the kernel's pointers: with it nvcc hoisted x's and a stream's loads (LDG.E.CONSTANT) above the PDL
wait, and DSV4_HC_PRE_POST failed at random (ERR 0.01-0.07).

GGML_CUDA_HC_POST_ROWS_LEGACY=1 runs the thread-an-output kernel and turns the fusion off with it;
GGML_CUDA_HC_POST_NORM_LEGACY=1 runs the norm on its own. test-backend-ops: DSV4_HC_POST gains 1024, 7168 and 8192
(and 9216, past the block kernel); DSV4_HC_PRE_FUSED gains the post with 64 and 257 tokens, where the fusion runs.
): post and norm 346.6 -> 199.4 us a layer, pp4096 +2.62 %, bit for bit
A speculation round at n_max=3 decodes the MTP head four times: once in
process() to carry the verified rows into the head's cache, then once per
draft step. Those rows are only needed as context for the first draft, so
they can ride in that decode instead of one of their own.

process() holds a generating sequence's rows with the target hidden state
each one needs, accept() cuts them to the sampled token plus the accepted
drafts, and draft() prepends what is left to the row the draft starts from.
end() flushes anything still held so a later prompt's cache is complete.
LLAMA_MTP_FOLD_LEGACY=1 restores the separate decode.

Measured on the real glm-5.3-flash pack, one RTX 5070 Ti, draft-mtp n_max 3,
2x320 tokens an arm, seed 1234:

  arm     decode calls  calls/round  acceptance  tok/s
  legacy          1664         4.92      30.1 %  119.8, 123.3
  fold            1472         3.98      24.7 %  108.8, 107.4

The fold engages exactly as predicted, 5 decodes a round down to 4, and is
11 % slower. The first draft decode is n_accepted+2 rows wide, so its shape
changes every round and llama.cpp cannot reuse the graph: rebuilds go from
1.7 % of decodes to 8.8 %. Losing reuse entirely costs 16 % upstream
(ggml-org#20605), while the fold's whole prize is one NextN block,
0.280 GB against the trunk's 11.860 GB a token, or 2.2 % of a round.

The next commit reverts it. Padding the batch to a constant width would
recover that 2.2 %, but the unrolled draft chain builds one constant-width
graph for the whole round and drops the catch-up decode as a side effect,
so the padding work would be thrown away.
This reverts f911f3e, which measured 11 % slower than the separate
catch-up decode on the real glm-5.3-flash pack: its first draft decode
varies in width with the accepted count, which costs more graph reuse
than the one head decode it saves is worth. The numbers and the ceiling
arithmetic are in that commit's message.

One git revert restores it, should the unrolled draft chain not subsume it.
@marcospaulo
marcospaulo merged commit 81aca75 into main Oct 2, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant