Found 2026-10-05 by the GLM decode work (strixhalo, performance mode). GLM-4.7-Flash Q4_K_M on the HRX decode path (-ub 1) at c4096: PPL 11.79 vs CPU 10.60. On the same build with the new decode kernels, the result is 10.64. So the existing HRX MLA decode path (absorbed MLA, dk 576 / dv 512) drifts past about 512 KV cells. Every MLA model that still takes that path is affected (deepseek2-family, GLM-4.7-Flash without the new kernels).
Related, also found there: in a flash-attention PV product, masked keys (P exactly 0) still changed the output by about 1e-6 through their V values, so stale KV rows from earlier requests leaked into decode and identical requests diverged. Other HRX flash-attention kernels may have the same determinism problem.
Next: bisect which op in the old path drifts (decode-split FA at dk 576 vs the MLA absorb matmuls), and add a repeated-request check with random data in masked KV cells.
Found 2026-10-05 by the GLM decode work (strixhalo, performance mode). GLM-4.7-Flash Q4_K_M on the HRX decode path (-ub 1) at c4096: PPL 11.79 vs CPU 10.60. On the same build with the new decode kernels, the result is 10.64. So the existing HRX MLA decode path (absorbed MLA, dk 576 / dv 512) drifts past about 512 KV cells. Every MLA model that still takes that path is affected (deepseek2-family, GLM-4.7-Flash without the new kernels).
Related, also found there: in a flash-attention PV product, masked keys (P exactly 0) still changed the output by about 1e-6 through their V values, so stale KV rows from earlier requests leaked into decode and identical requests diverged. Other HRX flash-attention kernels may have the same determinism problem.
Next: bisect which op in the old path drifts (decode-split FA at dk 576 vs the MLA absorb matmuls), and add a repeated-request check with random data in masked KV cells.