Split off from engine#123. The HRX0,Vulkan0 -ts 1,0 -ot exps=Vulkan0 split fails above ~4750 tokens, and it is not the decode-split multipass.
Repro (gfx1151, pinned f5b7f4ad, Qwen3-Coder-30B-A3B, -ngl 99 -fa 1 -c 8192)
A ~4800-token prompt (a 86-section synthetic doc):
-dev HRX0,Vulkan0 -ts 1,0 -ot exps=Vulkan0: at ~2128/2512/3000 tokens the code word ZX-4718-QQ is exact (18/18); at ~4800 it fails.
-dev HRX0,Vulkan0 (no -ot exps=Vulkan0): same failure at ~4800.
-dev HRX0 alone: 0 faults, but the code word comes back wrong at ~4800 (2/3), while at ~4704 it is exact.
-dev Vulkan0 alone: clean.
Recorded with the env + dispatch per condition in strixhalo:~/goal-verify/split-inspect/{mp-on,mp-off}.{env,srv} and ~/goal-verify/sv-*.srv: with flash_attention_decode_split enabled the dispatch is flash_attention_decode_split_next_q8; with GGML_HRX_DISABLE_DISPATCH=flash_attention_decode_split the dispatch is flash_attention_f32_f16_wmma only — i.e. no multipass runs — yet the failure is the same. So the cause is upstream of / independent of the decode-split (a context-size, cross-device, or server-side issue).
Evidence
strixhalo:~/goal-verify/cov2/split-depths.txt, cov2/split.server
strixhalo:~/goal-verify/cor responses -> ~/goal-verify/corr-{mp,fb}.srv (a server_context_impl::post_decode/update_slots abort)
llama-bench cannot run this split at all: both split benchmarks abort with ggml_backend_sched: pre-allocated tensor (blk.0.ffn_down_exps.weight) in a buffer (Vulkan0) that cannot run the operation (NONE).
Not part of #123's multipass fix. Filed so #123 can close on its own defect.
Split off from engine#123. The
HRX0,Vulkan0 -ts 1,0 -ot exps=Vulkan0split fails above ~4750 tokens, and it is not the decode-split multipass.Repro (gfx1151, pinned f5b7f4ad, Qwen3-Coder-30B-A3B, -ngl 99 -fa 1 -c 8192)
A ~4800-token prompt (a 86-section synthetic doc):
-dev HRX0,Vulkan0 -ts 1,0 -ot exps=Vulkan0: at ~2128/2512/3000 tokens the code wordZX-4718-QQis exact (18/18); at ~4800 it fails.-dev HRX0,Vulkan0(no-ot exps=Vulkan0): same failure at ~4800.-dev HRX0alone: 0 faults, but the code word comes back wrong at ~4800 (2/3), while at ~4704 it is exact.-dev Vulkan0alone: clean.Recorded with the env + dispatch per condition in
strixhalo:~/goal-verify/split-inspect/{mp-on,mp-off}.{env,srv}and~/goal-verify/sv-*.srv: withflash_attention_decode_splitenabled the dispatch isflash_attention_decode_split_next_q8; withGGML_HRX_DISABLE_DISPATCH=flash_attention_decode_splitthe dispatch isflash_attention_f32_f16_wmmaonly — i.e. no multipass runs — yet the failure is the same. So the cause is upstream of / independent of the decode-split (a context-size, cross-device, or server-side issue).Evidence
strixhalo:~/goal-verify/cov2/split-depths.txt,cov2/split.serverstrixhalo:~/goal-verify/cor responses->~/goal-verify/corr-{mp,fb}.srv(aserver_context_impl::post_decode/update_slotsabort)llama-benchcannot run this split at all: both split benchmarks abort withggml_backend_sched:pre-allocated tensor (blk.0.ffn_down_exps.weight) in a buffer (Vulkan0) that cannot run the operation (NONE).Not part of #123's multipass fix. Filed so #123 can close on its own defect.