Repository navigation
Eval bug: draft-mtp DeviceLost during *prompt* on AMD RADV — common_speculative_process runs llama_decode(ctx_dft) after every prefill ubatch #27306
Description
Activity
I am seeing a similar issue even with MTP disabled.
On Radeon 8060S / RADV, with MTP disabled, flash attention enabled, and
-b 2048 -ub 2048, Qwen3.8-27B reproducibly resets at the 53248-token ubatch boundary. Reducing both batch and ubatch to 1024 allows the same target model to continue past this region.Here is the apparent failure mechanism:
With
GGML_VK_SERIALIZE_SUBMISSIONS=1, llama.cpp isolates a single guilty submission:device lost on Vulkan0, likely caused by previous submission (nodes 3635 to 3635)GGML_VK_SYNC_LOGGER=1maps node 3635 to the final full-attention layer'sFLASH_ATTN_EXToperation (blk.63).At that point, the FA node represents approximately:
2 * 2048 query rows * 24 heads * (256 K + 256 V) * 53248 KV ~= 2.68 TFLOPllama.cpp's Vulkan scheduler has an approximately 200-GFLOP upper cap per submission, but it can currently split only between graph nodes. Consequently,
GGML_VK_MAX_NODES_PER_SUBMIT=1cannot help when a single FA node is itself too large.The evidence indicates that this unsplit FA dispatch exceeds AMDGPU's compute-job watchdog, causing the compute ring to be timed out and reset. After booting with:
amdgpu.lockup_timeout=10000,60000,10000,10000the otherwise identical unsplit workload crossed 53248 tokens and completed a 60000-token prompt successfully. This strongly supports a watchdog timeout.
I also made an AI-assisted prototype patch that divides FA query rows across multiple actual Vulkan submissions, using the existing FLOP cap. Without increasing
amdgpu.lockup_timeout, it completed:- 60K tokens with MTP disabled and
-b 2048 -ub 2048 - 60K tokens with MTP enabled and
-b 8192 -ub 2048
These results suggest that graph-level submission splitting is insufficient for large long-context FA nodes; the operation itself needs to be divided into bounded submissions.
- 60K tokens with MTP disabled and
Note: The text of this post was structured and drafted with the help of an AI assistant (Claude), based exclusively on my own measurements and technical data.
Confirming this exact mechanism produces a second, distinct symptom on a different backend: silent content divergence at temp=0, not a crash.
Setup: Qwen3.8-27B (
qwen35,nextn_predict_layers=1), self-speculative--spec-type draft-mtp(no-md), mainline build offdfa0c0fee(~b10591).Environment: Apple M5 Pro, 64 GB unified memory, macOS 26.5.2 (25F84), Metal backend.
At temp=0, greedy decode is bit-reproducible run-to-run without spec (verified: two separate
--spec-type noneruns, identical output). With--spec-type draft-mtpactive, output diverges from baseline mid-generation, same prompt/seed. Direct logprob comparison (n_probsviallama-server) at the divergence point: baseline top token logprob -0.630, second candidate -0.762 (gap 0.132 nats, ~1.14x). MTP selects the second candidate, not something outside the distribution. Small perturbation flipping a close decision, consistent with!is_mem_sharedrunning an extrallama_decode(ctx_dft)per ubatch and perturbing the target's own logits, not a broken accept/reject.Not quantization-specific: reproduced on both Q4_K_M and Q5_K_M (Q4 diverges on the first prompt tried, Q5 on 2/10 across a 10-prompt battery, same failure shape both times). Consistent with the in-tree TODO above the call site (
[TAG_SPEC_AVOID_DRAFT_REEVAL], "for now, always re-evaluate for simplicity"), matches what you'd expect from that code path being unfinished, not a Vulkan-specific issue.Happy to share the exact repro script/prompts if useful.
CUDA data point: the same
process()-per-ubatch path costs ~38% throughput on a backend that never crashesAdding a third symptom class to this issue. On CUDA the extra
llama_decode(ctx_dft)does not
produce a DeviceLost (Vulkan/RADV) or a silent content divergence (Metal, @ovidiu-morar above) —
it just costs a lot of throughput, quietly.Setup
- 2× NVIDIA RTX PRO Blackwell (32.6 GB + 16.3 GB), 121 GB DDR5, CUDA toolkit 12.8.93, driver 610.43.02
qwen4exp(Qwen3.8-Flash-Next), our own Q4_K_M conversion, 130.7 GiB,nextn_predict_layers = 1- Disclosure: this is a fork of llama.cpp. For these two runs the fork's own offload path is
compiled out — experts are placed with plain-otonto the two devices and CPU, so the data
path is upstream's. The fork is based on the model: add Qwen3.8-Flash-Next (qwen4exp) #27742 merge, socommon_speculative_processis
the code discussed here. --parallel 1,-fa off, KV f16,-c 8192, static expert split
(blk.0-14 → CUDA0,blk.15-21 → CUDA1, rest CPU,per_layer_token_embd=CPU)
A/B — same harness, same four rotating prompts,
temperature 0, 4 measured rounds each after
a discarded warm-up, server restarted between arms so the only change is--spec-type. Output text
was checked at 22 / 414 / 986 / 2584 prompt tokens in both arms and was correct in both.arm --spec-typemin / median / max t/s spread draft n/accepted A none26.85 / 26.93 / 27.09 0.9 % – B draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.515.82 / 16.71 / 17.53 10.2 % 116 / 43 = 37 % Speculation costs 38 % of throughput here, and multiplies run-to-run spread by 11.
Why we think this is worth adding rather than being just a slow config
The failure is invisible if you only look at t/s inside one arm. We spent two days treating 17 t/s
as the model's speed and looking for the bottleneck in the backbone — attention, expert placement,
transfer granularity — because the number was stable and the output was correct. Only a same-harness
A/B against--spec-type noneshowed the drafter itself was the cost. The 0.9 % vs 10.2 % spread
is the other tell: arm B's timing is dominated by something whose cost varies per ubatch.That matches
[TAG_SPEC_AVOID_DRAFT_REEVAL]directly: for!is_mem_sharedevery target prefill
ubatch is followed by a second, equally widellama_decode(ctx_dft)through the whole MTP block.
At 37 % acceptance the draft does not earn back a full extra pass.We have not tried arm C from the opening post (skip
process()unless a slot is
SLOT_STATE_GENERATING) yet. If a CUDA measurement of that arm is useful for sizing the fix, we can
run the same A/B with it and report.One methodological note that may save someone else the two days. During the same period this
model produced correct-looking, fluent output while a separate expert-offload bug corrupted it above
~296 prompt tokens, and generation speed stayed within 23–24 t/s in both states — the draft
overhead roughly cancelled the lost acceptance. A throughput-only test cannot see either problem.
We now run a fixed text ladder at four prompt lengths before every measurement and refuse to record
a number if any length degrades.Correction to my comment above — two things I got wrong, one of which weakens a number I quoted.
1. The acceptance figure and the throughput figure are not from the same request.
I put
draft 116 / 43 = 37 %in the same table row as the throughput, which reads as if both
describe the same measurement. They do not. In our harness the throughput is the median of four
rounds (rotating ~25–30 token prompts, 1200 tokens generated each), whiledraft_n/
draft_n_acceptedare read from one separate 9-token request withn_predict 200, issued
after those rounds.So
37 %is a single short-prompt sample, not the acceptance during the measured rounds. It is
indicative at best and should not be read as characterising arm B. I should have either measured
it in the same requests or left it out.The throughput A/B itself is unaffected — both arms ran the identical four rounds, same
prompts, same seed conditions, server restarted between arms:arm --spec-typemin / median / max t/s spread A none26.85 / 26.93 / 27.09 0.9 % B draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.515.82 / 16.71 / 17.53 10.2 % 2. "the fork's own offload path is compiled out" is wrong. It is compiled in; it is simply not
selected, because the-otexpression maps experts toCUDA0/CUDA1/CPUrather than to the
fork's own buffer type. No fork-specific code runs in the expert path for these two runs, but the
reason is placement, not compilation. The distinction matters if anyone tries to reproduce.One further disclosure while I am at it, since I checked it after posting: our fork carries an
optional adaptive-speculation gate (FLE_SPEC_ADAPTIVE, skips draft rounds after consecutive
zero-accept rounds). It is off unless the environment variable is set, and it was not set for
either arm — verified in the run logs, zero occurrences. So it did not influence the A/B.Apologies for the noise; I would rather correct the record than leave a number standing that looks
better sourced than it is.Maybe this will help
Same failure also. Strix Halo / RADV, later llama.cpp build, heavily cached ~91K context
Environment:
- Ryzen AI Max+ 395
- Radeon 8060S / RADV STRIX_HALO (
gfx1151) - 128 GB unified RAM
- Ubuntu Server
- kernel
7.0.0-30-generic - Mesa/RADV
26.0.8-1ubuntu0.3 - llama.cpp build
10419, commitaee56b3ab - Qwen3.8-27B UD-Q8_K_L (Unsloth)
- Vulkan
- --spec-type draft-mtp
- -c 262144
- one slot
Slightly different from the large unique-prefix prefill in the original report. This crashed during a already-long, heavy cached conversation.
Last successful response before the crash was:
prompt_tokens: 90951 cached_tokens: 86222 prompt_n: 4729 completion_tokens: 2852 prompt_per_second: 19.8571 predicted_per_second: 12.0972 draft_n: 3213 draft_n_accepted: 1782About 94.8% of the ~91K-token prompt was already cached and only 4,729 prompt tokens were newly evaluated.
Next turn was relatively small and died with
decode() failed: vk::Queue::submit: ErrorDeviceLostKernel showed the same compute-ring hang:
Dumping IP State AMDGPU device coredump file has been created ring comp_1.1.0 timeout Process llama-server reset compute queue (1:1:0) Ring comp_1.1.0 reset failed GPU reset begin! MODE2 reset GPU reset succeeded, trying to resume GPU reset(1) succeeded! device wedged, but recovered through resetRestarted the server/proxy only, no config changes, went back to the same conversation, and used Regenerate on the same message. The retry ran for 2050.52 s and then failed again with
decode() failed: vk::Queue::submit: ErrorDeviceLostSecond kernel failure:
Dumping IP State AMDGPU device coredump file has been created ring comp_1.1.0 timeout Process llama-server reset compute queue (1:1:0) Ring comp_1.1.0 reset succeeded device wedged, but recovered through resetThis time the ring reset succeeded, and it didn't escalate to a MODE2 reset.
There were also two earlier warnings during the second request:
Fence fallback timer expired on ring comp_1.2.0 Fence fallback timer expired on ring comp_1.2.0They occurred roughly 16 min. and 10 min. before the final comp_1.1.0 timeout. FWIW
Host memory was fine during the retry: roughly 52 GB of 124 GB in use, 0% swap.
No further attempts or changes to MTP/context/Vulkan/watchdog settings.
Summary
On AMD gfx1151 / Vulkan / RADV,
llama-serverwith--spec-type draft-mtpdies mid-prompt (decode() failed: vk::Queue::submit: ErrorDeviceLost) at tens of thousands of tokens. The same argv with MTP off survives past 125k. A one-site diagnostic that keeps MTP flags on but skipscommon_speculative_processwhile no slot isSLOT_STATE_GENERATINGalso survives past 125k on both 1-seat and 4-seat topologies; generate still drafts.This is not a checkpoint /
get_tensordeath (--ctx-checkpoints 0, this-bootcreate_checkpoint=0get_tensor=0update_tgt=0). It is not “Vulkan FA at 114k without MTP” on one seat (MTP-off 1×262144 finished 125948). It is the extra draft-context decode thatimpl_draft_mtp::processruns after every target ubatch on!is_mem_sharedQwen3.8.GGML_VK_MAX_NODES_PER_SUBMIT=1is already set (the UMA submit-frequency workaround from #24872 / the turboquant fork). It is not enough.Name and Version
/propsbuild_info=b10352-4dee52f82.Operating systems
Linux (AMD Strix Halo box).
GGML backends
Vulkan (RADV).
Hardware
gfx1151), 128 GB unified LPDDR5XGGML_VK_MAX_NODES_PER_SUBMIT=1,queue_preemption_timeout_ms=60000Models
Qwen3.8-27BUD-Q8_K_XL +mmproj-F16(Unsloth)qwen35-family): most layers recurrent, a minority full-attention. Qwen3.8 MTP is!is_mem_shared, so the draft context is a separatellama_decode(ctx_dft, same n_tokens)after the target ubatch.Problem description
With
--spec-type draft-mtp --spec-draft-n-max 2, long unique-prefix prefill dies invk::Queue::submitduring prompt processing:dmesg is a compute-ring timeout then reset (
ring comp_1.2.0/comp_1.0.1), process survives as a zombie (/v1/models200, later/completion500).Call site (b10352
server-context.cppdecode(), thencommon/speculative.cppimpl_draft_mtp::process): after every successful targetllama_decode(ctx_tgt, ubatch),common_speculative_processruns. On this model that is a second same-width draft decode.--ubatch-size 1024⇒ a 1024-wide draft graph after every 1024-wide target graph, at risingpos. The outerdecode()catch does not say which of the twollama_decodes threw; the A/B below isolates the extra one.There is an in-tree TODO immediately above the call:
A/B (same box, same GGUF, same day)
Shared flags on every arm unless noted:
Hop =
/completionunique prefix, tokenized ~125948,n_predict=1.-c/-np/ slot)--spec-type)get_tensor=0--spec-type draft-mtp --spec-draft-n-max 2decode() / vk::Queue::submit, this-bootcreate_check=0get_tensor=0update_tgt=0SLOT_STATE_GENERATINGdraft_n=6accepted 2-c 1048576 -np 4)draft_n=6accepted 2Production 4-seat MTP-on stock process-every-ubatch (same argv as D, unpatched binary) died three times in the 114–117k band (
n_tokens=114025,114688,116736), same raise, same zero copy counters. Do not collapse the 1-seat 55k death into those 4-seat 114k deaths — same class, different CONFIG. Arm D is the 4-seat isolation: skip the extra decode and the 114–117k band is crossed.After C/D hop 1,
begin()warnedctx_dft pos_max=-1(the skip did what it said). Hop 2 was a short unique prompt,n_predict=32: generate still ran draft-mtp (draft_n=6, accepted 2). Accept is degraded vs a filled draft KV — expected with an emptyctx_dft, not a reason to drop--spec-type draft-mtp.Minimal diagnostic patch (not proposed as the final ship)
In
tools/server/server-context.cpp, around the existingcommon_speculative_process(spec.get(), batch_view)after the target ubatch:That is what arms C and D ran. A production-shaped fix is probably tail-only / last-ubatch process during prompt (so
pending_h/ draft KV is warm for generate) rather than skip-all. Windowed / smaller draft ubatch is another lever.--spec-type noneis a diagnostic, not a fix.What this is not
--ctx-checkpoints/create_checkpoint/vk_buffer_get_tensor. Those are a separate AMD DeviceLost class (54–62k on this box when checkpoints are armed). This repro has checkpoints off and the copy counters stay 0.GGML_VK_MAX_NODES_PER_SUBMIT=1is on.--ubatch-sizeas the product fix. Those are known timeout knobs (Eval bug: AMD Vulkan vk::DeviceLostError crash sensitive to ubatch-size and context length #20515 / Vulkan: vk::DeviceLostError crash in ggml_vk_flash_attn on AMD gfx1102 (RADV PHOENIX) #20889) and they trade prefill speed.Related (read; none of these is this pin)
impl_draft_mtp::process→llama_decode(ctx_dft)during prefill, MTP-off survives.nodes_per_submit=1/ validated to 70k. That knob is already on here; 1-seat MTP-on still dies at 55k. Never filed on ggml-org/llama.cpp.GGML_VK_MAX_NODES_PER_SUBMIT=1).common_speculative_process.--spec-type draft-mtpDeviceLost on RADV (RX 6900 XT); also reported on Strix Halo.create_checkpoint/update_tgt. No process-during-prefill A/B.--spec-typein the argv. Model list happens to include an MTP GGUF.draft-mtpcrash under KV saturation.cublasSgemm, not Vulkan.Expected
--spec-type draft-mtpcan stay on for generate without a second full-widthllama_decode(ctx_dft)after every prompt ubatch on a backend whose single-submit budget cannot swallow that extra graph at depth.First Bad Commit
Unknown. Reproduced on
4dee52f82(b10352). Not bisected.