Skip to content

Eval bug: draft-mtp DeviceLost during *prompt* on AMD RADV — common_speculative_process runs llama_decode(ctx_dft) after every prefill ubatch #27306

Description

@Biggles10-claude

Summary

On AMD gfx1151 / Vulkan / RADV, llama-server with --spec-type draft-mtp dies mid-prompt (decode() failed: vk::Queue::submit: ErrorDeviceLost) at tens of thousands of tokens. The same argv with MTP off survives past 125k. A one-site diagnostic that keeps MTP flags on but skips common_speculative_process while no slot is SLOT_STATE_GENERATING also survives past 125k on both 1-seat and 4-seat topologies; generate still drafts.

This is not a checkpoint / get_tensor death (--ctx-checkpoints 0, this-boot create_checkpoint=0 get_tensor=0 update_tgt=0). It is not “Vulkan FA at 114k without MTP” on one seat (MTP-off 1×262144 finished 125948). It is the extra draft-context decode that impl_draft_mtp::process runs after every target ubatch on !is_mem_shared Qwen3.8.

GGML_VK_MAX_NODES_PER_SUBMIT=1 is already set (the UMA submit-frequency workaround from #24872 / the turboquant fork). It is not enough.

Name and Version

b10352-4dee52f82
Vulkan backend, llama-server

/props build_info = b10352-4dee52f82.

Operating systems

Linux (AMD Strix Halo box).

GGML backends

Vulkan (RADV).

Hardware

  • AMD Ryzen AI Max+ 395 / Radeon 8060S (gfx1151), 128 GB unified LPDDR5X
  • GTT pool 124 GiB (UMA frame buffer ~2 GiB)
  • RADV / Mesa last recorded 2026-08-16: Mesa 25.2.8 (not re-probed on the day of this write-up)
  • Env on every arm: GGML_VK_MAX_NODES_PER_SUBMIT=1, queue_preemption_timeout_ms=60000

Models

  • Qwen3.8-27B UD-Q8_K_XL + mmproj-F16 (Unsloth)
  • Hybrid (qwen35-family): most layers recurrent, a minority full-attention. Qwen3.8 MTP is !is_mem_shared, so the draft context is a separate llama_decode(ctx_dft, same n_tokens) after the target ubatch.

Problem description

With --spec-type draft-mtp --spec-draft-n-max 2, long unique-prefix prefill dies in vk::Queue::submit during prompt processing:

radv/amdgpu: The CS has been cancelled because the context is lost. This context is innocent.
ggml_vulkan: device lost on Vulkan0
srv  update_slots: decode() failed: vk::Queue::submit: ErrorDeviceLost

dmesg is a compute-ring timeout then reset (ring comp_1.2.0 / comp_1.0.1), process survives as a zombie (/v1/models 200, later /completion 500).

Call site (b10352 server-context.cpp decode(), then common/speculative.cpp impl_draft_mtp::process): after every successful target llama_decode(ctx_tgt, ubatch), common_speculative_process runs. On this model that is a second same-width draft decode. --ubatch-size 1024 ⇒ a 1024-wide draft graph after every 1024-wide target graph, at rising pos. The outer decode() catch does not say which of the two llama_decodes threw; the A/B below isolates the extra one.

There is an in-tree TODO immediately above the call:

// TODO: avoid restoring the draft context and re-evaluating the drafted tokens when not needed [TAG_SPEC_AVOID_DRAFT_REEVAL]
//       for now, always re-evaluate for simplicity

A/B (same box, same GGUF, same day)

Shared flags on every arm unless noted:

-ngl 999 --flash-attn on --ubatch-size 1024 --batch-size 2048
--cache-type-k q4_0 --cache-type-v q4_0
--cache-prompt --cache-reuse 256 --cache-ram 0 --ctx-checkpoints 0
--image-min-tokens 1024 --threads 16 --jinja
GGML_VK_MAX_NODES_PER_SUBMIT=1

Hop = /completion unique prefix, tokenized ~125948, n_predict=1.

Arm MTP seats (-c / -np / slot) process() during prompt last healthy n result
A off (no --spec-type) 1 × 262144 n/a 125948 HTTP 200, 992.6 s, 126.90 t/s, DeviceLost=0, get_tensor=0
B on --spec-type draft-mtp --spec-draft-n-max 2 1 × 262144 stock (every ubatch) 55296 / 0.44 HTTP 500 at t=6:16.751, decode() / vk::Queue::submit, this-boot create_check=0 get_tensor=0 update_tgt=0
C on n-max 2 1 × 262144 skipped unless a slot is SLOT_STATE_GENERATING 125948 HTTP 200, 993.2 s, 126.81 t/s, DeviceLost=0. Skip log counts 1/16/32/48. Hop 2: draft_n=6 accepted 2
D on n-max 2 4 × 262144 (-c 1048576 -np 4) skipped as C 125949 HTTP 200, 994.6 s, 126.64 t/s, DeviceLost=0. Same skip counts. Hop 2: draft_n=6 accepted 2

Production 4-seat MTP-on stock process-every-ubatch (same argv as D, unpatched binary) died three times in the 114–117k band (n_tokens=114025, 114688, 116736), same raise, same zero copy counters. Do not collapse the 1-seat 55k death into those 4-seat 114k deaths — same class, different CONFIG. Arm D is the 4-seat isolation: skip the extra decode and the 114–117k band is crossed.

After C/D hop 1, begin() warned ctx_dft pos_max=-1 (the skip did what it said). Hop 2 was a short unique prompt, n_predict=32: generate still ran draft-mtp (draft_n=6, accepted 2). Accept is degraded vs a filled draft KV — expected with an empty ctx_dft, not a reason to drop --spec-type draft-mtp.

Minimal diagnostic patch (not proposed as the final ship)

In tools/server/server-context.cpp, around the existing common_speculative_process(spec.get(), batch_view) after the target ubatch:

const bool any_generating = std::any_of(slots.begin(), slots.end(), [](const server_slot & s) {
    return s.state == SLOT_STATE_GENERATING;
});
if (any_generating) {
    if (!common_speculative_process(spec.get(), batch_view)) {
        throw std::runtime_error("failed to process speculative batch");
    }
}
// else: skip draft-context fill during prompt. MTP flags stay on.

That is what arms C and D ran. A production-shaped fix is probably tail-only / last-ubatch process during prompt (so pending_h / draft KV is warm for generate) rather than skip-all. Windowed / smaller draft ubatch is another lever. --spec-type none is a diagnostic, not a fix.

What this is not

Related (read; none of these is this pin)

Ticket State Why it is related Why it is not this issue
TheTom/llama-cpp-turboquant#185 closed (fork) Same chip, same impl_draft_mtp::process → llama_decode(ctx_dft) during prefill, MTP-off survives. Closed as fixed by UMA nodes_per_submit=1 / validated to 70k. That knob is already on here; 1-seat MTP-on still dies at 55k. Never filed on ggml-org/llama.cpp.
#24872 / #21724 merged / closed UMA submit-frequency vs amdgpu lockup_timeout. The raise is the same class (ring timeout → DeviceLost). Does not mention MTP draft-fill. Already applied (GGML_VK_MAX_NODES_PER_SUBMIT=1).
#20515 closed stale Same chip (8060S / gfx1151), DeviceLost vs ubatch × depth. No MTP A/B. No common_speculative_process.
#20889 closed AMD RADV FA DeviceLost (gfx1102). Different GPU; FA-off workaround.
#24623 closed stale Gemma 4 MTP + DeviceLost on AMD 680M Vulkan; MTP-off works. Generate-time crash on short prompts; no skip-during-prompt isolation.
#27076 open Qwen3.8-27B + --spec-type draft-mtp DeviceLost on RADV (RX 6900 XT); also reported on Strix Halo. Next-turn after a 29k generate; second abort is create_checkpoint / update_tgt. No process-during-prefill A/B.
#26447 open Vega 8 / other AMD iGPU DeviceLost ~50k; timeout workaround. No --spec-type in the argv. Model list happens to include an MTP GGUF.
#25664 open gfx1151 DeviceLost on linux-7.x; DS4 fused-LID fallback / 2s lockup_timeout. Different model / op. Comments mention Qwen3.8-27B as another timeout, not this pin.
#23286 closed, not merged Would skip draft generation during prefill. Behaviour matrix still runs process/sync on text-only prompt prefill. That is the call that kills us. Closed after author could not repro #22867.
#26558 open draft-mtp crash under KV saturation. CUDA cublasSgemm, not Vulkan.
discussion #27154 open gfx1151 Qwen3.8-27B; a commenter DeviceLosts on Vulkan at ~50% of 262k. Throughput thread. No MTP-prefill isolation.

Expected

--spec-type draft-mtp can stay on for generate without a second full-width llama_decode(ctx_dft) after every prompt ubatch on a backend whose single-submit budget cannot swallow that extra graph at depth.

First Bad Commit

Unknown. Reproduced on 4dee52f82 (b10352). Not bisected.

Activity

  1. olympichek commented on Aug 22, 2026

    @olympichek

    I am seeing a similar issue even with MTP disabled.

    On Radeon 8060S / RADV, with MTP disabled, flash attention enabled, and -b 2048 -ub 2048, Qwen3.8-27B reproducibly resets at the 53248-token ubatch boundary. Reducing both batch and ubatch to 1024 allows the same target model to continue past this region.

    Here is the apparent failure mechanism:

    With GGML_VK_SERIALIZE_SUBMISSIONS=1, llama.cpp isolates a single guilty submission:

    device lost on Vulkan0, likely caused by previous submission
    (nodes 3635 to 3635)
    

    GGML_VK_SYNC_LOGGER=1 maps node 3635 to the final full-attention layer's FLASH_ATTN_EXT operation (blk.63).

    At that point, the FA node represents approximately:

    2 * 2048 query rows * 24 heads * (256 K + 256 V) * 53248 KV ~= 2.68 TFLOP
    

    llama.cpp's Vulkan scheduler has an approximately 200-GFLOP upper cap per submission, but it can currently split only between graph nodes. Consequently, GGML_VK_MAX_NODES_PER_SUBMIT=1 cannot help when a single FA node is itself too large.

    The evidence indicates that this unsplit FA dispatch exceeds AMDGPU's compute-job watchdog, causing the compute ring to be timed out and reset. After booting with:

    amdgpu.lockup_timeout=10000,60000,10000,10000
    

    the otherwise identical unsplit workload crossed 53248 tokens and completed a 60000-token prompt successfully. This strongly supports a watchdog timeout.

    I also made an AI-assisted prototype patch that divides FA query rows across multiple actual Vulkan submissions, using the existing FLOP cap. Without increasing amdgpu.lockup_timeout, it completed:

    • 60K tokens with MTP disabled and -b 2048 -ub 2048
    • 60K tokens with MTP enabled and -b 8192 -ub 2048

    These results suggest that graph-level submission splitting is insufficient for large long-context FA nodes; the operation itself needs to be divided into bounded submissions.

  2. ovidiu-morar commented on Aug 27, 2026

    @ovidiu-morar

    Note: The text of this post was structured and drafted with the help of an AI assistant (Claude), based exclusively on my own measurements and technical data.

    Confirming this exact mechanism produces a second, distinct symptom on a different backend: silent content divergence at temp=0, not a crash.

    Setup: Qwen3.8-27B (qwen35, nextn_predict_layers=1), self-speculative --spec-type draft-mtp (no -md), mainline build off dfa0c0fee (~b10591).

    Environment: Apple M5 Pro, 64 GB unified memory, macOS 26.5.2 (25F84), Metal backend.

    At temp=0, greedy decode is bit-reproducible run-to-run without spec (verified: two separate --spec-type none runs, identical output). With --spec-type draft-mtp active, output diverges from baseline mid-generation, same prompt/seed. Direct logprob comparison (n_probs via llama-server) at the divergence point: baseline top token logprob -0.630, second candidate -0.762 (gap 0.132 nats, ~1.14x). MTP selects the second candidate, not something outside the distribution. Small perturbation flipping a close decision, consistent with !is_mem_shared running an extra llama_decode(ctx_dft) per ubatch and perturbing the target's own logits, not a broken accept/reject.

    Not quantization-specific: reproduced on both Q4_K_M and Q5_K_M (Q4 diverges on the first prompt tried, Q5 on 2/10 across a 10-prompt battery, same failure shape both times). Consistent with the in-tree TODO above the call site ([TAG_SPEC_AVOID_DRAFT_REEVAL], "for now, always re-evaluate for simplicity"), matches what you'd expect from that code path being unfinished, not a Vulkan-specific issue.

    Happy to share the exact repro script/prompts if useful.

  3. sirfyyn commented on Aug 27, 2026

    @sirfyyn

    CUDA data point: the same process()-per-ubatch path costs ~38% throughput on a backend that never crashes

    Adding a third symptom class to this issue. On CUDA the extra llama_decode(ctx_dft) does not
    produce a DeviceLost (Vulkan/RADV) or a silent content divergence (Metal, @ovidiu-morar above) —
    it just costs a lot of throughput, quietly.

    Setup

    • 2× NVIDIA RTX PRO Blackwell (32.6 GB + 16.3 GB), 121 GB DDR5, CUDA toolkit 12.8.93, driver 610.43.02
    • qwen4exp (Qwen3.8-Flash-Next), our own Q4_K_M conversion, 130.7 GiB, nextn_predict_layers = 1
    • Disclosure: this is a fork of llama.cpp. For these two runs the fork's own offload path is
      compiled out — experts are placed with plain -ot onto the two devices and CPU, so the data
      path is upstream's. The fork is based on the model: add Qwen3.8-Flash-Next (qwen4exp) #27742 merge, so common_speculative_process is
      the code discussed here.
    • --parallel 1, -fa off, KV f16, -c 8192, static expert split
      (blk.0-14 → CUDA0, blk.15-21 → CUDA1, rest CPU, per_layer_token_embd=CPU)

    A/B — same harness, same four rotating prompts, temperature 0, 4 measured rounds each after
    a discarded warm-up, server restarted between arms so the only change is --spec-type. Output text
    was checked at 22 / 414 / 986 / 2584 prompt tokens in both arms and was correct in both.

    arm --spec-type min / median / max t/s spread draft n/accepted
    A none 26.85 / 26.93 / 27.09 0.9 % –
    B draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.5 15.82 / 16.71 / 17.53 10.2 % 116 / 43 = 37 %

    Speculation costs 38 % of throughput here, and multiplies run-to-run spread by 11.

    Why we think this is worth adding rather than being just a slow config

    The failure is invisible if you only look at t/s inside one arm. We spent two days treating 17 t/s
    as the model's speed and looking for the bottleneck in the backbone — attention, expert placement,
    transfer granularity — because the number was stable and the output was correct. Only a same-harness
    A/B against --spec-type none showed the drafter itself was the cost. The 0.9 % vs 10.2 % spread
    is the other tell: arm B's timing is dominated by something whose cost varies per ubatch.

    That matches [TAG_SPEC_AVOID_DRAFT_REEVAL] directly: for !is_mem_shared every target prefill
    ubatch is followed by a second, equally wide llama_decode(ctx_dft) through the whole MTP block.
    At 37 % acceptance the draft does not earn back a full extra pass.

    We have not tried arm C from the opening post (skip process() unless a slot is
    SLOT_STATE_GENERATING) yet. If a CUDA measurement of that arm is useful for sizing the fix, we can
    run the same A/B with it and report.

    One methodological note that may save someone else the two days. During the same period this
    model produced correct-looking, fluent output while a separate expert-offload bug corrupted it above
    ~296 prompt tokens, and generation speed stayed within 23–24 t/s in both states — the draft
    overhead roughly cancelled the lost acceptance. A throughput-only test cannot see either problem.
    We now run a fixed text ladder at four prompt lengths before every measurement and refuse to record
    a number if any length degrades.

  4. sirfyyn commented on Aug 27, 2026

    @sirfyyn

    Correction to my comment above — two things I got wrong, one of which weakens a number I quoted.

    1. The acceptance figure and the throughput figure are not from the same request.

    I put draft 116 / 43 = 37 % in the same table row as the throughput, which reads as if both
    describe the same measurement. They do not. In our harness the throughput is the median of four
    rounds (rotating ~25–30 token prompts, 1200 tokens generated each), while draft_n /
    draft_n_accepted are read from one separate 9-token request with n_predict 200, issued
    after those rounds.

    So 37 % is a single short-prompt sample, not the acceptance during the measured rounds. It is
    indicative at best and should not be read as characterising arm B. I should have either measured
    it in the same requests or left it out.

    The throughput A/B itself is unaffected — both arms ran the identical four rounds, same
    prompts, same seed conditions, server restarted between arms:

    arm --spec-type min / median / max t/s spread
    A none 26.85 / 26.93 / 27.09 0.9 %
    B draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.5 15.82 / 16.71 / 17.53 10.2 %

    2. "the fork's own offload path is compiled out" is wrong. It is compiled in; it is simply not
    selected, because the -ot expression maps experts to CUDA0 / CUDA1 / CPU rather than to the
    fork's own buffer type. No fork-specific code runs in the expert path for these two runs, but the
    reason is placement, not compilation. The distinction matters if anyone tries to reproduce.

    One further disclosure while I am at it, since I checked it after posting: our fork carries an
    optional adaptive-speculation gate (FLE_SPEC_ADAPTIVE, skips draft rounds after consecutive
    zero-accept rounds). It is off unless the environment variable is set, and it was not set for
    either arm — verified in the run logs, zero occurrences. So it did not influence the A/B.

    Apologies for the noise; I would rather correct the record than leave a number standing that looks
    better sourced than it is.

  5. Mathajas commented on Aug 29, 2026

    @Mathajas
  6. Wrigleysc commented on Sep 8, 2026

    @Wrigleysc

    Same failure also. Strix Halo / RADV, later llama.cpp build, heavily cached ~91K context

    Environment:

    • Ryzen AI Max+ 395
    • Radeon 8060S / RADV STRIX_HALO (gfx1151)
    • 128 GB unified RAM
    • Ubuntu Server
    • kernel 7.0.0-30-generic
    • Mesa/RADV 26.0.8-1ubuntu0.3
    • llama.cpp build 10419, commit aee56b3ab
    • Qwen3.8-27B UD-Q8_K_L (Unsloth)
    • Vulkan
    • --spec-type draft-mtp
    • -c 262144
    • one slot

    Slightly different from the large unique-prefix prefill in the original report. This crashed during a already-long, heavy cached conversation.

    Last successful response before the crash was:

    prompt_tokens: 90951
    cached_tokens: 86222
    prompt_n: 4729
    completion_tokens: 2852
    
    prompt_per_second: 19.8571
    predicted_per_second: 12.0972
    
    draft_n: 3213
    draft_n_accepted: 1782
    

    About 94.8% of the ~91K-token prompt was already cached and only 4,729 prompt tokens were newly evaluated.

    Next turn was relatively small and died with decode() failed: vk::Queue::submit: ErrorDeviceLost

    Kernel showed the same compute-ring hang:

    Dumping IP State
    AMDGPU device coredump file has been created
    ring comp_1.1.0 timeout
    Process llama-server
    reset compute queue (1:1:0)
    Ring comp_1.1.0 reset failed
    GPU reset begin!
    MODE2 reset
    GPU reset succeeded, trying to resume
    GPU reset(1) succeeded!
    device wedged, but recovered through reset
    

    Restarted the server/proxy only, no config changes, went back to the same conversation, and used Regenerate on the same message. The retry ran for 2050.52 s and then failed again with decode() failed: vk::Queue::submit: ErrorDeviceLost

    Second kernel failure:

    Dumping IP State
    AMDGPU device coredump file has been created
    ring comp_1.1.0 timeout
    Process llama-server
    reset compute queue (1:1:0)
    Ring comp_1.1.0 reset succeeded
    device wedged, but recovered through reset
    

    This time the ring reset succeeded, and it didn't escalate to a MODE2 reset.

    There were also two earlier warnings during the second request:

    Fence fallback timer expired on ring comp_1.2.0
    Fence fallback timer expired on ring comp_1.2.0
    

    They occurred roughly 16 min. and 10 min. before the final comp_1.1.0 timeout. FWIW

    Host memory was fine during the retry: roughly 52 GB of 124 GB in use, 0% swap.

    No further attempts or changes to MTP/context/Vulkan/watchdog settings.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions