Repository navigation
Misc. bug: Vulkan: DeviceLost on AMD APUs (gfx90c) due to GPU job timeout from command batch size #21724
Description
Activity
A note on why a code-level fix is preferable to increasing the kernel timeout:
As far as I know, the
amdgpu.lockup_timeoutparameter can only be set as a boot parameter (requires GRUB config change + reboot), and increasing could have side effects. The timeout exists to detect real GPU hangs (driver bugs, infinite shader loops, etc.). On an APU where the display compositor shares the same GPU, a longer timeout means a longer desktop freeze when an actual hang occurs — a 60-second timeout means a 60-second frozen screen.For reference,
nodes_per_submit=1showed no performance regression on the gfx90c APU (actually ~3% faster at 512×512), but faster discrete GPUs may benefit from batching — which is why an environment variable override or device-based heuristic would be better than a hard-coded global change. Other APUs with more CUs or different architectures may have different performance characteristics where some batching is still beneficial. More testing across a range of integrated GPUs would be needed to find the right balance. An environment variable would at least let users on other hardware experiment without code changes.Reacted by dsignarius and Ross RosarioI'm observing the same thing on Vega 64 and Qwen3.6. Normal text processing works all right. Only when I load up some bigger image (with mmproj), llama-server crashes with:
radv/amdgpu: The CS has been cancelled because the context is lost. This context is innocent. [...] terminate called after throwing an instance of 'vk::DeviceLostError' what(): vk::Queue::submit: ErrorDeviceLostSimilar error in dmesg:
[15356.800251] amdgpu 0000:2f:00.0: amdgpu: Dumping IP State [15356.803152] amdgpu 0000:2f:00.0: amdgpu: Dumping IP State Completed [15356.803166] amdgpu 0000:2f:00.0: amdgpu: [drm] AMDGPU device coredump file has been created [15356.803168] amdgpu 0000:2f:00.0: amdgpu: [drm] Check your /sys/class/drm/card1/device/devcoredump/data [15356.808591] amdgpu 0000:2f:00.0: amdgpu: ring comp_1.1.0 timeout, but soft recovered [15358.848250] amdgpu 0000:2f:00.0: amdgpu: Dumping IP State [15358.851346] amdgpu 0000:2f:00.0: amdgpu: Dumping IP State Completed [15358.851360] amdgpu 0000:2f:00.0: amdgpu: [drm] AMDGPU device coredump file has been created [15358.851362] amdgpu 0000:2f:00.0: amdgpu: [drm] Check your /sys/class/drm/card1/device/devcoredump/data [15358.861368] amdgpu 0000:2f:00.0: amdgpu: ring comp_1.1.0 timeout, signaled seq=1947328, emitted seq=1947331 [15358.861371] amdgpu 0000:2f:00.0: amdgpu: Process llama-server pid 189134 thread llama-server pid 189134 [15358.861373] amdgpu 0000:2f:00.0: amdgpu: Starting comp_1.1.0 ring reset [15358.861457] amdgpu 0000:2f:00.0: amdgpu: Ring comp_1.1.0 reset succeeded [15358.861478] amdgpu 0000:2f:00.0: [drm] device wedged, but recovered through reset [15360.895269] amdgpu 0000:2f:00.0: amdgpu: Dumping IP State [15360.898375] amdgpu 0000:2f:00.0: amdgpu: Dumping IP State Completed [15360.898388] amdgpu 0000:2f:00.0: amdgpu: [drm] AMDGPU device coredump file has been created [15360.898390] amdgpu 0000:2f:00.0: amdgpu: [drm] Check your /sys/class/drm/card1/device/devcoredump/data [15360.908394] amdgpu 0000:2f:00.0: amdgpu: ring comp_1.1.0 timeout, signaled seq=1947332, emitted seq=1947334 [15360.908398] amdgpu 0000:2f:00.0: amdgpu: Process llama-server pid 189134 thread llama-server pid 189134 [15360.908401] amdgpu 0000:2f:00.0: amdgpu: Starting comp_1.1.0 ring reset [15360.908482] amdgpu 0000:2f:00.0: amdgpu: Ring comp_1.1.0 reset succeeded [15360.908484] amdgpu 0000:2f:00.0: [drm] device wedged, but recovered through resetApplying mentioned workaround (
nodes_per_submit = 1or10) allows normal processing without any crashes.Reacted by Terry Lentz Jrmollahmasud977-sys commented
on Apr 17, 2026 on Apr 17, 2026 via email · Hidden as spamshow commentMore actionsWitnessing the same issue, albeit occasionally but definitely recently, with GLM4.6V IQ4XS quant from unsloth, even without loading the mmproj file or passing images via prompt. It happens immediately after pp is done.
Running on 3x 7900xtx gpus.
Command:
llama-server --host 0.0.0.0 --port 5805 -dev Vulkan1,Vulkan2,Vulkan3 --no-warmup -ngl all -fa on --sampling-seq k --top-k 1 --parallel 1 --predict 49152 --cache-ram -1 --spec-default --fit-target 64 --reasoning on -m /home/user/models/large/glm46v.unsloth.iq4xs.ggufradv/amdgpu: The CS has been cancelled because the context is lost. This context is innocent. /usr/lib/libggml-base.so.0(+0x1734b) [0x7faa0234934b] /usr/lib/libggml-base.so.0(ggml_print_backtrace+0x20f) [0x7faa0234971f] /usr/lib/libggml-base.so.0(+0x31fae) [0x7faa02363fae] /usr/lib/libstdc++.so.6(+0xb1eba) [0x7faa01cb1eba] /usr/lib/libstdc++.so.6(_ZSt10unexpectedv+0x0) [0x7faa01c975d9] /usr/lib/libstdc++.so.6(+0xb2176) [0x7faa01cb2176] /usr/lib/libggml-vulkan.so.0(+0x73b41) [0x7fa9fd473b41] /usr/lib/libggml-vulkan.so.0(+0x1b977f) [0x7fa9fd5b977f] /usr/lib/libggml-vulkan.so.0(+0x1c0101) [0x7fa9fd5c0101] /usr/lib/libggml-base.so.0(ggml_backend_sched_graph_compute_async+0x902) [0x7faa023693c2] /usr/lib/libllama.so.0(_ZN13llama_context13graph_computeEP11ggml_cgraphb+0xb2) [0x7faa0209f9d2] /usr/lib/libllama.so.0(_ZN13llama_context14process_ubatchERK12llama_ubatch14llm_graph_typeP22llama_memory_context_iR11ggml_status+0x12b) [0x7faa02093bcb] /usr/lib/libllama.so.0(_ZN13llama_context6decodeERK11llama_batch+0x40a) [0x7faa02097a9a] /usr/lib/libllama.so.0(llama_decode+0x16) [0x7faa020a24d6] llama-server(+0xd3af9) [0x55ad9840caf9] llama-server(+0x1668b3) [0x55ad9849f8b3] llama-server(+0x6a59a) [0x55ad983a359a] /usr/lib/libc.so.6(+0x276c1) [0x7faa018276c1] /usr/lib/libc.so.6(__libc_start_main+0x89) [0x7faa018277f9] llama-server(+0x6bf75) [0x55ad983a4f75] terminate called after throwing an instance of 'vk::DeviceLostError' what(): vk::Queue::submit: ErrorDeviceLost build_info: b8892-df978aaWitnessing the same issue, albeit occasionally but definitely recently, with GLM4.6V IQ4XS quant from unsloth, even without loading the mmproj file or passing images via prompt. It happens immediately after pp is done.
BTW I'm pretty sure this is a regression after downgrading llama.cpp to b8270 as suggested in #21811
I'm getting a very similar issue on my Strix Halo (gfx1151) setup:
[7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] radv/amdgpu: The CS has been cancelled because the context is lost. This context is innocent. [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libggml-base.so.0(+0x1979a) [0x75237995179a] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libggml-base.so.0(ggml_print_backtrace+0x204) [0x752379951c64] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libggml-base.so.0(+0x2e189) [0x752379966189] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/si4q3zks5mn5jhzzyri9hhd3cv789vlm-gcc-15.2.0-lib/lib/libstdc++.so.6(+0xc539a) [0x752378cc539a] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/si4q3zks5mn5jhzzyri9hhd3cv789vlm-gcc-15.2.0-lib/lib/libstdc++.so.6(_ZSt10unexpectedv+0x0) [0x752378cb286e] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/si4q3zks5mn5jhzzyri9hhd3cv789vlm-gcc-15.2.0-lib/lib/libstdc++.so.6(+0xc5637) [0x752378cc5637] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libggml-vulkan.so.0(+0x9eead) [0x752379a9eead] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libggml-vulkan.so.0(+0x19f387) [0x752379b9f387] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libggml-vulkan.so.0(+0x19f728) [0x752379b9f728] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libggml-base.so.0(ggml_backend_sched_graph_compute_async+0x92c) [0x7523799702fc] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libllama.so.0(_ZN13llama_context13graph_computeEP11ggml_cgraphb+0xa1) [0x75237d4d2271] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libllama.so.0(_ZN13llama_context14process_ubatchERK12llama_ubatch14llm_graph_typeP22llama_memory_context_iR11ggml_status+0x11f) [0x75237d4d52ff] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libllama.so.0(_ZN13llama_context6decodeERK11llama_batch+0x3c9) [0x75237d4dc0e9] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/lib/libllama.so.0(llama_decode+0x11) [0x75237d4dddc1] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/bin/llama-server(+0x11d0c8) [0x5795312c40c8] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/bin/llama-server(+0x1b75c1) [0x57953135e5c1] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/bin/llama-server(+0x72b3f) [0x579531219b3f] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/fjkx1l5cnskzrqacf08z7i8z17256w0j-glibc-2.42-61/lib/libc.so.6(+0x2b285) [0x75237882b285] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/fjkx1l5cnskzrqacf08z7i8z17256w0j-glibc-2.42-61/lib/libc.so.6(__libc_start_main+0x88) [0x75237882b338] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/bin/llama-server(+0x73215) [0x57953121a215] [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] terminate called after throwing an instance of 'vk::DeviceLostError' [7:35:56 PM] maj 08 19:35:55 RX-78-FPC llm-router[15732]: [38593] what(): vk::Queue::submit: ErrorDeviceLost [7:36:00 PM] maj 08 19:36:00 RX-78-FPC llm-router[15732]: srv operator(): http client error: Failed to read connectionThis happens on IQ4_NL Qwen 3.6 27B model, b9075 release (latest
masterbranch), but i've noticed it on previous releases too. I'm using Vulkan back-end with RADV on latest NixOS unstable.Some more info from llama-server logs:
[7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: spawning server instance with name=coder-smart on port 38593 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: spawning server instance with args: [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: /nix/store/s2jab8bjlxc1h503ylkjcs8y4b52p1m3-llama-cpp-mpi-vulkan-4.2.0/bin/llama-server [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --chat-template-kwargs [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: {"enable_thinking":false} [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --host [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 127.0.0.1 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --jinja [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --metrics [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --min-p [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 0.0 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --mlock [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --no-mmap [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --offline [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --port [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 38593 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --presence-penalty [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 1.5 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --prio [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 2 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --props [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --repeat-penalty [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 1.0 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --slots [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --temperature [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 1.0 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --top-k [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 20 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --top-p [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 0.95 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --warmup [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --webui [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --alias [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: coder-smart [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --batch-size [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 4096 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --cont-batching [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --cache-ram [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 40960 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --swa-checkpoints [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 100 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --direct-io [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --flash-attn [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: on [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --fit [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: off [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --model [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: /home/LLMs/llama-models/Qwen3.6-27B-IQ4_NL.gguf [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --n-gpu-layers [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: all [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --parallel [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 1 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --threads [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 32 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --threads-batch [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 32 [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: --ubatch-size [7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: srv load: 2048[7:28:26 PM] maj 08 19:28:26 RX-78-FPC llm-router[15732]: [38593] system_info: n_threads = 32 (n_threads_batch = 32) / 32 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |And
journalctl -k:maj 08 19:27:35 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Dumping IP State maj 08 19:27:35 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Dumping IP State Completed maj 08 19:27:35 RX-78-FPC kernel: amdgpu 0000:c2:00.0: [drm] AMDGPU device coredump file has been created maj 08 19:27:35 RX-78-FPC kernel: amdgpu 0000:c2:00.0: [drm] Check your /sys/class/drm/card1/device/devcoredump/data maj 08 19:27:35 RX-78-FPC kernel: amdgpu 0000:c2:00.0: ring gfx_0.0.0 timeout, signaled seq=166926, emitted seq=166928 maj 08 19:27:35 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Process llama-server pid 12519 thread llama-server pid 12519 maj 08 19:27:35 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Starting gfx_0.0.0 ring reset maj 08 19:27:35 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Ring gfx_0.0.0 reset succeeded maj 08 19:27:35 RX-78-FPC kernel: amdgpu 0000:c2:00.0: [drm] device wedged, but recovered through reset maj 08 19:28:59 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Fence fallback timer expired on ring gfx_0.0.0 maj 08 19:35:55 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Dumping IP State maj 08 19:35:55 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Dumping IP State Completed maj 08 19:35:55 RX-78-FPC kernel: amdgpu 0000:c2:00.0: [drm] AMDGPU device coredump file has been created maj 08 19:35:55 RX-78-FPC kernel: amdgpu 0000:c2:00.0: [drm] Check your /sys/class/drm/card1/device/devcoredump/data maj 08 19:35:55 RX-78-FPC kernel: amdgpu 0000:c2:00.0: ring gfx_0.0.0 timeout, signaled seq=168485, emitted seq=168487 maj 08 19:35:55 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Process llama-server pid 15778 thread llama-server pid 15778 maj 08 19:35:55 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Starting gfx_0.0.0 ring reset maj 08 19:35:55 RX-78-FPC kernel: amdgpu 0000:c2:00.0: Ring gfx_0.0.0 reset succeeded maj 08 19:35:55 RX-78-FPC kernel: amdgpu 0000:c2:00.0: [drm] device wedged, but recovered through resetWhat else do we need to confirm this bug?
Reacted by Ross Rosario and System64What else do we need to confirm this bug?
Thank you. Working on a fix is another story, but this issue should be at least marked as confirmed on
AMD+Vulkanby now.With
nodes_per_submit = 1I still get:[Thread debugging using libthread_db enabled] Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1". __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/x86_64/syscall_cancel.S:56 warning: 56 ../sysdeps/unix/sysv/linux/x86_64/syscall_cancel.S: No such file or directory #0 __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/x86_64/syscall_cancel.S:56 56 in ../sysdeps/unix/sysv/linux/x86_64/syscall_cancel.S #1 0x00007f050269b668 in __internal_syscall_cancel (a1=<optimized out>, a2=<optimized out>, a3=<optimized out>, a4=<optimized out>, a5=a5@entry=0, a6=a6@entry=0, nr=61) at ./nptl/cancellation.c:49 warning: 49 ./nptl/cancellation.c: No such file or directory #2 0x00007f050269b6ad in __syscall_cancel (a1=<optimized out>, a2=<optimized out>, a3=<optimized out>, a4=<optimized out>, a5=a5@entry=0, a6=a6@entry=0, nr=61) at ./nptl/cancellation.c:75 75 in ./nptl/cancellation.c #3 0x00007f05027067c7 in __GI___wait4 (pid=<optimized out>, stat_loc=<optimized out>, options=<optimized out>, usage=<optimized out>) at ../sysdeps/unix/sysv/linux/wait4.c:30 warning: 30 ../sysdeps/unix/sysv/linux/wait4.c: No such file or directory #4 0x00007f050751b50b in ggml_print_backtrace () at /home/daniel/Build/llama.cpp/ggml/src/ggml.c:234 234 waitpid(child_pid, NULL, 0); #5 0x00007f0507530f7d in ggml_uncaught_exception () at /home/daniel/Build/llama.cpp/ggml/src/ggml.cpp:9 9 ggml_print_backtrace(); #6 0x00007f05028b344a in ?? () from /lib/x86_64-linux-gnu/libstdc++.so.6 #7 0x00007f05028a15e9 in std::terminate() () from /lib/x86_64-linux-gnu/libstdc++.so.6 #8 0x00007f05028b36c8 in __cxa_throw () from /lib/x86_64-linux-gnu/libstdc++.so.6 #9 0x00007f0502d84fcd in vk::detail::throwResultException (result=vk::Result::eErrorDeviceLost, message=0x7f0502f65838 "vk::Queue::submit") at /home/daniel/Build/vulkansdk/1.4.341.1/x86_64/include/vulkan/vulkan.hpp:8232 8232 case Result::eErrorDeviceLost : throw DeviceLostError( message ); #10 vk::detail::resultCheck (result=vk::Result::eErrorDeviceLost, message=0x7f0502f65838 "vk::Queue::submit") at /home/daniel/Build/vulkansdk/1.4.341.1/x86_64/include/vulkan/vulkan.hpp:8498 8498 throwResultException( result, message ); #11 vk::Queue::submit<vk::detail::DispatchLoaderDynamic, true> (this=0x55b2ec25a360, submits=..., fence=..., d=...) at /home/daniel/Build/vulkansdk/1.4.341.1/x86_64/include/vulkan/vulkan_funcs.hpp:937 937 detail::resultCheck( result, VULKAN_HPP_NAMESPACE_STRING "::Queue::submit" ); #12 ggml_vk_submit (ctx=std::shared_ptr<vk_context_struct> (use count 4, weak count 2) = {...}, fence=...) at /home/daniel/Build/llama.cpp/ggml/src/ggml-vulkan/ggml-vulkan.cpp:2493 2493 ctx->p->q->queue.submit(submit_infos, fence); #13 0x00007f0502e8c2b4 in ggml_vk_compute_forward (ctx=0x55b2e903e7b0, cgraph=0x55b2ed0b47b8, tensor=0x55b2ec39b640, tensor_idx=0, almost_ready=false) at /home/daniel/Build/llama.cpp/ggml/src/ggml-vulkan/ggml-vulkan.cpp:13530 13530 ggml_vk_submit(subctx, {}); #14 0x00007f0502e8c073 in ggml_vk_build_graph (ctx=0x55b2e903e7b0, cgraph=0x55b2ed0b47b8, node_idx=2, node_begin=0x55b2ec39b640, node_idx_begin=0, last_node=false, almost_ready=false, submit=true) at /home/daniel/Build/llama.cpp/ggml/src/ggml-vulkan/ggml-vulkan.cpp:13498 13498 ggml_vk_compute_forward(ctx, cgraph, node_begin, node_idx_begin, almost_ready); #15 0x00007f0502e985a9 in ggml_backend_vk_graph_compute (backend=0x55b2ea417a00, cgraph=0x55b2ed0b47b8) at /home/daniel/Build/llama.cpp/ggml/src/ggml-vulkan/ggml-vulkan.cpp:14898 14898 bool enqueued = ggml_vk_build_graph(ctx, cgraph, i, cgraph->nodes[submit_node_idx], submit_node_idx, i + ctx->num_additional_fused_ops >= last_node, almost_ready, submit); #16 0x00007f0507537604 in ggml_backend_graph_compute_async (backend=0x55b2ea417a00, cgraph=0x55b2ed0b47b8) at /home/daniel/Build/llama.cpp/ggml/src/ggml-backend.cpp:452 452 return backend->iface.graph_compute(backend, cgraph); #17 0x00007f050753c36d in ggml_backend_sched_compute_splits (sched=0x55b2e9014710) at /home/daniel/Build/llama.cpp/ggml/src/ggml-backend.cpp:1678 1678 enum ggml_status ec = ggml_backend_graph_compute_async(split_backend, &split->graph); #18 0x00007f050753d213 in ggml_backend_sched_graph_compute_async (sched=0x55b2e9014710, graph=0x55b2ec318350) at /home/daniel/Build/llama.cpp/ggml/src/ggml-backend.cpp:1901 1901 return ggml_backend_sched_compute_splits(sched); #19 0x00007f0507072fc5 in llama_context::graph_compute (this=0x55b2e8f426b0, gf=0x55b2ec318350, batched=true) at /home/daniel/Build/llama.cpp/src/llama-context.cpp:2191 2191 auto status = ggml_backend_sched_graph_compute_async(sched.get(), gf); #20 0x00007f050706e8f8 in llama_context::process_ubatch (this=0x55b2e8f426b0, ubatch=..., gtype=LLM_GRAPH_TYPE_DECODER, mctx=0x55b2edfd9670, ret=@0x7ffe692e6ecc: GGML_STATUS_SUCCESS) at /home/daniel/Build/llama.cpp/src/llama-context.cpp:1231 1231 const auto status = graph_compute(res->get_gf(), ubatch.n_tokens > 1); #21 0x00007f05070706b8 in llama_context::decode (this=0x55b2e8f426b0, batch_inp=...) at /home/daniel/Build/llama.cpp/src/llama-context.cpp:1692 1692 const auto * res = process_ubatch(ubatch, LLM_GRAPH_TYPE_DECODER, mctx.get(), status); #22 0x00007f050707736b in llama_decode (ctx=0x55b2e8f426b0, batch=...) at /home/daniel/Build/llama.cpp/src/llama-context.cpp:3510 3510 const int ret = ctx->decode(batch); #23 0x000055b2e6bbcc79 in server_context_impl::update_slots (this=0x55b2e8f5a390) at /home/daniel/Build/llama.cpp/tools/server/server-context.cpp:2797 2797 const int ret = llama_decode(ctx, batch_view); #24 0x000055b2e6bb0ce9 in server_context_impl::init()::{lambda()#1}::operator()() const (__closure=0x55b2e8f5a4d0) at /home/daniel/Build/llama.cpp/tools/server/server-context.cpp:997 997 update_slots(); #25 0x000055b2e6be8cca in std::__invoke_impl<void, server_context_impl::init()::{lambda()#1}&>(std::__invoke_other, server_context_impl::init()::{lambda()#1}&) (__f=...) at /usr/include/c++/14/bits/invoke.h:61 61 { return std::forward<_Fn>(__f)(std::forward<_Args>(__args)...); } #26 0x000055b2e6bdded4 in std::__invoke_r<void, server_context_impl::init()::{lambda()#1}&>(server_context_impl::init()::{lambda()#1}&) (__fn=...) at /usr/include/c++/14/bits/invoke.h:111 111 std::__invoke_impl<__type>(__tag{}, std::forward<_Callable>(__fn), #27 0x000055b2e6bd346e in std::_Function_handler<void (), server_context_impl::init()::{lambda()#1}>::_M_invoke(std::_Any_data const&) (__functor=...) at /usr/include/c++/14/bits/std_function.h:290 290 return std::__invoke_r<_Res>(*_Base::_M_get_pointer(__functor), #28 0x000055b2e6adb87e in std::function<void()>::operator() (this=0x55b2e8f5a4d0) at /usr/include/c++/14/bits/std_function.h:591 591 return _M_invoker(_M_functor, std::forward<_ArgTypes>(__args)...); #29 0x000055b2e6cbd039 in server_queue::start_loop (this=0x55b2e8f5a3a8, idle_sleep_ms=-1000) at /home/daniel/Build/llama.cpp/tools/server/server-queue.cpp:163 163 callback_update_slots(); #30 0x000055b2e6b8ab01 in server_context::start_loop (this=0x7ffe692ee5c8) at /home/daniel/Build/llama.cpp/tools/server/server-context.cpp:3097 3097 impl->queue_tasks.start_loop(params.sleep_idle_seconds * 1000); #31 0x000055b2e6acc0aa in main (argc=7, argv=0x7ffe692f1248) at /home/daniel/Build/llama.cpp/tools/server/server.cpp:336 336 ctx_server.start_loop(); [Inferior 1 (process 7532) detached] terminate called after throwing an instance of 'vk::DeviceLostError' what(): vk::Queue::submit: ErrorDeviceLostThough I am not sure it is the same bug. Also occurs with b9049.
Wondering if this (from the SYCL toolkit install) could be an answer:
GRUB_CMDLINE_LINUX_DEFAULT:
i915.enable_hangcheck=0@danielfdickinson
i915? That's an Intel driver.$ modinfo i915 | grep hangcheck parm: enable_hangcheck:Periodically check GPU activity for detecting hangs. WARNING: Disabling this can cause system wide hangs. (default: true) (bool)I think that's the same as
amdgpu.lockup_timeoutbut without changeable timeout.Sorry missed the AMD. Obviously not the same bug, though same error.
4 remaining items
Confirmed on AMD 760M (RADV PHOENIX, gfx1102) with llama.cpp — same root cause (amdgpu.lockup_timeout exceeded by large vkQueueSubmit batches).
I implemented the auto-detect approach @terry-lentz suggested in the issue body: nodes_per_submit = ctx->device->uma ? 10 : 100, using the existing eIntegratedGpu detection in the Vulkan backend. This keeps 100 for discrete GPUs (zero impact) and uses 10 for UMA/iGPUs.
Tested with Gemma-4 26B IQ4_NL, turbo3 KV, FlashAttention on:
pp16384: GPU hang → 122 t/s
tg32 @ 188k context: unusable (0.099 t/s) → 22.09 t/s
Full context range up to 262k now stable
I also tested nodes_per_submit=1 (as suggested in the issue) — it fixes the hang but has significant submit overhead at large prompt sizes (pp16384 took >30min vs 122 t/s with nps=10). 10 is a better tradeoff for UMA devices.Note: This only addresses UMA/APU devices. The discrete GPU reports (7900XTX, Vega 64, RX 7600 XT) may need a different approach since eIntegratedGpu won't match them — possibly the env variable override @terry-lentz suggested, or a broader heuristic.
Isolated commit (cherry-pickable, based on recent upstream): https://github.com/fukuro-kun/fukuro-llama-cpp-turboquant/tree/fix/vulkan-uma-nodes-per-submit
- added 3 commits that reference this issue
on Jun 21, 2026 - added a commit that references this issue
on Jul 14, 2026 AMD 780M Vulkan DeviceLost with long context — workaround
I’m running
llama.cppVulkan on an AMD Ryzen 7 8845HS / Radeon 780M iGPU with a 64 GB shared-memory system.With very large contexts,
llama-serverconsistently failed around 90–103k tokens:radv/amdgpu: The CS has been cancelled because the context is lost. ggml_vulkan: device lost on Vulkan0 vk::Queue::submit: ErrorDeviceLostKernel log showed:
amdgpu: ring comp_1.3.0 timeout amdgpu: Starting comp_1.3.0 ring resetReducing:
GGML_VK_MAX_NODES_PER_SUBMIT=1
only moved the failure threshold (~90.7k → ~103k); it did not eliminate it.
The actual workaround was increasing the AMDGPU watchdog timeout from its 2s default:
sudo grubby --update-kernel=ALL --args="amdgpu.lockup_timeout=10000"After reboot, the long-context Vulkan workload can proceed past the previous failure point.
Hardware/software: Ryzen 7 8845HS / Radeon 780M, Fedora 44, kernel 7.1.8, RADV, llama.cpp
a94d563ed.Reacted by Furkan and KarakurtReacted by fairydreaming and FurkanReacted by Noel Maersk, Andrea Santoro and Furkan- added a commit that references this issue
on Sep 22, 2026
Name and Version
ggml commit 58c3805 (via stable-diffusion.cpp, not llama-cli directly)
Reproduced using the shared ggml Vulkan backend (ggml-vulkan.cpp)
Operating systems
Linux
Which llama.cpp modules do you know to be affected?
No response
Command line
Problem description & steps to reproduce
On AMD Renoir APU (gfx90c, RADV), Vulkan compute jobs crash with
vk::Queue::submit: ErrorDeviceLostwhen running large compute graphs.Root Cause
The default Linux
amdgpu.lockup_timeoutis 2000ms.ggml_backend_vk_graph_computebatches up to 100 nodes pervkQueueSubmit. On slow integrated GPUs/APUs, the accumulated GPU work in a single submission exceeds this timeout, causing the kernel to reset the compute ring.The relevant code in
ggml-vulkan.cpp:At smaller workloads each batch completes in time. At larger workloads (e.g., 1024×1024 diffusion with seq_len ≈ 4608), the quadratic attention scaling pushes batches past the 2-second limit.
Steps to Reproduce
Encountered via stable-diffusion.cpp (which vendors ggml, commit
58c38058) running Flux 2 Klein 4B at 1024×1024 on AMD Renoir APU:sd-cli --diffusion-model flux-2-klein-4b-Q4_0.gguf --vae ae.safetensors \ --llm Qwen3-4B-Q4_K_M.gguf --cfg-scale 1.0 --steps 4 --vae-tiling \ -H 1024 -W 1024 -p "a cat sitting in Taipei" -o output.pngThe same issue would affect llama.cpp with sufficiently long contexts on APU hardware, since the Vulkan backend code is shared.
Validated Fix
Changing
nodes_per_submitfrom100to1resolves the crash with no measurable performance regression (actually ~3% faster at 512×512):Suggested Improvement
Rather than hardcoding
nodes_per_submit = 1globally, a proper fix could:GGML_VK_NODES_PER_SUBMIT)Note
We encountered this in stable-diffusion.cpp's vendored ggml (commit
58c38058). We have a working local fix, but are filing here since llama.cpp is where the Vulkan backend is actively maintained and changes sync downstream.Kernel timeout can also be increased as a user-side workaround:
sudo sh -c 'echo 60000 > /sys/module/amdgpu/parameters/lockup_timeout'First Bad Commit
No response
Relevant log output
From kernel log (journalctl -k):
amdgpu: ring comp_1.2.0 timeout, signaled seq=12481, emitted seq=12485
amdgpu: Process sd-cli pid 33328 thread sd-cli pid 33328
amdgpu: Starting comp_1.2.0 ring reset
amdgpu: Ring comp_1.2.0 reset succeeded
[drm] device wedged, but recovered through reset
From sd-cli stderr:
radv/amdgpu: The CS has been cancelled because the context is lost. This context is guilty of a hard recovery.
terminate called after throwing an instance of 'vk::DeviceLostError'
what(): vk::Queue::submit: ErrorDeviceLost