Repository navigation
Misc. bug: test-backend-ops not working correctly on vulkan (radv) #20249
Copy link
Copy link
Closed
Labels
VulkanIssues specific to the Vulkan backendIssues specific to the Vulkan backendbugSomething isn't workingSomething isn't working
Description
Activity
Do the correctness tests also fail, or is it just the perf tests?
Only the perf tests are failing.
I can reproduce it.
ggml_vk_submit(0x55b2b80022b0, 0) ggml_backend_vk_synchronize() ggml_vk_synchronize() ggml_vk_graph_cleanup() ggml_vk_command_pool_cleanup() ggml_vk_command_pool_cleanup() IM2COL(type_input=f32,type_kernel=f16,dst_type=f32,ne_input=[256,256,256,1],ne_kernel=[3,3,256,1],s0=1,s1=1,p0=1,p1=1,d0=1,d1=1,is_2D=1): 52 runs - 398217.73 us/run - 655364 kB/run - 1.57 GB/s ~vk_buffer_struct(0xe350000000e35, 671093248) ggml_vk_get_device(0) ggml_backend_vk_buffer_type(0) ggml_vk_get_device(0) ggml_vk_create_buffer(Vulkan0, 104903680, { DeviceLocal }, { HostVisible | HostCoherent }) ggml_backend_vk_buffer_init_tensor(0x55b2b8276890 (0x55b2b84f78b0), 0x55b2b93b6f20) ggml_backend_vk_buffer_init_tensor(0x55b2b8276890 (0x55b2b84f78b0), 0x55b2b93b7090) ggml_backend_vk_buffer_init_tensor(0x55b2b8276890 (0x55b2b84f78b0), 0x55b2b93b7200) ggml_backend_vk_buffer_set_tensor(0x55b2b8276890, 0x55b2b93b6f20, 0x7fbcdb5fe010, 0, 10485760) ggml_vk_buffer_write(10485760) ggml_vk_buffer_write_2d(10485760, 1) ggml_vk_create_temporary_context(0x55b2b7ff7b30) ggml_vk_ctx_begin(Vulkan0) ggml_vk_create_cmd_buffer() ggml_vk_buffer_write_2d_async(10485760, 1) STAGING ggml_vk_sync_buffers() ggml_vk_ctx_end(0x55b2b7ff7b30, 1) ggml_vk_submit(0x55b2b7ff7b30, 0x40000000004) ggml_vk_queue_command_pools_cleanup() ggml_vk_command_pool_cleanup() ggml_backend_vk_buffer_set_tensor(0x55b2b8276890, 0x55b2b93b7090, 0x55b2b8069350, 0, 46080) ggml_vk_buffer_write(46080) ggml_vk_buffer_write_2d(46080, 1) ggml_vk_create_temporary_context(0x55b2b8651df0) ggml_vk_ctx_begin(Vulkan0) ggml_vk_create_cmd_buffer() ggml_vk_buffer_write_2d_async(46080, 1) STAGING ggml_vk_sync_buffers() ggml_vk_ctx_end(0x55b2b8651df0, 1) ggml_vk_submit(0x55b2b8651df0, 0x40000000004) ggml_vk_queue_command_pools_cleanup() ggml_backend_vk_buffer_set_tensor(0x55b2b8276890, 0x55b2b93b7200, 0x7fbcc65ff010, 0, 94371840) ggml_vk_buffer_write(94371840) ggml_vk_buffer_write_2d(94371840, 1) ggml_vk_create_temporary_context(0x55b2b8651df0) ggml_vk_ctx_begin(Vulkan0) ggml_vk_create_cmd_buffer() ggml_vk_buffer_write_2d_async(94371840, 1) STAGING ggml_vk_sync_buffers() ggml_vk_ctx_end(0x55b2b8651df0, 1) ggml_vk_submit(0x55b2b8651df0, 0x40000000004) ggml_vk_queue_command_pools_cleanup() ggml_backend_vk_graph_compute(2 nodes) ggml_vk_create_context(0x55b2b8651df0) ggml_vk_ctx_begin(Vulkan0) ggml_vk_create_cmd_buffer() ggml_vk_buffer_memset_async(0, 0, 1024) ggml_vk_sync_buffers() ggml_vk_build_graph(0x55b2b93b7200, IM2COL) ggml_vk_op_f32((0x55b2b93b7090, name=kernel, type=1, ne0=3, ne1=3, ne2=2560, ne3=1, nb0=2, nb1=6, nb2=18, nb3=46080), (0x55b2b93b6f20, name=input, type=0, ne0=32, ne1=32, ne2=2560, ne3=1, nb0=4, nb1=128, nb2=4096, nb3=10485760), (0x55b2b93b7200, name=out, type=0, ne0=23040, ne1=32, ne2=32, ne3=1, nb0=4, nb1=92160, nb2=2949120, nb3=94371840), IM2COL) ggml_pipeline_request_descriptor_sets(im2col_f32, 1) ggml_vk_dispatch_pipeline(im2col_f32, {(0xe3b0000000e3b, 0, 10485760), (0xe3b0000000e3b, 10531840, 1), }, (1,32,2560)) ggml_vk_ctx_end(0x55b2b8651df0, 1) ggml_vk_compute_forward(0x55b2b93b7200, name=out, op=IM2COL, type=0, ne0=23040, ne1=32, ne2=32, ne3=1, nb0=4, nb1=92160, nb2=2949120, nb3=94371840, view_src=0, view_offs=0) ggml_vk_submit(0x55b2b8651df0, 0) radv/amdgpu: The CS has been cancelled because the context is lost. This context is guilty of a hard recovery. Validation Warning: [ BestPractices-Error-Result ] | MessageID = 0x53c1342f vkQueueSubmit(): Returned error VK_ERROR_DEVICE_LOST. Objects: 1 [0] VkQueue 0x55b2b7f6cb50It's caused by input width and height 256. The test does 32, 64 and 256, the third one crashes.
Maybe the command buffer just runs for too long?
- addedbugSomething isn't workingSomething isn't workingVulkanIssues specific to the Vulkan backendIssues specific to the Vulkan backendand removed
on Apr 9, 2026
Metadata
Metadata
Assignees
Labels
VulkanIssues specific to the Vulkan backendIssues specific to the Vulkan backendbugSomething isn't workingSomething isn't working
Name and Version
./llama-cli --version 5.546s
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon RX 7800 XT (RADV NAVI32) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 1 | matrix cores: KHR_coopmat
version: 8244 (35bee03)
built with GNU 15.2.1 for Linux x86_64
Operating systems
Linux
Which llama.cpp modules do you know to be affected?
Test code
Command line
Problem description & steps to reproduce
Starting with commit ad51c0a, using the mentioned command results in a crash. This has been tested on Arch Linux on different versions of mesa (old ones and git). I haven't run all the supported operations but at least one (IM2COL) is currently unusable, others like CONV_2D or ADD work correctly.
First Bad Commit
ad51c0a
Relevant log output
Logs