Name and Version
$ ./llama --version
version: 0.5.0-dev (build 11188, commit e85e15c)
built with GNU 15.1.1 for Linux aarch64
Operating systems
Linux
Which llama.cpp modules do you know to be affected?
Test code
Command line
./test-backend-ops -o 'ARGSORT(type=f32,ne=[2048,1,1,1],order=0)'
Problem description & steps to reproduce
- On Adreno A750
maxComputeWorkGroupInvocations is 2048
maxComputeWorkGroupSize: [1024, 1024, 1024]
use_small is true in ggml_vk_argsort()
pipeline_argsort_f32[10] is used, which has BLOCK_SIZE of 1024
- 2 workgroups are dispatched
- Only half of the array is sorted
- Test fails
I think the 2 workgroups end up sorting the first 1024 elements (so those end up being sorted twice while the other 1024 half remains unsorted).
First Bad Commit
No response
Relevant log output
[ARGSORT] ERR = 0.999998560 > 0.000000100 ARGSORT(type=f32,ne=[2048,1,1,1],order=0): FAIL
[ARGSORT] ERR = 0.999998601 > 0.000000100 ARGSORT(type=f32,ne=[2048,1,1,1],order=0): FAIL
Name and Version
$ ./llama --version
version: 0.5.0-dev (build 11188, commit e85e15c)
built with GNU 15.1.1 for Linux aarch64
Operating systems
Linux
Which llama.cpp modules do you know to be affected?
Test code
Command line
./test-backend-ops -o 'ARGSORT(type=f32,ne=[2048,1,1,1],order=0)'Problem description & steps to reproduce
maxComputeWorkGroupInvocationsis2048maxComputeWorkGroupSize:[1024, 1024, 1024]use_smallistrueinggml_vk_argsort()pipeline_argsort_f32[10]is used, which hasBLOCK_SIZEof1024I think the 2 workgroups end up sorting the first 1024 elements (so those end up being sorted twice while the other 1024 half remains unsorted).
First Bad Commit
No response
Relevant log output
[ARGSORT] ERR = 0.999998560 > 0.000000100 ARGSORT(type=f32,ne=[2048,1,1,1],order=0): FAIL
[ARGSORT] ERR = 0.999998601 > 0.000000100 ARGSORT(type=f32,ne=[2048,1,1,1],order=0): FAIL