Conversation
Assisted-by: Codex
|
Using wave64 mode in rdna gpus is not supported from hip. I know its tempting, the gpus can dual issue valu instructions in this mode without requireing the Compiler to find opertunities for Packed fp32 math. See There are various mines to step on. |
|
@IMbackK I know. But the issue is, HIP is behind due to this. So the question is - do we take the performance loss by following the guidelines or do we just risk it and optimize for actual architectures? The more I'm looking into the drivers, the more I'm convinced that AMD is shooting themselves in the foot like this. The Vulkan backend is using wavefront64 and it gives them numerous decode advantages already. |
|
But of course if you say we're not using wave64, I'll revert to the "safe" wave32 kernels. |
|
I dont know wtf they are doing either, but in this case we have to do what they say as i have tried this for work and it was an absolute maintenance nightmare over time with: header implemented cuda compat intrinsics not working correctly for the extra threads, random hip headers having a preprocessor guard preventing wave64 with gfx10+ etc. I get that if its just one file (top_k) its probably not going to hurt too much, but using it helps in alot of places for gfx11 (indeed see ggml-vulkan or mesa/aco generally it uses wave64 almost exclusively) so i dont think we can take the precedence. |
|
Aight, reverting to wave32 then. Argh. |
Assisted-by: OpenAI Codex
RDNA4 supports wave64, but it operates differently on the hardware and I almost believe you loose performance from what I remember. |
From the Vulkan experience, we'ved switched RDNA1/2 to wave32 in some cases, but not RDNA3/4, there wave64 seemed advantageous for most kernels. |
|
@IMbackK any chance you could look at it on your end? Would be nice to have the TOP_K issue resolved :) |
|
I have pretty limited time to work on llamacpp but its on my todo list, i will get around to it.
|
|
@pwilkin are you still perusing getting this merged? |
|
Yeah, will check your remarks. |
Overview
The current TOP_K kernels are still far from optimal, @Geramy 's native library fix resolves some of the cases, but not all of them.
Additional information
Added a small-case kernel (ported @0cc4m 's Vulkan kernel basically) and sorted out the optimal cases. Comparison table (tests disabled HIP graphs because the graph update algorithm is currently broken in upstream code pending ROCm/rocm-systems#11069):
type=f32,ne=[2,1,1,1],k=1,ties=0type=f32,ne=[4096,1,1,1],k=16,ties=0type=f32,ne=[4096,16,1,1],k=16,ties=0type=f32,ne=[8192,1,1,1],k=16,ties=0type=f32,ne=[8192,16,1,1],k=16,ties=0type=f32,ne=[12288,1,1,1],k=16,ties=0type=f32,ne=[12288,16,1,1],k=16,ties=0type=f32,ne=[16384,1,1,1],k=16,ties=0type=f32,ne=[16384,16,1,1],k=16,ties=0type=f32,ne=[24576,1,1,1],k=16,ties=0type=f32,ne=[24576,16,1,1],k=16,ties=0type=f32,ne=[32768,1,1,1],k=16,ties=0type=f32,ne=[32768,16,1,1],k=16,ties=0type=f32,ne=[65536,1,1,1],k=16,ties=0type=f32,ne=[65536,16,1,1],k=16,ties=0type=f32,ne=[131072,1,1,1],k=16,ties=0type=f32,ne=[131072,16,1,1],k=16,ties=0type=f32,ne=[1,1,1,1],k=1,ties=0type=f32,ne=[1000,1,1,1],k=1,ties=0type=f32,ne=[65000,1,1,1],k=1,ties=0type=f32,ne=[200000,1,1,1],k=1,ties=0type=f32,ne=[1,16,1,1],k=1,ties=0type=f32,ne=[1000,16,1,1],k=1,ties=0type=f32,ne=[65000,16,1,1],k=1,ties=0type=f32,ne=[200000,16,1,1],k=1,ties=0type=f32,ne=[4,1,1,1],k=4,ties=0type=f32,ne=[1000,1,1,1],k=4,ties=0type=f32,ne=[65000,1,1,1],k=4,ties=0type=f32,ne=[200000,1,1,1],k=4,ties=0type=f32,ne=[4,16,1,1],k=4,ties=0type=f32,ne=[1000,16,1,1],k=4,ties=0type=f32,ne=[65000,16,1,1],k=4,ties=0type=f32,ne=[200000,16,1,1],k=4,ties=0type=f32,ne=[8,1,1,1],k=8,ties=0type=f32,ne=[1000,1,1,1],k=8,ties=0type=f32,ne=[65000,1,1,1],k=8,ties=0type=f32,ne=[200000,1,1,1],k=8,ties=0type=f32,ne=[8,16,1,1],k=8,ties=0type=f32,ne=[1000,16,1,1],k=8,ties=0type=f32,ne=[65000,16,1,1],k=8,ties=0type=f32,ne=[200000,16,1,1],k=8,ties=0type=f32,ne=[10,1,1,1],k=10,ties=0type=f32,ne=[1000,1,1,1],k=10,ties=0type=f32,ne=[65000,1,1,1],k=10,ties=0type=f32,ne=[200000,1,1,1],k=10,ties=0type=f32,ne=[10,16,1,1],k=10,ties=0type=f32,ne=[1000,16,1,1],k=10,ties=0type=f32,ne=[65000,16,1,1],k=10,ties=0type=f32,ne=[200000,16,1,1],k=10,ties=0type=f32,ne=[16,1,1,1],k=16,ties=0type=f32,ne=[1000,1,1,1],k=16,ties=0type=f32,ne=[65000,1,1,1],k=16,ties=0type=f32,ne=[200000,1,1,1],k=16,ties=0type=f32,ne=[16,16,1,1],k=16,ties=0type=f32,ne=[1000,16,1,1],k=16,ties=0type=f32,ne=[65000,16,1,1],k=16,ties=0type=f32,ne=[200000,16,1,1],k=16,ties=0type=f32,ne=[32,1,1,1],k=32,ties=0type=f32,ne=[1000,1,1,1],k=32,ties=0type=f32,ne=[65000,1,1,1],k=32,ties=0type=f32,ne=[200000,1,1,1],k=32,ties=0type=f32,ne=[32,16,1,1],k=32,ties=0type=f32,ne=[1000,16,1,1],k=32,ties=0type=f32,ne=[65000,16,1,1],k=32,ties=0type=f32,ne=[200000,16,1,1],k=32,ties=0type=f32,ne=[40,1,1,1],k=40,ties=0type=f32,ne=[1000,1,1,1],k=40,ties=0type=f32,ne=[65000,1,1,1],k=40,ties=0type=f32,ne=[200000,1,1,1],k=40,ties=0type=f32,ne=[40,16,1,1],k=40,ties=0type=f32,ne=[1000,16,1,1],k=40,ties=0type=f32,ne=[65000,16,1,1],k=40,ties=0type=f32,ne=[200000,16,1,1],k=40,ties=0type=f32,ne=[400,1,1,1],k=400,ties=0type=f32,ne=[1000,1,1,1],k=400,ties=0type=f32,ne=[65000,1,1,1],k=400,ties=0type=f32,ne=[200000,1,1,1],k=400,ties=0type=f32,ne=[400,16,1,1],k=400,ties=0type=f32,ne=[1000,16,1,1],k=400,ties=0type=f32,ne=[65000,16,1,1],k=400,ties=0type=f32,ne=[200000,16,1,1],k=400,ties=0Requirements