Skip to content

ggml-cuda : add flash attention support for head size 88 (Llama 4 Vision) - #20375

Open
caffeinatedbits wants to merge 1 commit into
ggml-org:masterfrom
caffeinatedbits:fix-llama4-vision
Open

caffeinatedbits wants to merge 1 commit into
ggml-org:masterfrom
caffeinatedbits:fix-llama4-vision

ggml-cuda : add flash attention support for head size 88

41e6c8c
Select commit
Loading
Failed to load commit list.
Sign in for the full log view

The logs for this run have expired and are no longer available.