CUDA: add FP32 FlashAttention vector kernel (#7188)
* CUDA: add FP32 FlashAttention vector kernel * fixup! CUDA: add FP32 FlashAttention vector kernel * fixup! fixup! CUDA: add FP32 FlashAttention vector kernel * fixup! fixup! fixup! CUDA: add FP32 FlashAttention vector kernel
J
Johannes Gäßler committed
dc685be46622a8fabfd57cfa804237c8f15679b8
Parent: 6f1b636
Committed by GitHub <noreply@github.com>
on 5/12/2024, 5:40:45 PM