whisper.cpp
f8a83177 - CUDA: use mma PTX instructions for FlashAttention (llama/11583)

Commit
1 year ago
CUDA: use mma PTX instructions for FlashAttention (llama/11583) * CUDA: use mma PTX instructions for FlashAttention * __shfl_sync workaround for movmatrix * add __shfl_sync to HIP Co-authored-by: Diego Devesa <slarengh@gmail.com>
Committer
Parents
Loading