whisper.cpp
f8a83177
- CUDA: use mma PTX instructions for FlashAttention (llama/11583)
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Commit
View On
GitHub
Commit
1 year ago
CUDA: use mma PTX instructions for FlashAttention (llama/11583) * CUDA: use mma PTX instructions for FlashAttention * __shfl_sync workaround for movmatrix * add __shfl_sync to HIP Co-authored-by: Diego Devesa <slarengh@gmail.com>
References
#2779 - sync : ggml
Author
JohannesGaessler
Committer
ggerganov
Parents
85451e36
Loading