llama.cpp
d9b6be07 - ggml-cuda: provide static workspace for cuBLAS handles (#26574)

Commit
29 days ago
ggml-cuda: provide static workspace for cuBLAS handles (#26574) * provide static workspace for cuBLAS handles * account for concurrent streams when using GGML_CUDA_GRAPH_OPT * drop cublas_handle overloads and remove direct cublasSetStream calls * Update ggml/src/ggml-cuda/common.cuh --------- Co-authored-by: Oliver Simons <osimons@nvidia.com>
Author
Parents
Loading