llama.cpp
CUDA: add bf16 and f32 support to cublas_mul_mat_batched
#14361
Merged

CUDA: add bf16 and f32 support to cublas_mul_mat_batched #14361

am17an
am17an am17an marked this pull request as ready for review 1 year ago
am17an am17an requested a review from JohannesGaessler JohannesGaessler 1 year ago
github-actions github-actions added testing
github-actions github-actions added Nvidia GPU
github-actions github-actions added ggml
am17an am17an force pushed 1 year ago
JohannesGaessler
JohannesGaessler commented on 2025-06-24
am17an am17an force pushed 1 year ago
am17an am17an requested a review from JohannesGaessler JohannesGaessler 1 year ago
am17an
JohannesGaessler
JohannesGaessler commented on 2025-06-25
am17an am17an force pushed 1 year ago
am17an am17an requested a review from JohannesGaessler JohannesGaessler 1 year ago
JohannesGaessler
JohannesGaessler approved these changes on 2025-06-25
am17an
JohannesGaessler
jeffbolznv
jeffbolznv
jeffbolznv
am17an CUDA: add bf16 and f32 support to cublas_mul_mat_batched
4887fa57
am17an Review: add type traits and make function more generic
526fc4e0
am17an Review: make check more explicit, add back comments, and fix formatting
c02cd2fc
am17an Review: fix formatting, remove useless type conversion, fix naming fo…
2c4e42ed
am17an am17an force pushed to 2c4e42ed 1 year ago
am17an am17an merged 27208bf6 into master 1 year ago
am17an am17an deleted the add_bp16_fp32_to_cublas_batched branch 347 days ago

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone