llama.cpp
HIP: use hipBLAS for dense prefill on gfx900, keep MMQ for MoE
#24588
Merged

HIP: use hipBLAS for dense prefill on gfx900, keep MMQ for MoE #24588

DEV-DUFORD
DEV-DUFORD HIP: keep MMQ for gfx900 MoE and Q8_0, use hipBLAS for dense K-quants
25c55df1
DEV-DUFORD HIP: tighten conditional block to be explicitly for gfx900
717920bf
DEV-DUFORD HIP: Further simplified gfx900 conditional block
76fd3124
DEV-DUFORD DEV-DUFORD requested a review 60 days ago
github-actions github-actions added Nvidia GPU
github-actions github-actions added ggml
sanmai
sanmai commented on 2026-06-15
IMbackK
IMbackK requested changes on 2026-06-15
DEV-DUFORD removed unnecessary comment
5c54430b
DEV-DUFORD
DEV-DUFORD DEV-DUFORD requested a review from IMbackK IMbackK 58 days ago
github-actions github-actions added CUDA
slavap
DEV-DUFORD
IMbackK
IMbackK approved these changes on 2026-06-25
IMbackK
DEV-DUFORD
phil2sat
JohannesGaessler
IMbackK
JohannesGaessler
JohannesGaessler approved these changes on 2026-06-30
IMbackK
JohannesGaessler JohannesGaessler merged d9df1100 into master 43 days ago

Login to write a write a comment.

Login via GitHub

Assignees
No one assigned
Labels
Milestone