llama.cpp
ggml-cuda : disable MMQ on devices with less than 48 KiB shared memory
#26141
Merged

ggml-cuda : disable MMQ on devices with less than 48 KiB shared memory #26141

KakaruHayate
KakaruHayate KakaruHayate requested a review 48 days ago
github-actions github-actions added ggml
github-actions github-actions added CUDA
ggml-gh-bot
KakaruHayate KakaruHayate changed the title CUDA/MUSA: fall back to cuB when no MMQ tile size fits in smem CUDA/MUSA: fall back to cuBlass when no MMQ tile size fits in smem 47 days ago
KakaruHayate KakaruHayate changed the title CUDA/MUSA: fall back to cuBlass when no MMQ tile size fits in smem ggml-cuda : fall back to cuBLAS when no MMQ tile size fits in shared memory 47 days ago
KakaruHayate KakaruHayate force pushed 47 days ago
am17an
KakaruHayate
am17an
KakaruHayate
am17an
KakaruHayate
KakaruHayate KakaruHayate force pushed 45 days ago
KakaruHayate KakaruHayate changed the title ggml-cuda : fall back to cuBLAS when no MMQ tile size fits in shared memory ggml-cuda : disable MMQ on devices with less than 48 KiB shared memory 45 days ago
KakaruHayate ggml-cuda : disable MMQ on devices with less than 48 KiB shared memory
0e27a00c
KakaruHayate KakaruHayate marked this pull request as draft 45 days ago
KakaruHayate KakaruHayate force pushed to 0e27a00c 45 days ago
KakaruHayate KakaruHayate marked this pull request as ready for review 45 days ago
am17an
am17an approved these changes on 2026-07-29
ORippler
ORippler approved these changes on 2026-07-29
am17an am17an merged caa596ab into master 45 days ago

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone