Optimized HIP Q2_0 dot-product path for gfx1201 by replacing HIP's byte permutation with native permutation instructions. #26753
hip/gfx1201: optimize q2_0 vec_dot_q2_0_q8_1 with native amdgcn perm
c43d584b
Broadened HIP's Q2_0 perm optimization
857d3378
Remove redundant HIP perm availability guard
ae5a832a
Optimize HIP Q2_0 MMQ unpack with native perm
8963530a
IMbackK
approved these changes
on 2026-08-26
cuda: label HIP preprocessor guard
48338513
cuda: label HIP preprocessor guard
480d7b37
Restore MMQ tile index handling
f66bcd34
IMbackK
approved these changes
on 2026-08-29
ggerganov
merged
e4221480
into master 2 days ago
Labels
ggml
merge ready
CUDA
Login to write a write a comment.
Login via GitHub