[CUDA] Speed up 8-bit MatMulNBits dequantization with byte permutes #31350
[CUDA] Speed up 8-bit MatMulNBits dequantization with byte permutes
972f27c1
[CUDA] Use safe bit casts in 8-bit dequantization
0405b748
tianleiwu
enabled auto-merge (squash) 7 days ago
tianleiwu
merged
5b249f65
into main 4 days ago
tianleiwu
deleted the tlwu/20260801/matmul_8bits_fast_dequant branch 4 days ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub