fix to avoid quantizing attention with varied q,k,v sizes #9357
fix to avoid quantizing attention with varied q,k,v sizes
ccc99af4
updated the changes to address the comments
0e0768a6
viboga
merged
4771256b
into master 4 years ago
viboga
deleted the Vish/noattnquant branch 4 years ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub