onnxruntime
c379a89b - [MLAS AArch64] SQNBitGemm optimization (#19272)

Commit
2 years ago
[MLAS AArch64] SQNBitGemm optimization (#19272) 1. Add support for packing 4-bit values 32 at a time for CompInt8. 32 4-bit values can fit into a single 128-bit NEON register. For CompInt8, this enables a more efficient path for block sizes greater than or equal to 32. CompFp32 seems to do better with handling 16 elements at a time, so this 32-value packing is not used there. Pack differently based on compute type. Adjust APIs to handle this. 2. Introduce template argument for whether to handle zero-point. This results in less code for the no zero-point (symmetric) case. However, there is a binary size increase due to the additional template instantiations.
Author
Parents
Loading