onnxruntime
Adding cuda kernel (optimized for sm80) for block-wise 4b quantized float 16 GEMM.
#18619
Merged

Loading