onnxruntime
Use fp32 accumulation in SkipLayerNorm/EmbedLayerNorm CUDA kernels
#28682
Merged

Loading