onnxruntime
ce0025d3 - Fallback Pow op in layer norm to FP32 in TRT to avoid overflow (#13639)

Commit
3 years ago
Fallback Pow op in layer norm to FP32 in TRT to avoid overflow (#13639) Accuracy loss is observed when transformer models such as BERT, DeBERTa, ViT are running in TRT FP16 mode. The cause is that overflow happens at Pow op in layer norm. This PR provides the option to force Pow to run in TRT FP32 precision if overflow occurs. Co-authored-by: Ubuntu <azureuser@orteplinuxdev.bxgbzpva45kedp3rhbsbit4phb.jx.internal.cloudapp.net>
Author
Parents
Loading