text-generation-inference
6f15ac60 - feat: support force downcast after FastRMSNorm multiply for Gemma (#1658)

Commit

2 years ago

feat: support force downcast after FastRMSNorm multiply for Gemma (#1658) This PR adds `force_downcast_after` to `FastRMSNorm.forward` which is used in the Gemma model. References https://github.com/huggingface/transformers/pull/29402 and https://github.com/huggingface/transformers/pull/29729 Setting `force_downcast_after=True` will perform the `hidden_states * weight` multiplication in f32 and then downcast to half. This differs slightly from the current implementation which first casts the `hidden_states` to a half and then multiples.

References

#1658 - feat: support force downcast after FastRMSNorm multiply for Gemma

Author

drbh

Parents

dfbd9a39

text-generation-inference 6f15ac60 - feat: support force downcast after FastRMSNorm multiply for Gemma (#1658)

text-generation-inference
6f15ac60 - feat: support force downcast after FastRMSNorm multiply for Gemma (#1658)