[`GPTNeoX`] Flex Attention + Refactor #34896
gpt neox flex attention + refactor
65da2538
some formatting
ebfdc1f9
vasqu
commented
on 2024-11-23
small fix on dropout
b3c2b11b
add assertion on flex attn test
06685f82
flaky ci :(
53f73190
add head mask support
21941c5c
style
02089056
handle dtype, replace torch where
e689d28c
fixup flex with output attns
7a444f5f
code review and several other fixes
bd22b3d7
Update src/transformers/modeling_utils.py
94ae1949
style
4c8c9a2c
vasqu
commented
on 2024-11-28
remove unnecessary comment
96f85b04
remove incorrect comment
086d0a5e
make flex attn check more agnostic tor versions and centralized
bb3fb239
change peft input dtype check to value since q and k could be affecte…
5e6d904f
i forgor
e51aa860
flaky
1a326c1a
vasqu
commented
on 2024-11-29
code review and small fixes
e8094285
Update src/transformers/models/gpt_neox/modeling_gpt_neox.py
05f51fec
vasqu
commented
on 2024-12-02
vasqu
deleted the flex-gptneox branch 1 year ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub