text-generation-inference
Purely refactors paged/attention into `layers/attention` and make hardware differences more obvious with 1 file per hardware.
#1986
Merged

Loading