text-generation-inference
Purely refactors paged/attention into `layers/attention` and make hardware differences more obvious with 1 file per hardware.
#1986
Merged
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Overview
Commits
19
Changes
View On
GitHub
Loading