Purely refactors paged/attention into `layers/attention` and make hardware differences more obvious with 1 file per hardware. #1986
Using flash decoding
4fd3065d
Fix after rebase..
8171747e
Less intrusive.
be8c14be
REvert changes in modeling.
ed96a76d
Speedup flashdecoding.
6bbc8430
HHachweew
6aeb5a73
Fixing non flash decoding llama path.
7a29e826
Router logic knows about page size.
50d5c08b
Missing cohere.
a6f16035
Fixing cohere flash decoding.
7890cd66
Revamped all this architecture.
daddd2e9
Fix cohere.
a76e6502
Fixing falcon.
cf595934
Enabling custom block size schedule.
13caf958
Update router/src/infer.rs
be87c840
Removing flash decoding part so it gets merged.
91f55ea2
Narsil
commented
on 2024-05-31
Update server/text_generation_server/utils/import_utils.py
c67539fb
Adress comments + fix 2nd path in falcon.
d44688b6
danieldk
dismissed these changes
on 2024-05-31
Narsil
commented
on 2024-05-31
Update server/text_generation_server/layers/attention/xpu.py
b0c168d2
Narsil
dismissed their stale review
via b0c168d2
2 years ago
Narsil
merged
06edde94
into main 2 years ago
Narsil
deleted the rearchitecture_attention_code branch 2 years ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub