text-generation-inference
Purely refactors paged/attention into `layers/attention` and make hardware differences more obvious with 1 file per hardware.
#1986
Merged

Purely refactors paged/attention into `layers/attention` and make hardware differences more obvious with 1 file per hardware. #1986

Narsil merged 19 commits into main from rearchitecture_attention_code
Narsil
Narsil Using flash decoding
4fd3065d
Narsil Fix after rebase..
8171747e
Narsil Less intrusive.
be8c14be
Narsil REvert changes in modeling.
ed96a76d
Narsil Speedup flashdecoding.
6bbc8430
Narsil HHachweew
6aeb5a73
Narsil Fixing non flash decoding llama path.
7a29e826
Narsil Router logic knows about page size.
50d5c08b
Narsil Missing cohere.
a6f16035
Narsil Fixing cohere flash decoding.
7890cd66
Narsil Revamped all this architecture.
daddd2e9
Narsil Fix cohere.
a76e6502
Narsil Fixing falcon.
cf595934
Narsil Enabling custom block size schedule.
13caf958
Narsil Update router/src/infer.rs
be87c840
Narsil Removing flash decoding part so it gets merged.
91f55ea2
Narsil
Narsil commented on 2024-05-31
Narsil Update server/text_generation_server/utils/import_utils.py
c67539fb
danieldk
danieldk commented on 2024-05-31
Narsil Adress comments + fix 2nd path in falcon.
d44688b6
danieldk
danieldk dismissed these changes on 2024-05-31
Narsil
Narsil commented on 2024-05-31
Narsil Update server/text_generation_server/layers/attention/xpu.py
b0c168d2
Narsil Narsil dismissed their stale review via b0c168d2 2 years ago
Narsil Narsil merged 06edde94 into main 2 years ago
Narsil Narsil deleted the rearchitecture_attention_code branch 2 years ago

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone