text-generation-inference
Lots of improvements (Still 2 allocators)
#2449
Merged
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Overview
Commits
49
Changes
View On
GitHub
Commits
Making prefix/flashinfer the default and testing the full release tests.
Narsil
committed
1 year ago
Include flashinfer in the docker.
Narsil
committed
1 year ago
Using prebuilt.
Narsil
committed
1 year ago
Allowing window_left_size (dummy version).
Narsil
committed
1 year ago
Disabling flashinfer/prefix caching on odd head_dim
Narsil
committed
1 year ago
Disable prefix caching for lora.
Narsil
committed
1 year ago
More specific codes.
Narsil
committed
1 year ago
Update lock
Narsil
committed
1 year ago
Updating integration tests with new values with FI/FD.
Narsil
committed
1 year ago
Update cargo lock ?
Narsil
committed
1 year ago
Upgrade to 1.80 because of bitstream...
Narsil
committed
1 year ago
Everywhere 1.80
Narsil
committed
1 year ago
Forgot last default place.
Narsil
committed
1 year ago
Apply suggestions from code review
Narsil
committed
1 year ago
Updated flake lock
Narsil
committed
1 year ago
Tmp
Narsil
committed
1 year ago
Upgrade resolution system for less errors in resolution.
Narsil
committed
1 year ago
Remove lambda for cleaner function.
Narsil
committed
1 year ago
Handling debugger.
Narsil
committed
1 year ago
OVerride the env in server tests.
Narsil
committed
1 year ago
Is this enough to make it work ?
Narsil
committed
1 year ago
This seems to be working.
Narsil
committed
1 year ago
Downgrade some logs.
Narsil
committed
1 year ago
Fixing the default for vlm.
Narsil
committed
1 year ago
Don't enable prefix caching on VLM just yet.
Narsil
committed
1 year ago
Change `add_special_tokens` in order to have the correct tokens for chat
Narsil
committed
1 year ago
Fixing prefix caching for flashdecoding.
Narsil
committed
1 year ago
Update all models.
Narsil
committed
1 year ago
Fixed flashinfer version.
Narsil
committed
1 year ago
add_special_tokens is internal only
Narsil
committed
1 year ago
Fixing seqlen with the new vlms.
Narsil
committed
1 year ago
Fixing the issue with `add_special_tokens` not being passed around.
Narsil
committed
1 year ago
Fixing the test.
Narsil
committed
1 year ago
Removing encoder_decoder (seq2seq).
Narsil
committed
1 year ago
Update the chat test.
Narsil
committed
1 year ago
Fixing the batching tokenization in flash causal lm.
Narsil
committed
1 year ago
Truncating left for radix purposes.
Narsil
committed
1 year ago
Oops this doesn't belong here.
Narsil
committed
1 year ago
Put back default pure shell.
Narsil
committed
1 year ago
Update server tests
Narsil
committed
1 year ago
Only n_heads / process_group.size() are necessary.
Narsil
committed
1 year ago
Revert the integrationt tests change (seem linked to head_size
Narsil
committed
1 year ago
Adding error message when assert is violated.
Narsil
committed
1 year ago
Fixing the free algorithm to handle times where the common prefix is
Narsil
committed
1 year ago
Apply suggestions from code review
Narsil
committed
1 year ago
Update server/text_generation_server/layers/attention/common.py
Narsil
committed
1 year ago
Fix disabling prefix caching - Fix windowing checks.
Narsil
committed
1 year ago
Revert the Cohere tokenizer change (for now using a revision instead).
Narsil
committed
1 year ago
Fmt.
Narsil
committed
1 year ago
Loading