text-generation-inference
Lots of improvements (Still 2 allocators)
#2449
Merged

Commits
  • Making prefix/flashinfer the default and testing the full release tests.
    Narsil committed 1 year ago
  • Include flashinfer in the docker.
    Narsil committed 1 year ago
  • Using prebuilt.
    Narsil committed 1 year ago
  • Allowing window_left_size (dummy version).
    Narsil committed 1 year ago
  • Disabling flashinfer/prefix caching on odd head_dim
    Narsil committed 1 year ago
  • Disable prefix caching for lora.
    Narsil committed 1 year ago
  • More specific codes.
    Narsil committed 1 year ago
  • Update lock
    Narsil committed 1 year ago
  • Updating integration tests with new values with FI/FD.
    Narsil committed 1 year ago
  • Update cargo lock ?
    Narsil committed 1 year ago
  • Upgrade to 1.80 because of bitstream...
    Narsil committed 1 year ago
  • Everywhere 1.80
    Narsil committed 1 year ago
  • Forgot last default place.
    Narsil committed 1 year ago
  • Apply suggestions from code review
    Narsil committed 1 year ago
  • Updated flake lock
    Narsil committed 1 year ago
  • Tmp
    Narsil committed 1 year ago
  • Upgrade resolution system for less errors in resolution.
    Narsil committed 1 year ago
  • Remove lambda for cleaner function.
    Narsil committed 1 year ago
  • Handling debugger.
    Narsil committed 1 year ago
  • OVerride the env in server tests.
    Narsil committed 1 year ago
  • Is this enough to make it work ?
    Narsil committed 1 year ago
  • This seems to be working.
    Narsil committed 1 year ago
  • Downgrade some logs.
    Narsil committed 1 year ago
  • Fixing the default for vlm.
    Narsil committed 1 year ago
  • Don't enable prefix caching on VLM just yet.
    Narsil committed 1 year ago
  • Change `add_special_tokens` in order to have the correct tokens for chat
    Narsil committed 1 year ago
  • Fixing prefix caching for flashdecoding.
    Narsil committed 1 year ago
  • Update all models.
    Narsil committed 1 year ago
  • Fixed flashinfer version.
    Narsil committed 1 year ago
  • add_special_tokens is internal only
    Narsil committed 1 year ago
  • Fixing seqlen with the new vlms.
    Narsil committed 1 year ago
  • Fixing the issue with `add_special_tokens` not being passed around.
    Narsil committed 1 year ago
  • Fixing the test.
    Narsil committed 1 year ago
  • Removing encoder_decoder (seq2seq).
    Narsil committed 1 year ago
  • Update the chat test.
    Narsil committed 1 year ago
  • Fixing the batching tokenization in flash causal lm.
    Narsil committed 1 year ago
  • Truncating left for radix purposes.
    Narsil committed 1 year ago
  • Oops this doesn't belong here.
    Narsil committed 1 year ago
  • Put back default pure shell.
    Narsil committed 1 year ago
  • Update server tests
    Narsil committed 1 year ago
  • Only n_heads / process_group.size() are necessary.
    Narsil committed 1 year ago
  • Revert the integrationt tests change (seem linked to head_size
    Narsil committed 1 year ago
  • Adding error message when assert is violated.
    Narsil committed 1 year ago
  • Fixing the free algorithm to handle times where the common prefix is
    Narsil committed 1 year ago
  • Apply suggestions from code review
    Narsil committed 1 year ago
  • Update server/text_generation_server/layers/attention/common.py
    Narsil committed 1 year ago
  • Fix disabling prefix caching - Fix windowing checks.
    Narsil committed 1 year ago
  • Revert the Cohere tokenizer change (for now using a revision instead).
    Narsil committed 1 year ago
  • Fmt.
    Narsil committed 1 year ago
Loading