[VLM] Support multimodal inputs for Florence-2 models #13320
port vision tower
d3e6ff72
port projection layer
953b50eb
init mm processor
86029942
fix profiling
a1a441d7
fix encoder prompt
fe93dd53
fix decoder input embeds
4379c27d
fix pos enc
fb8e629f
fix decoding
ac58e4b8
Merge branch 'vllm-project:main' into florence-2
22fef6ad
add create_decoder_prompt
d55016cd
fix prompt format
a1cfbd58
update profiling
8cd06437
make task token work
e9a2a4a2
add processor test
dc877490
Merge branch 'vllm-project:main' into florence-2
dd125cb6
fix profiling
92c0fb8f
fix dtype
d556ce93
remove hardcoded value
edfed860
add model test
608d3408
ooops
aeb1d941
fix OOM
9a1e50a0
pass test
786e3080
make test pass
ac62ffeb
reduce test
3e295958
lb
b4ddf353
code format
18c3f7a4
update doc and example
c0ef7ed1
add core model
79eb5eee
Isotr0py
marked this pull request as ready for review 1 year ago
Update vllm/model_executor/models/florence2.py
3a00624c
cleanup
d161df81
fix tokenizer mode
f11ac647
Merge branch 'upstream/main' into florence-2
ac7f9db5
remove kv_caches
07969b4e
clean up
3771ab07
Revert test for now
71c7368c
vllm-bot
merged
edf309eb
into main 1 year ago
Isotr0py
deleted the florence-2 branch 1 year ago
Assignees
No one assigned
Labels
documentation
ready
Login to write a write a comment.
Login via GitHub