vllm
[Model] Reduce redundant computations in mamba2 blocks for Bamba-9B
#15423
Merged

[Model] Reduce redundant computations in mamba2 blocks for Bamba-9B #15423

cyang49
github-actions
fabianlim
fabianlim commented on 2025-03-25
fabianlim
fabianlim commented on 2025-03-25
fabianlim
fabianlim commented on 2025-03-25
yury-tokpanov
fabianlim
cyang49 cyang49 force pushed to 4c672ff2 1 year ago
tlrmchlsmth
tlrmchlsmth commented on 2025-03-25
fabianlim
yury-tokpanov
yury-tokpanov commented on 2025-03-25
cyang49 cyang49 force pushed from 9b3705b3 1 year ago
tlrmchlsmth
tlrmchlsmth commented on 2025-03-31
cyang49 cyang49 force pushed 1 year ago
cyang49 cyang49 force pushed 1 year ago
cyang49 cyang49 force pushed 1 year ago
cyang49
tlrmchlsmth
cyang49 Reduce cpu-gpu synchronization in mamba2 chunked prefill
2be3f161
cyang49 Reduce redundancy in mamba2 blocks
08df5615
fabianlim refactored mixer and bamba
a7b363cd
cyang49 lint
e06a170e
cyang49 disable assertions that causes d2h copies in mamba2
678e6910
cyang49 simplify seq_idx computation
21ebc660
cyang49 Revert "simplify seq_idx computation" as repeat_interleave
569185e8
cyang49 remove assertions that cause synchronizations
01d5a332
cyang49 Pack mamba2 metadata
017597e9
cyang49 pack more into mamba2 metadata
c0ead4d3
cyang49 Patching mamba2 and zamba2
4382192e
cyang49 cyang49 force pushed to 4382192e 1 year ago
cyang49 Add back torch.any on has_inital_states to recover original semantics
080eee92
cyang49 moving more redundant computation into metadata prep
7236a55b
cyang49
cyang49
cyang49 cyang49 marked this pull request as ready for review 1 year ago
tlrmchlsmth
tlrmchlsmth approved these changes on 2025-04-10
tlrmchlsmth tlrmchlsmth added ready
tlrmchlsmth tlrmchlsmth enabled auto-merge (squash) 1 year ago
tlrmchlsmth tlrmchlsmth merged daefed05 into main 1 year ago

Login to write a write a comment.

Login via GitHub

Assignees
No one assigned
Labels
Milestone