[Kernel] Enable Hybrid Model Support in Triton Unified Attention Kernel #21197
modify shape of kv cache
c3f8dd47
removed prefill support from split-kv attention
7d52de7e
added reorder_batch method
41a12543
added tiling to support large non-power-of-2 block sizes
6109a140
formatting
343ed93c
updated parameters
a96e2295
resolved conflicts
ffa29db3
removed unneeded files
6a375ee4
Merge branch 'vllm-project:main' into jvl-hybrid-support
39e352d5
add prefill support back to split-kv kernel
c4f7d677
removed changes to triton_attn.py
b4fd773d
Merge branch 'vllm-project:main' into jvl-hybrid-support
40218f46
restored changes to triton_attn.py
832ccb8d
formatting
d893bcb6
Merge branch 'main' into jvl-hybrid-support
02b04645
formatting
9b63155d
Merge branch 'main' into jvl-hybrid-support
d3979e9c
Merge branch 'main' into jvl-hybrid-support
771f5246
Merge branch 'main' into jvl-hybrid-support
ec62011b
temporarily revert changes to enable simpler merge with latest main
39ed6b6e
Merge branch 'vllm-project:main' into jvl-hybrid-support
eead28ea
Merge branch 'vllm-project:main' into jvl-hybrid-support
2b415b65
restore changes to triton_unified_attention.py
4254fad7
Merge branch 'main' into jvl-hybrid-support
db615014
Merge branch 'main' into jvl-hybrid-support
3605c65b
replaced block size check by tile size check
8a846bcf
Merge branch 'main' into jvl-hybrid-support
f41f90e2
assign default tile sizes for prefill and decode, remove check
3dedd209
Merge branch 'main' into jvl-hybrid-support
ee82dd53
tdoublep
approved these changes
on 2025-09-18
tdoublep
enabled auto-merge (squash) 317 days ago
Merge branch 'main' into jvl-hybrid-support
8da5f743
tdoublep
merged
01a583fe
into main 317 days ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub