Add Triton grouped-GEMM for MoE experts on Ampere/Ada #8180
delock
commented
on 2026-07-28
delock
commented
on 2026-07-28
delock
commented
on 2026-07-28
delock
commented
on 2026-07-28
A good working version for group gemm
123c0309
Refactor code
e7b962dd
change autotune config
a815d4cb
Update comments
0699071e
Add fused triton kernel for some elementwise kernel
ab3f2392
Add benchmark code
f09c1da9
Fix format error, and add doc
584236a0
Fix format error, address reviewer comments
ca2090ba
hwchen2017
force pushed
from
264b3990
to
ca2090ba
33 days ago
Fix review comments
ed93495e
delock
approved these changes
on 2026-07-31
Replace disable_triton_group_mm env variabel with config
58ddff65
Skip CUDA forward tests when transformer op is unavailable
51aa775a
Revert "Skip CUDA forward tests when transformer op is unavailable"
1cca5431
hwchen2017
force pushed
from
fd3a792e
to
1cca5431
31 days ago
tohtana
requested changes
on 2026-08-01
Address review comments
5ac1ad3a
tohtana
approved these changes
on 2026-08-03
hwchen2017
merged
df84f6d8
into master 28 days ago
hwchen2017
deleted the hongwei/group_gemm branch 28 days ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub