deep gemm
0ede1ba6
standardize
af716105
Merge branch 'main' into fp8-deep-gemm
11c8ca0d
clear deepgemm flow and blackwell optimization
86a2a0e7
Merge branch 'main' into fp8-deep-gemm
02075795
avoid unnecessary gathers in grouped mm
c6e7b1be
Merge branch 'main' into fp8-deep-gemm
48240d71
Merge branch 'main' into fp8-deep-gemm
b20ad183
assertions and drop the synced path
4142bac6
vasqu
commented
on 2026-03-26
use lazy load kernel
f907d4c5
style
ee5a88b5
vasqu
approved these changes
on 2026-03-30
add prefix and check for cuda runtime version
cce9786f
better names
e3df5511
exit on missing functions
f831631f
comment about why we use deepspeed cutlass single gemm
13af86ed
global statements
27cce9e5
add cuda check
8b5b2d59
Merge branch 'main' into fp8-deep-gemm
13389af9
fix deepgemm fastpath for models with bf16 scales
c50cb013
force fp32 scales in experts as well
e0170edc
vasqu
approved these changes
on 2026-03-31
Update src/transformers/integrations/finegrained_fp8.py
13a4d833
Update src/transformers/integrations/hub_kernels.py
f749fe22
fix
30132c16
style
4efc6c84
vasqu
merged
bc576731
into main 171 days ago
vasqu
deleted the fp8-deep-gemm branch 171 days ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub