Fix Native TP expert and lm_head sharding for Qwen3-VL-MoE
8cb29882
adapt checkpoint weights layout to transformers layout with conversion
39064711
revert
8491c808
cleaner refacto
818bd0d4
cleaning
a9f6dea7
cleaner
608a7b34
3outeille
changed the title Fix Native TP expert and lm_head sharding for Qwen3-VL-MoE handle transpose during model loading with TP (Qwen3-VL-MoE)23 days ago
guarding
0ae45659
3outeillemarked this pull request as ready for review 23 days ago
Login to write a write a comment.
Login via GitHub