transformers
ffe7a339 - tp loading: single-batch packed scatter, document NCCL transfer floor

Commit
164 days ago
tp loading: single-batch packed scatter, document NCCL transfer floor BATCH_SIZE=all mappings so each source does ONE scatter per model load instead of one per batch. Reduces scatter calls from 96 → 8 for 70B. Extensively benchmarked scatter vs all_to_all_single vs batch_isend_irecv vs CUDA IPC on 70B Llama 3.1 with 8×B200. All produce the same ~11s of NCCL transfer — the floor is NVLink bisection bandwidth with 8 ranks each sending 17.5 GB of cross-traffic, not per-op overhead. NCCL flags tested (no significant effect): NCCL_MAX_NCHANNELS=32, NCCL_PROTO=Simple, NCCL_P2P_NET_CHUNKSIZE=4M, NCCL_NCHANNELS_PER_NET_PEER=8 Results (8×B200, tp_plan=auto): Llama-3.1-70B: main 16.49s refactor 25.96s (1.57×) Qwen2.5-72B: main 17.62s refactor 26.92s (1.53×) Generate 2.5× faster on refactor (both 70B models)
Author
Parents
Loading