tp loading: single-batch packed scatter, document NCCL transfer floor
BATCH_SIZE=all mappings so each source does ONE scatter per model load
instead of one per batch. Reduces scatter calls from 96 → 8 for 70B.
Extensively benchmarked scatter vs all_to_all_single vs batch_isend_irecv
vs CUDA IPC on 70B Llama 3.1 with 8×B200. All produce the same ~11s of
NCCL transfer — the floor is NVLink bisection bandwidth with 8 ranks
each sending 17.5 GB of cross-traffic, not per-op overhead.
NCCL flags tested (no significant effect):
NCCL_MAX_NCHANNELS=32, NCCL_PROTO=Simple,
NCCL_P2P_NET_CHUNKSIZE=4M, NCCL_NCHANNELS_PER_NET_PEER=8
Results (8×B200, tp_plan=auto):
Llama-3.1-70B: main 16.49s refactor 25.96s (1.57×)
Qwen2.5-72B: main 17.62s refactor 26.92s (1.53×)
Generate 2.5× faster on refactor (both 70B models)