arXiv · 2608.15118
Collective Communication for Distributed LLM Systems: Planning, Runtime Adaptation, and Computation Coordination
Abstract
Distributed large language model (LLM) systems increasingly rely on collective communication primitives such as AllReduce (AR), ReduceScatter (RS), AllGather (AG), and AlltoAll (A2A). In modern LLM training and serving clusters, heterogeneous GPU interconnects, multi-NIC networking, mixed parallelism strategies, low-latency inference requests, and high-throughput training pipelines have motivated increasingly diverse ways to plan, execute, and overlap collective communication. This paper presents a tutorial-style, collective-centric taxonomy for collective communication. We organize recent advances into three layers: communication planning, which generates topology-aware collective schedules; communication execution and adaptation, which maps these schedules onto GPU runtimes and hardware in real clusters; and computation-communication coordination, which turns collective optimization into end-to-end training and inference benefits. We further discuss open challenges and future opportunities for collective communication in distributed LLM systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xuebin Song, Menghao Zhang, Yuezheng Liu, Jinyi Xia, Shucan Yang, Xiaohe Hu, Chunming Hu, Mingwei Xu. 2026-08-15. Collective Communication for Distributed LLM Systems: Planning, Runtime Adaptation, and Computation Coordination. https://arxiv.org/abs/2608.15118
Cite the original work for its findings. Save a collection to share your selection of sources.