arXiv · 2608.01707
On Topology's Role in ML Training Performance
Abstract
Modern machine learning training workloads run on large-scale networks of compute accelerators. The networks commonly deployed in these systems are typically variations of two basic topologies: the fat-tree Clos and the torus. In this paper, we derive analytical results the elucidate how the choice of topology shapes achievable performance for the small set of collective communication operations that underlies modern machine learning workloads. We also consider how these results change when we include additional factors such as network failures and job placement strategies. Overall, we find that one topology does not dominate in all cases, but that the Clos achieves better collective completion time in most cases and provides benefits in resilience and flexibility.
Explore related subjects
Keep this discovery
Sarah McClure, Tegan Wilson, Brad Karp, Michael Mitzenmacher, Sylvia Ratnasamy, Scott Shenker, Minlan Yu. 2026-08-03. On Topology's Role in ML Training Performance. https://arxiv.org/abs/2608.01707
Cite the original work for its findings. Save a collection to share your selection of sources.