arXiv · 2601.17606
Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
Abstract
Performant all-to-all collective operations in MPI are critical to fast Fourier transforms, transposition, and machine learning applications. There are many existing implementations for all-to-all exchanges on emerging systems, with the achieved performance dependent on many factors, including message size, process count, architecture, and parallel system partition. This paper presents novel all-to-all algorithms for emerging many-core systems. Further, the paper presents a performance analysis against existing algorithms and system MPI, with novel algorithms achieving up to 3x speedup over system MPI at 32 nodes of state-of-the-art Sapphire Rapids systems.
Explore related subjects
Keep this discovery
Shannon Kinkead, Jackson Wesley, Whit Schonbein, David DeBonis, Matthew G. F. Dosanjh, Amanda Bienz. 2026-01-24. Scaling All-to-all Operations Across Emerging Many-Core Supercomputers. https://arxiv.org/abs/2601.17606
Cite the original work for its findings. Save a collection to share your selection of sources.