arXiv · 2609.39138
MoSE: Mode-Switching Expander for Mixed LLM Training and Inference
Abstract
AI clusters increasingly run large language model (LLM) inference and training on the same fabric. Prefill-decode (P-D) disaggregation creates key-value (KV) cache transfers between prefill and decode groups, whereas training collectives and all-to-all traffic benefit from near-uniform global connectivity. A static sparse topology can therefore be poorly matched to one of the two traffic patterns. We present Mode-Switching Expander (MoSE), a reconfigurable expander that treats topology design as a fixed-degree edge-allocation problem. MoSE reallocates the same sparse edge budget toward direct P-D connectivity in inference-heavy modes and restores a uniform random regular expander in training-heavy modes. We evaluate MoSE using a 1024-group flow-level topology model, shortest-path routing, and two mixed workloads. Across 20 seeds, MoSE reduces average and 95th-percentile (P95) load-aware KV communication cost by 90.8\% and 91.9\% relative to Static-Training in the inference-heavy mode. In the training-heavy mode, it reduces average and P95 training communication cost by 22.7\% and 27.6\% relative to stale Static-Inference. These results show that coarse-grained topology switching can support both workload modes without additional ports or routing changes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Fan Yang, Ying Zhou, Binglei Wang, Zhenjie Zhou, Jialong Li. 2026-09-30. MoSE: Mode-Switching Expander for Mixed LLM Training and Inference. https://arxiv.org/abs/2609.39138
Cite the original work for its findings. Save a collection to share your selection of sources.