arXiv · 2609.36724
Routing in Gradient Space: Balanced Usage Is Not Expert Specialization
Abstract
Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping observations with aligned gradients. On five multi-task text-classification mixtures, we compare GAR with task-loss-only routing, gradient-combination and gradient-conflict methods, and load-balancing losses. With a fully trainable RoBERTa backbone and classification-head experts, GAR has the highest aggregate validation accuracy, 1.07 percentage points above task-loss-only routing. With frozen DeBERTa and Qwen3-1.7B backbones and low-rank adapter experts, it again ranks first, 1.10 points above task-loss-only routing, with better-balanced expert load and higher gradient-mass purity, the share of each expert's gradient-norm mass from its dominant task; the load-balancing losses flatten load further but leave this purity near its task-loss-only level. Top-1 routing, trainable full-parameter feed-forward network (FFN) experts, and a larger backbone also show positive aggregate gains. The results distinguish expert-load balance from gradient-based routing organization and indicate the predictive value of gradient-informed routing in multi-task text classification.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuchen Li, Mingyu Du, Zongqi Fan, Nguyen H. Tran, Ken-Tye Yong. 2026-09-29. Routing in Gradient Space: Balanced Usage Is Not Expert Specialization. https://arxiv.org/abs/2609.36724
Cite the original work for its findings. Save a collection to share your selection of sources.