arXiv · 2509.13093
GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR
Abstract
End-to-end multi-talker automatic speech recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech. A critical bottleneck is that speaker-specific acoustic characteristics, which are essential for distinguishing overlapping speech, are often diluted in deep network layers. To address this, we propose the Global-Local Aware Dynamic Mixture-of-Experts (GLAD) architecture. GLAD introduces a novel routing mechanism that dynamically fuses speaker-aware global context with fine-grained local acoustic details to adaptively guide expert selection. Experiments on the LibriSpeechMix and CH109 datasets demonstrate that GLAD significantly outperforms existing Serialized Output Training (SOT)-based MTASR approaches, exhibiting exceptional robustness in challenging, high-overlap scenarios. To the best of our knowledge, this is the first work to apply a global-local fusion MoE strategy to MTASR.
Explore related subjects
Keep this discovery
Yujie Guo, Jiaming Zhou, Yuhang Jia, Shiwan Zhao, Yong Qin. 2025-09-16. GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR. https://arxiv.org/abs/2509.13093
Cite the original work for its findings. Save a collection to share your selection of sources.