arXiv · 2508.08967
Revealing the Role of Audio Channels in ASR Performance Degradation
Abstract
Pre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the input audio comes from different recording channels. While previous studies have demonstrated this phenomenon, it is often attributed to the mismatch between training and testing corpora. This study argues that variations in speech characteristics caused by different recording channels can fundamentally harm ASR performance. To address this limitation, we propose a normalization technique designed to mitigate the impact of channel variation by aligning internal feature representations in the ASR model with those derived from a clean reference channel. This approach significantly improves ASR performance on previously unseen channels and languages, highlighting its ability to generalize across channel and language differences.
Explore related subjects
Keep this discovery
Kuan-Tang Huang, Li-Wei Chen, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang. 2025-08-12. Revealing the Role of Audio Channels in ASR Performance Degradation. https://arxiv.org/abs/2508.08967
Cite the original work for its findings. Save a collection to share your selection of sources.