arXiv · 2407.18571
Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks
Abstract
Speech bandwidth expansion is crucial for expanding the frequency range of low-bandwidth speech signals, thereby improving audio quality, clarity and perceptibility in digital applications. Its applications span telephony, compression, text-to-speech synthesis, and speech recognition. This paper presents a novel approach using a high-fidelity generative adversarial network, unlike cascaded systems, our system is trained end-to-end on paired narrowband and wideband speech signals. Our method integrates various bandwidth upsampling ratios into a single unified model specifically designed for speech bandwidth expansion applications. Our approach exhibits robust performance across various bandwidth expansion factors, including those not encountered during training, demonstrating zero-shot capability. To the best of our knowledge, this is the first work to showcase this capability. The experimental results demonstrate that our method outperforms previous end-to-end approaches, as well as interpolation and traditional techniques, showcasing its effectiveness in practical speech enhancement applications.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mahmoud Salhab, Haidar Harmanani. 2024-07-26. Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks. https://arxiv.org/abs/2407.18571
Cite the original work for its findings. Save a collection to share your selection of sources.