arXiv · 2609.31041
Room Impulse Response Embeddings for Speech Enhancement in Noisy and Reverberant Environments
Abstract
We propose a self-supervised approach for learning room impulse response (RIR) representations from single-channel noisy-reverberant speech. It consists of first training on reverberant data, then on noisy-reverberant data, and finally with a teacher-student approach, where the student learns to replicate the teacher's embeddings when given a noisy version of the reverberant input. We assess their representational capabilities by estimating acoustic room parameters from them. Conditioning a discriminative speech enhancement model on the derived embeddings yields consistent gains across all evaluated metrics, including downstream word error rate, for both reverberant and noisy-reverberant speech.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Adrian Meise, Reinhold Haeb-Umbach. 2026-09-25. Room Impulse Response Embeddings for Speech Enhancement in Noisy and Reverberant Environments. https://arxiv.org/abs/2609.31041
Cite the original work for its findings. Save a collection to share your selection of sources.