arXiv · 2310.11713
Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation
Abstract
The audio-visual sound separation field assumes visible sources in videos, but this excludes invisible sounds beyond the camera's view. Current methods struggle with such sounds lacking visible cues. This paper introduces a novel "Audio-Visual Scene-Aware Separation" (AVSA-Sep) framework. It includes a semantic parser for visible and invisible sounds and a separator for scene-informed separation. AVSA-Sep successfully separates both sound types, with joint training and cross-modal alignment enhancing effectiveness.
Explore related subjects
Keep this discovery
Yiyang Su, Ali Vosoughi, Shijian Deng, Yapeng Tian, Chenliang Xu. 2023-10-18. Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation. https://arxiv.org/abs/2310.11713
Cite the original work for its findings. Save a collection to share your selection of sources.