arXiv · 2504.02988
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
Abstract
We present SELDVisualSynth, a tool for generating synthetic videos for audio-visual sound event localization and detection (SELD). Our approach incorporates real-world background images to improve realism in synthetic audio-visual SELD data while also ensuring audio-visual spatial alignment. The tool creates 360 synthetic videos where objects move matching synthetic SELD audio data and its annotations. Experimental results demonstrate that a model trained with this data attains performance gains across multiple metrics, achieving superior localization recall (56.4 LR) and competitive localization error (21.9deg LE). We open-source our data generation tool for maximal use by members of the SELD research community.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Adrian S. Roman, Aiden Chang, Gerardo Meza, Iran R. Roman. 2025-04-03. Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection. https://arxiv.org/abs/2504.02988
Cite the original work for its findings. Save a collection to share your selection of sources.