arXiv · 2608.24579
Visually-Guided Spatial Audio Generation for $360^\circ$ In-the-Wild Speech Scenes
Abstract
Spatial audio is a key component of immersive $360^\circ$ media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned $360^\circ$ video and an omnidirectional audio track, we recover the missing directional FOA components. To support this task, we introduce YT-SPEECH, a speech-oriented $360^\circ$ video-FOA dataset curated from YouTube. We propose a two-stage Localizer-Renderer framework, where an audio-visual segmentation backbone provides frame-wise spatial heatmaps and a conditional complex-domain U-Net reconstructs directional FOA signals from the omnidirectional channel. A confidence-based gating strategy stabilizes conditioning under ambiguous acoustic conditions. Experiments show improved reconstruction fidelity, spatial accuracy, and perceptual speech quality relative to ablated variants and prior approaches.
Explore related subjects
Keep this discovery
Qingyu Luo, Peng Zhang, Wenwu Wang, Philip J. B. Jackson. 2026-08-25. Visually-Guided Spatial Audio Generation for $360^\circ$ In-the-Wild Speech Scenes. https://arxiv.org/abs/2608.24579
Cite the original work for its findings. Save a collection to share your selection of sources.