arXiv · 2609.39732
Bin2Ambi: Learning Ambisonic Soundfield Reconstruction from Head-Tracked Binaural Audio
Abstract
User-generated content has become one of the most-consumed content types. However, capturing spatial audio with consumer hardware is still challenging. Given the widespread success of smart earbuds, binaural audio could be a promising option to capture spatial audio on consumer devices, but its inherent signal characteristic limits its usability as a recording format. In this paper, we propose and define a new task, Binaural to Ambisonics conversion (Bin2Ambi). In our proposed system, we exploit simultaneously captured head-tracking data provided from the motion sensors in smart earbuds. We show that this motion data help resolve the inherent directional uncertainty of two-channel binaural audio due to front-back localization ambiguities and lateral errors in the cone of confusion. Our results show that our system learns directional and diffuse-field information and that head-tracking especially reduces extreme localization errors. Objective metrics and a subjective listening test suggest that the converted Ambisonics soundfield achieves an average directional error of up to $11.8^\circ$ and a perceived spatial quality similar to a DirAC ground-truth model. The proposed algorithm can serve as a baseline for future improvements to this novel Bin2Ambi task.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gavin Milner, Nils Peters. 2026-09-30. Bin2Ambi: Learning Ambisonic Soundfield Reconstruction from Head-Tracked Binaural Audio. https://arxiv.org/abs/2609.39732
Cite the original work for its findings. Save a collection to share your selection of sources.