arXiv · 2609.36441
Perception-Inspired Bayesian Causal Fusion for Audiovisual Source Localization
Abstract
Multimodal fusion promises more accurate perception but only when the modalities share a common cause. When they do not, the second modality carries no information about the target, and fusing it can only corrupt the estimate. We cast this whether-to-fuse decision as Bayesian causal inference, following the optimal-observer model of human multisensory perception, and implement it as a plug-and-play layer on top of frozen audio and visual models for sound event localization and detection. The model infers a common-cause posterior over visible candidates, then gates precision-weighted fusion accordingly. Fusing unconditionally more than doubles the direction error, whereas the causal gate improves on-screen localization while limiting off-screen degradation, without any joint network retraining.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kyung Yun Lee, Sungnyun Kim, Sebastian J. Schlecht, Tae-Hyun Oh, Vesa Välimäki. 2026-09-29. Perception-Inspired Bayesian Causal Fusion for Audiovisual Source Localization. https://arxiv.org/abs/2609.36441
Cite the original work for its findings. Save a collection to share your selection of sources.