arXiv · 2509.13767
VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI
Abstract
Accurate segmentation of articulatory structures in real-time MRI (rtMRI) remains challenging, as existing methods rely primarily on visual cues and overlook complementary information from synchronized speech signals. We propose VocSegMRI, a multimodal framework integrating video, audio, and phonological inputs via cross-attention fusion and a contrastive learning objective that improves cross-modal alignment and segmentation precision. Evaluated on USC-75 and further validated via zero-shot transfer on USC-TIMIT, VocSegMRI outperforms unimodal and multimodal baselines, with ablations confirming the contribution of each component.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Daiqi Liu, Johannes Enk, Maureen Stone, Fangxu Xing, Tomás Arias-Vergara, Jerry L. Prince, Jana Hutter, Jonghye Woo, Andreas Maier, Paula Andrea Pérez-Toro. 2025-09-17. VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI. https://arxiv.org/abs/2509.13767
Cite the original work for its findings. Save a collection to share your selection of sources.