Searcharxiv⌕ Search

arXiv · 2609.31898

MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus

Abstract

Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Speech-TRacking Open (MAESTRO) corpus, the first AAD dataset to simultaneously record EEG, eye gaze, pupillometry, egocentric video, and head inertial measurement unit (IMU) data. MAESTRO includes four competing speakers and background noise across multiple signal-to-noise ratio (SNR) conditions, enabling attention decoding under realistic listening scenarios. Through a four-speaker attention decoding benchmark, we show that combining behavioral and physiological signals improves decoding performance over EEG-only approaches, enabling future advances in multimodal auditory attention decoding. These findings open the door to new applications, analyses, and methodological advances in multimodal AAD. The complete dataset is publicly available at https://huggingface.co/datasets/aspire-osu/maestro-eeg-dataset . The official code repository is available at https://github.com/ASPIRE-OSU/MAESTRO .

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

K M Naimul Hassan, Ali Alavi, Donald S. Williamson. 2026-09-25. MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus. https://arxiv.org/abs/2609.31898

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Teacher-Free Self-Distilled Consistency Trajectory Learning for Fast Speech Enhancement

Consistency trajectory models offer a route to fast, high- quality speech enhancement, collapsing the many reverse steps of diffusion-based enhancers into a handful. When instantiated on a Schrödinger bridge (SB), which pins the generative process to fixed clean and noisy endpoints, exist- ing consistency-trajectory enhancers (SBCTMs) still require a pretrained teacher to supply trajectory supervision, which raises training cost and ties the final quality to that of the teacher. We propose a teacher-free, self-distilled consistency- trajectory framework that removes the external teacher result- ing in a 5X reduction in per epoch training time. Our model is trained with a three-stage curriculum of clean speech pre- diction, a self-distilled shortcut objective, and perceptual fine-tuning with a multi-resolution short-time Fourier trans- form (MR-STFT) loss. Using the same NCSN++ backbone as SBCTM, our model attains a wide-band PESQ of 3.01, ES- TOI 0.87, and SI-SDR 19.07 dB on VoiceBank+DEMAND compared to 3.57, 0.87 and 12.8 dB for the teacher based model. Further, we find that a geometric schedule at low reverse step count maximizes perceptual quality, while a higher-step uniform schedule favors signal fidelity.

eess.AS↗

Perception-Inspired Bayesian Causal Fusion for Audiovisual Source Localization

Multimodal fusion promises more accurate perception but only when the modalities share a common cause. When they do not, the second modality carries no information about the target, and fusing it can only corrupt the estimate. We cast this whether-to-fuse decision as Bayesian causal inference, following the optimal-observer model of human multisensory perception, and implement it as a plug-and-play layer on top of frozen audio and visual models for sound event localization and detection. The model infers a common-cause posterior over visible candidates, then gates precision-weighted fusion accordingly. Fusing unconditionally more than doubles the direction error, whereas the causal gate improves on-screen localization while limiting off-screen degradation, without any joint network retraining.

eess.AS↗

Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT

Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBERT backbone and a 64-dimensional representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt. We compare a frozen-backbone recognizer (Baseline), trainable fusion with zero auxiliary input (Null), and the same fusion supplied with the learned representation (Learned). Across Buckeye, Switchboard, and AMI IHM, Null reduces WER by 0.71-1.45 points over Baseline, whereas Learned differs from Null by +0.07, -0.09, and +0.00 points, with no significant differences. However, removing or mismatching the representation at inference increases Learned WER. Thus, Learned depends on the representation yet shows no measurable incremental WER benefit over the parameter-matched control.

eess.AS↗