Searcharxiv⌕ Search

arXiv · 2609.30774

Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information

Abstract

In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of real conversations. We therefore introduce, to our knowledge, the first benchmark for online audio-visual TSE (AV-TSE), built from intact dyadic interactions with independent third-party interference. Observing that anticipating upcoming activity from semantic, acoustic, and facial cues benefits online AV-TSE, we propose an LLM-based target-speaker voice activity projection (TS-VAP) module. Unlike conventional VAP with separated speaker channels, it forecasts the future activity of the target and conversational partner directly from the overlapping mixture, drawing on the linguistic and conversational knowledge of a speech-LLM, and uses this prediction to guide a low-latency separator. We further combine this predictive context with historical and synchronous speaker context. Experiments show that TS-VAP consistently improves streaming extraction across multiple AV-TSE backbones, and that further combining historical, synchronous, and predictive context yields nearly 1 dB gain on real AV conversations. Project page: https://jjjjiaozi.github.io/TS-VAP/.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Shuhan Zhang, Wenxuan Wu, Haizhou Li. 2026-09-25. Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information. https://arxiv.org/abs/2609.30774

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-level feature spaces, we propose Audio-JEPA (Audio Joint-Embedding Predictive Architecture), tailored specifically for audio data. Audio-JEPA uses a simple Vision Transformer backbone to predict latent representations of masked spectrogram patches rather than reconstructing raw audio. We pre-train on unlabeled AudioSet clips (10s, 32kHz) with random patch masking on mel-spectrograms. We evaluate on the X-ARES suite covering speech, music, and environmental sound tasks. Although our implementation is a straightforward translation of the original model to audio, the results still show comparable performance to wav2vec 2.0 and data2vec while using less than one-fifth of their training data and with no hyper-parameter tuning. All code and pretrained checkpoints are available on GitHub.

cs.SD↗

Channel-Preserving Representation Alignment for EEG-to-Music Reconstruction

Reconstructing naturalistic music from EEG requires extracting music-specific information from weak signals distributed across electrodes. Existing spectrogram regression and diffusion approaches provide audio synthesis mechanisms, but learning the EEG-to-music mapping faces an information--estimation tradeoff. Combining electrode measurements into fewer features can ease estimation from limited recordings while obscuring distinctions between music segments. We propose channel-preserving representation alignment, separating electrode information retention from prediction regularization. Per-electrode tokenization preserves spatial identity for learned aggregation, while multi-view self-distillation and channel dropout encourage consistency across electrode subsets during pretraining and music alignment. The aligned representation conditions a pretrained audio generator. Our theory explains these complementary choices. A finite-sample regression analysis identifies when retained task information outweighs estimation cost. For nonlinear alignment, an exact decomposition reveals a prediction-consistency penalty in the masked contrastive objective, with margin-dependent bounds connecting prediction variability to ranking errors across electrode subsets. On NMED-T/H, our method improves within-subject 50-way identification from 0.388 to 0.485 and generated-audio CLAP similarity from 0.635 to 0.683 compared with previous work.

cs.SD↗

Patient-level validation of a foundation-model pipeline for pediatric lung sounds: physician-annotated adventitious events are recognized, disease groups are not reliably predicted

Background: Lung-sound classifiers are usually evaluated on recording- or event-level splits, although each child contributes many recordings. We evaluated a foundation-model pipeline (PulmoVec) with the patient as the unit of partitioning and asked whether disease group can be predicted beyond age and sex. Methods: We analyzed 19693 physician-annotated respiratory events from 736 children in the public SPRSound database. A frozen Health Acoustic Representations (HeAR) encoder with low-rank adapters was trained for screening (normal versus adventitious), sound pattern (normal, crackles, wheeze/rhonchi) and disease group (pneumonia, bronchial disease, normal/other); event probabilities were stacked with age, sex and auscultation site. The primary analysis was nested patient-grouped cross-validation; comparators were the majority class, annotated event duration and demographics. Results: The area under the receiver operating characteristic curve (AUC) was 0.95 (95% CI 0.94 to 0.96) for both screening and sound pattern, against 0.78 and 0.75 for event duration alone; the positive predictive value for adventitious events was 0.72. Disease-group prediction reached an AUC of 0.58 (95% CI 0.54 to 0.62), and its accuracy was below the majority-class rate at event level (0.44 versus 0.62) and at patient level (0.46 versus 0.60). Conclusions: Recognition of physician-annotated adventitious events holds under patient-level validation and is not explained by event duration or demographics, although its positive predictive value would limit unaided use. Prediction of disease group, a clinical diagnosis, from lung-sound events alone was not demonstrated. Lung-sound studies should partition data by patient and report non-acoustic comparators.

cs.SD↗