arXiv · 2608.10106
BiTSE: Binaural Target Speaker Extraction in Noisy Multi-Talker Environments for AR Glass Arrays
Abstract
Isolating a desired speech signal in noisy multi-talker conversational scenarios is a key requirement for augmented reality (AR) wearable microphone array systems. In this work, a binaural target speaker extraction (TSE) framework, termed BiTSE, is proposed. It leverages both spatial and temporal cues, specifically the direction-of-arrival (DoA) of the target speaker and corresponding voice activity information, to guide the extraction process. Built upon a binaural signal denoising architecture, our model integrates three key enhancements: (i) a DoA-aware attention mechanism using cyclic positional embeddings, (ii) a timestamp-based masking strategy that utilizes speaker activity to suppress non-target segments, and (iii) a novel two-stage loss optimization strategy that first trains the model for robust denoising and then fine-tunes it to improve perceptual quality. Evaluations on the SPeech Enhancement for Augmented Reality (SPEAR) challenge dataset demonstrate that the proposed BiTSE consistently improves upon conventional approaches, leading to enhanced signal fidelity and perceptual quality.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Selani A. Indrapala, Wageesha N. Manamperi. 2026-08-10. BiTSE: Binaural Target Speaker Extraction in Noisy Multi-Talker Environments for AR Glass Arrays. https://arxiv.org/abs/2608.10106
Cite the original work for its findings. Save a collection to share your selection of sources.