arXiv · 2610.11507
Monaural Continuous-Radius Regional Speech Extraction with Cross-Radius Consistency Learning
Abstract
Source-to-microphone distance enables speaker-independent speech extraction without prior enrollment. We propose a monaural continuous-radius regional speech extraction method that directly models the cumulative speech target within a queried radius. To realize continuous region control, the query radius is encoded as a continuous scalar and injected into a time-frequency extraction network, allowing a single model to operate over arbitrary radii within the trained range. Exploiting the nested structure of target-speaker sets across query radii, we introduce cross-radius consistency learning to stabilize predictions for adjacent radii sharing the same nonempty target-speaker set. Experiments on measured RIRs show that continuous-radius conditioning improves selective extraction over discrete conditioning. The proposed method achieves 29.25~dB SI-SDR, 4.40~dB SI-SDRi, 62.01~dB attenuation, and 0.46\% RCE, outperforming fixed-threshold and local-range baselines. It also remains effective with more speakers and additive noise.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Biao Dong, Jie Chen, Jianwei Fang, Wei Xiao, Jiqing Han, Yongjun He. 2026-10-08. Monaural Continuous-Radius Regional Speech Extraction with Cross-Radius Consistency Learning. https://arxiv.org/abs/2610.11507
Cite the original work for its findings. Save a collection to share your selection of sources.