arXiv · 2606.16115
Stabilizing Short Duration Speaker Verification through Neural Re-scoring with Hybrid Enrollment
Abstract
Short-duration speaker verification (SDSV) is crucial for personalized keyword spotting, where test utterances are typically shorter than three seconds. Limited speech duration results in unstable speaker representations and increased sensitivity to noise and phoneme variations, thereby degrading performance. To investigate this issue, we construct VoxPhrase, a large-scale SDSV corpus automatically segmented from the VoxCeleb dataset. Our analysis shows that text-dependent (TD) enrollment is constrained by duration and yields unstable speaker representations. In contrast, although text-independent (TI) enrollment introduces content mismatch, its representations become more stable as the enrollment duration increases. Accordingly, we propose a hybrid-enrollment neural re-scoring framework that combines TD and TI enrollment and performs frame-level comparison via parallel cross-attention. Experiments on VoxPhrase demonstrate consistent improvements across multiple speaker models.
Explore related subjects
Keep this discovery
Zhiqi Ai, Han Cheng, Shiyi Mu, Zhiyong Chen, Yongjin Zhou, Shugong Xu. 2026-06-15. Stabilizing Short Duration Speaker Verification through Neural Re-scoring with Hybrid Enrollment. https://arxiv.org/abs/2606.16115
Cite the original work for its findings. Save a collection to share your selection of sources.