arXiv · 2308.07595
The DKU-MSXF Diarization System for the VoxCeleb Speaker Recognition Challenge 2023
Abstract
This paper describes the DKU-MSXF submission to track 4 of the VoxCeleb Speaker Recognition Challenge 2023 (VoxSRC-23). Our system pipeline contains voice activity detection, clustering-based diarization, overlapped speech detection, and target-speaker voice activity detection, where each procedure has a fused output from 3 sub-models. Finally, we fuse different clustering-based and TSVAD-based diarization systems using DOVER-Lap and achieve the 4.30% diarization error rate (DER), which ranks first place on track 4 of the challenge leaderboard.
Explore related subjects
Keep this discovery
Ming Cheng, Weiqing Wang, Xiaoyi Qin, Yuke Lin, Ning Jiang, Guoqing Zhao, Ming Li. 2023-08-15. The DKU-MSXF Diarization System for the VoxCeleb Speaker Recognition Challenge 2023. https://arxiv.org/abs/2308.07595
Cite the original work for its findings. Save a collection to share your selection of sources.