arXiv · 1906.10555
Naver at ActivityNet Challenge 2019 -- Task B Active Speaker Detection (AVA)
Abstract
This report describes our submission to the ActivityNet Challenge at CVPR 2019. We use a 3D convolutional neural network (CNN) based front-end and an ensemble of temporal convolution and LSTM classifiers to predict whether a visible person is speaking or not. Our results show significant improvements over the baseline on the AVA-ActiveSpeaker dataset.
Explore related subjects
Keep this discovery
Joon Son Chung. 2019-06-25. Naver at ActivityNet Challenge 2019 -- Task B Active Speaker Detection (AVA). https://arxiv.org/abs/1906.10555
Cite the original work for its findings. Save a collection to share your selection of sources.