arXiv · 1703.04105
Combining Residual Networks with LSTMs for Lipreading
Abstract
We propose an end-to-end deep learning architecture for word-level visual speech recognition. The system is a combination of spatiotemporal convolutional, residual and bidirectional Long Short-Term Memory networks. We train and evaluate it on the Lipreading In-The-Wild benchmark, a challenging database of 500-size target-words consisting of 1.28sec video excerpts from BBC TV broadcasts. The proposed network attains word accuracy equal to 83.0, yielding 6.8 absolute improvement over the current state-of-the-art, without using information about word boundaries during training or testing.
Explore related subjects
Keep this discovery
Themos Stafylakis, Georgios Tzimiropoulos. 2017-03-12. Combining Residual Networks with LSTMs for Lipreading. https://arxiv.org/abs/1703.04105
Cite the original work for its findings. Save a collection to share your selection of sources.