arXiv · 1810.12411
Counting in Language with RNNs
Abstract
In this paper we examine a possible reason for the LSTM outperforming the GRU on language modeling and more specifically machine translation. We hypothesize that this has to do with counting. This is a consistent theme across the literature of long term dependence, counting, and language modeling for RNNs. Using the simplified forms of language -- Context-Free and Context-Sensitive Languages -- we show how exactly the LSTM performs its counting based on their cell states during inference and why the GRU cannot perform as well.
Explore related subjects
Keep this discovery
Heng xin Fun, Sergiy V Bokhnyak, Francesco Saverio Zuppichini. 2018-10-29. Counting in Language with RNNs. https://arxiv.org/abs/1810.12411
Cite the original work for its findings. Save a collection to share your selection of sources.