arXiv · 1608.03995
An Analysis of Lemmatization on Topic Models of Morphologically Rich Language
Abstract
Topic models are typically represented by top-$m$ word lists for human interpretation. The corpus is often pre-processed with lemmatization (or stemming) so that those representations are not undermined by a proliferation of words with similar meanings, but there is little public work on the effects of that pre-processing. Recent work studied the effect of stemming on topic models of English texts and found no supporting evidence for the practice. We study the effect of lemmatization on topic models of Russian Wikipedia articles, finding in one configuration that it significantly improves interpretability according to a word intrusion metric. We conclude that lemmatization may benefit topic models on morphologically rich languages, but that further investigation is needed.
Explore related subjects
Keep this discovery
Chandler May, Ryan Cotterell, Benjamin Van Durme. 2016-08-13. An Analysis of Lemmatization on Topic Models of Morphologically Rich Language. https://arxiv.org/abs/1608.03995
Cite the original work for its findings. Save a collection to share your selection of sources.