arXiv · 1803.06390
Corpus Statistics in Text Classification of Online Data
Abstract
Transformation of Machine Learning (ML) from a boutique science to a generally accepted technology has increased importance of reproduction and transportability of ML studies. In the current work, we investigate how corpus characteristics of textual data sets correspond to text classification results. We work with two data sets gathered from sub-forums of an online health-related forum. Our empirical results are obtained for a multi-class sentiment analysis application.
Explore related subjects
Keep this discovery
Marina Sokolova, Victoria Bobicev. 2018-03-16. Corpus Statistics in Text Classification of Online Data. https://arxiv.org/abs/1803.06390
Cite the original work for its findings. Save a collection to share your selection of sources.