arXiv · 1707.00621
Including Dialects and Language Varieties in Author Profiling
Abstract
This paper presents a computational approach to author profiling taking gender and language variety into account. We apply an ensemble system with the output of multiple linear SVM classifiers trained on character and word $n$-grams. We evaluate the system using the dataset provided by the organizers of the 2017 PAN lab on author profiling. Our approach achieved 75% average accuracy on gender identification on tweets written in four languages and 97% accuracy on language variety identification for Portuguese.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alina Maria Ciobanu, Marcos Zampieri, Shervin Malmasi, Liviu P. Dinu. 2017-07-03. Including Dialects and Language Varieties in Author Profiling. https://arxiv.org/abs/1707.00621
Cite the original work for its findings. Save a collection to share your selection of sources.