arXiv · 1804.05388
Introducing two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness
Abstract
We present two novel datasets for the low-resource language Vietnamese to assess models of semantic similarity: ViCon comprises pairs of synonyms and antonyms across word classes, thus offering data to distinguish between similarity and dissimilarity. ViSim-400 provides degrees of similarity across five semantic relations, as rated by human judges. The two datasets are verified through standard co-occurrence and neural network models, showing results comparable to the respective English datasets.
Explore related subjects
Keep this discovery
Kim Anh Nguyen, Sabine Schulte im Walde, Ngoc Thang Vu. 2018-04-15. Introducing two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness. https://arxiv.org/abs/1804.05388
Cite the original work for its findings. Save a collection to share your selection of sources.