arXiv · 2109.11969
Rethinking Crowd Sourcing for Semantic Similarity
Abstract
Estimation of semantic similarity is crucial for a variety of natural language processing (NLP) tasks. In the absence of a general theory of semantic information, many papers rely on human annotators as the source of ground truth for semantic similarity estimation. This paper investigates the ambiguities inherent in crowd-sourced semantic labeling. It shows that annotators that treat semantic similarity as a binary category (two sentences are either similar or not similar and there is no middle ground) play the most important role in the labeling. The paper offers heuristics to filter out unreliable annotators and stimulates further discussions on human perception of semantic similarity.
Explore related subjects
Keep this discovery
Shaul Solomon, Adam Cohn, Hernan Rosenblum, Chezi Hershkovitz, Ivan P. Yamshchikov. 2021-09-24. Rethinking Crowd Sourcing for Semantic Similarity. https://arxiv.org/abs/2109.11969
Cite the original work for its findings. Save a collection to share your selection of sources.