arXiv · 2009.00672
Document Similarity from Vector Space Densities
Abstract
We propose a computationally light method for estimating similarities between text documents, which we call the density similarity (DS) method. The method is based on a word embedding in a high-dimensional Euclidean space and on kernel regression, and takes into account semantic relations among words. We find that the accuracy of this method is virtually the same as that of a state-of-the-art method, while the gain in speed is very substantial. Additionally, we introduce generalized versions of the top-k accuracy metric and of the Jaccard metric of agreement between similarity models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ilia Rushkin. 2020-09-01. Document Similarity from Vector Space Densities. https://doi.org/10.1007/978-3-030-55187-2_14
Cite the original work for its findings. Save a collection to share your selection of sources.