arXiv · math/0701347
Efficient estimation of the cardinality of large data sets
Abstract
F.Giroire has recently proposed an algorithm which returns the approximate number of distincts elements in a large sequence of words, under strong constraints coming from the analysis of large data bases. His estimation is based on statistical properties of uniform random variables in $[0,1]$. In this note we propose an optimal estimation, using Kullback information and estimation theory.
Explore related subjects
Keep this discovery
Philippe Chassaing, Lucas Gerin. 2011-04-22. Efficient estimation of the cardinality of large data sets. https://arxiv.org/abs/math/0701347
Cite the original work for its findings. Save a collection to share your selection of sources.