arXiv · 1908.10833
Data ultrametricity and clusterability
Abstract
The increasing needs of clustering massive datasets and the high cost of running clustering algorithms poses difficult problems for users. In this context it is important to determine if a data set is clusterable, that is, it may be partitioned efficiently into well-differentiated groups containing similar objects. We approach data clusterability from an ultrametric-based perspective. A novel approach to determine the ultrametricity of a dataset is proposed via a special type of matrix product, which allows us to evaluate the clusterability of the dataset. Furthermore, we show that by applying our technique to a dissimilarity space will generate the sub-dominant ultrametric of the dissimilarity.
Explore related subjects
Keep this discovery
Dan Simovici, Kaixun Hua. 2019-08-28. Data ultrametricity and clusterability. https://doi.org/10.1088/1742-6596/1334/1/012002
Cite the original work for its findings. Save a collection to share your selection of sources.