arXiv · 2203.15267
Selective inference for k-means clustering
Abstract
We consider the problem of testing for a difference in means between clusters of observations identified via k-means clustering. In this setting, classical hypothesis tests lead to an inflated Type I error rate. To overcome this problem, we take a selective inference approach. We propose a finite-sample p-value that controls the selective Type I error for a test of the difference in means between a pair of clusters obtained using k-means clustering, and show that it can be efficiently computed. We apply our proposal in simulation, and on hand-written digits data and single-cell RNA-sequencing data.
Explore related subjects
Keep this discovery
Yiqun T. Chen, Daniela M. Witten. 2022-03-29. Selective inference for k-means clustering. https://arxiv.org/abs/2203.15267
Cite the original work for its findings. Save a collection to share your selection of sources.