arXiv · 2607.26202
Randomizing the Number of Centers in k-means++
Abstract
The $k$-means++ algorithm is a standard and widely used seeding method for $k$-means clustering, but for a fixed number $k$ of centers its worst-case expected approximation ratio is $\Theta(\log k)$. We consider the same algorithm when an adversary first fixes the dataset and some $K$; the number of centers $k$ is then chosen uniformly from $\{K,\ldots,2K-1\}$. We prove that $k$-means++ is an $O(1)$-approximation with constant probability in this budget-smoothed setup.
Explore related subjects
Keep this discovery
Vaclav Rozhon. 2026-07-28. Randomizing the Number of Centers in k-means++. https://arxiv.org/abs/2607.26202
Cite the original work for its findings. Save a collection to share your selection of sources.