arXiv · 2606.00342
PE-means: Improved Differentially Private $k$-means Clustering through Private Evolution
Abstract
We study the problem of differentially private (DP) $k$-means clustering in Euclidean space. Previous solutions rely on summing the private data directly, which induces a sensitivity proportional to the domain. We introduce PE-means, an extension of the private evolution (PE) algorithm (an increasingly popular method for synthetic data generation), to the problem of $k$-means clustering. The key advantage of PE is that it only computes a private histogram with constant sensitivity to guide the evolution. Our adaptation of PE includes new evolutionary operators for clustering, as well as other algorithmic improvements of independent interest. Overall, PE-means achieves an average improvement of 26% in clustering loss over state-of-the-art baselines such as Google's LSH-based algorithm and DP-Lloyd variants.
Explore related subjects
Keep this discovery
Thomas Humphries, Zinan Lin, Sergey Yekhanin. 2026-05-29. PE-means: Improved Differentially Private $k$-means Clustering through Private Evolution. https://arxiv.org/abs/2606.00342
Cite the original work for its findings. Save a collection to share your selection of sources.