Searcharxiv⌕ Search

arXiv · 2609.35453

Hermite Brings a Laptop: Analyzing Frieze-Jerrum Rounding Yields Improved Approximations for Clustering Problems

Abstract

The Frieze-Jerrum rounding is a standard tool for rounding SDP relaxations of graph partitioning and clustering problems, assigning nodes to at most $k$ clusters using $k$ independent Gaussian vectors. Its analysis hinges on the collision probability $P_k(ρ)$ that two nodes whose SDP vectors have inner product $ρ$ are assigned to the same cluster. No tractable closed form for $P_k$ is known for $k\geq 4$, making it difficult to certify approximation guarantees and hindering the systematic search for better algorithms. We develop a Hermite-coefficient certification framework to derive accurate and tractable bounds on $P_k$. Using the Hermite expansion of Gaussian noise stability, we express $P_k$ as a power series with nonnegative coefficients, reduce these coefficients to one-dimensional Gaussian integrals, and certify finitely many of them, yielding rigorous bounds on $P_k$ over the entire correlation range. Our framework yields strengthened polynomial-time approximations for several clustering problems. For MaxAgree Correlation Clustering, we derive a $0.7818$-approximation, the first improvement in two decades over the $0.7666$ ratio of Swamy (2004). On the hardness side, we show that the integrality ratio of the standard SDP relaxation is at most $0.802$, and that approximation beyond that is Unique Games-hard. We also improve the best known ratios for the variant with at most $K$ clusters, MaxAgree$[K]$ (e.g., from $0.77$ to $0.8151$ for $K=3$). For Max $K$-Cut we resolve, via a structural property of the Hermite expansion, a conjecture of de Klerk et al. (2004) characterizing the Frieze--Jerrum approximation ratio for every $K\ge3$; we show that this ratio is tight, and determine it to within $10^{-6}$ accuracy for $K\le16$. Finally, we reduce the additive approximation error for modularity maximization from $0.42084$ (Kawase et al., 2021) to $0.3790$.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

David García-Soriano, Atsushi Miyauchi. 2026-09-28. Hermite Brings a Laptop: Analyzing Frieze-Jerrum Rounding Yields Improved Approximations for Clustering Problems. https://arxiv.org/abs/2609.35453

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Capacitated Partition Vertex Cover and Partition Edge Cover

Our first focus is the Capacitated Partition Vertex Cover (C-PVC) problem in hypergraphs. In C-PVC, we are given a hypergraph with capacities on its vertices and a partition of the hyperedge set into $ω$ distinct groups. The objective is to select a minimum size subset of vertices that satisfies two main conditions: (1) in each group, the total number of covered hyperedges meets a specified threshold, and (2) the number of hyperedges assigned to any vertex respects its capacity constraint. A covered hyperedge is required to be assigned to a selected vertex that belongs to the hyperedge. This formulation generalizes classical Vertex Cover, Partial Vertex Cover, and Partition Vertex Cover. We investigate two variants: soft capacitated (multiple copies of a vertex are allowed) and hard capacitated (each vertex can be chosen at most once). Let $f$ denote the rank of the hypergraph (i.e., the maximum number of vertices contained in any single hyperedge). Our main contributions are: $(i)$ an $(f+1)$-approximation algorithm for the weighted soft-capacitated C-PVC problem, which runs in polynomial time for constant \(ω\), and $(ii)$ an $(f+ε)$-approximation algorithm for the unweighted hard-capacitated C-PVC problem, which runs in $n^{O(ω/ε)}$ time. We also study a natural generalization of the edge cover problem, the \emph{Weighted Partition Edge Cover} (W-PEC) problem, where each edge has an associated weight, and the vertex set is partitioned into groups. For each group, the goal is to cover at least a specified number of vertices using incident edges, while minimizing the total weight of the selected edges. We present the first exact polynomial-time algorithm for the weighted case, improving runtime from $O(ωn^3)$ to $O(mn+n^2 \log n)$ and simplifying the algorithmic structure over prior unweighted approaches.

cs.DS↗

Randomization for Faster Exact Optimization of Discounted Markov Decision Processes

We provide faster deterministic and randomized algorithms for exactly solving discounted Markov Decision Processes (DMDPs). We obtain our results by efficiently reducing computing optimal values and policies in DMDPs to the easier tasks of policy evaluation and computing approximately optimal values in DMDPs. We provide both a straightforward deterministic reduction and a more efficient randomized variant that, together with advances in approximately solving DMDPs, yield our results.

cs.DS↗

An ETH-Tight, Constructive FPT Algorithm for the Cone and Polytope Intersection Problem

In a landmark paper, Goemans and Rothvoss (2020) established an XP algorithm running in time $\text{enc}(P)^{2^{O(d)}} \cdot \text{enc}(Q)^{O(1)}$ for the Cone and Polytope Intersection problem: finding a vector $y \in \text{int.cone}(P \cap \mathbb{Z}^d) \cap Q$ together with a sparse certificate $λ\in \mathbb{Z}_{\ge 0}^{P \cap \mathbb{Z}^d}$ supported on at most $2^{2d+1}$ generators, where $P \subseteq \mathbb{R}^d$ is a bounded rational polyhedron and $Q \subseteq \mathbb{R}^d$ is an arbitrary rational polyhedron. For high-multiplicity bin packing, this gives a running time of ${|I|}^{2^{O(d)}}$, where $|I|$ denotes the encoding length of the input. Recently, Koana and Kumabe (2026) proved that the decision variant of this problem is fixed-parameter tractable (FPT) parameterized by the number of item types $d$ with running time $2^{d^{O(d)}} \cdot {|I|}^{O(1)} = 2^{2^{O(d \log d)}} \cdot {|I|}^{O(1)}$. In this work, we generalize the framework of Koana and Kumabe from standard bin packing to the full Cone and Polytope Intersection Problem of Goemans and Rothvoss, directly encompassing high-multiplicity bin packing, point-in-cone, and scheduling. Secondly, by combining Carathéodory-type integer cone bounds (Eisenbrand and Shmonin, 2006) with active support enumeration, we reduce the running time to: $$2^{2^{O(d)}} \cdot (\text{enc}(P) + \text{enc}(Q))^{O(1)}.$$ Under the Exponential Time Hypothesis (ETH), the double-exponential lower bound of Kowalik, Lassota, Majewski, Pilipczuk, and Sokołowski (2024) for point-in-cone and Jansen, Ohnesorge, and Pirotton (2026) for high-multiplicity bin packing implies that this parameter dependence is asymptotically optimal. Finally, we provide an explicit decompression algorithm that extracts a solution with sparse support $|\text{supp}(λ)| \le 2^{2d+1}$ in single-exponential FPT time.

cs.DS↗