Search arXiv⌕ Search

arXiv · 2609.35453

Hermite Brings a Laptop: Analyzing Frieze-Jerrum Rounding Yields Improved Approximations for Clustering Problems

Abstract

The Frieze-Jerrum rounding is a standard tool for rounding SDP relaxations of graph partitioning and clustering problems, assigning nodes to at most $k$ clusters using $k$ independent Gaussian vectors. Its analysis hinges on the collision probability $P_k(ρ)$ that two nodes whose SDP vectors have inner product $ρ$ are assigned to the same cluster. No tractable closed form for $P_k$ is known for $k\geq 4$, making it difficult to certify approximation guarantees and hindering the systematic search for better algorithms. We develop a Hermite-coefficient certification framework to derive accurate and tractable bounds on $P_k$. Using the Hermite expansion of Gaussian noise stability, we express $P_k$ as a power series with nonnegative coefficients, reduce these coefficients to one-dimensional Gaussian integrals, and certify finitely many of them, yielding rigorous bounds on $P_k$ over the entire correlation range. Our framework yields strengthened polynomial-time approximations for several clustering problems. For MaxAgree Correlation Clustering, we derive a $0.7818$-approximation, the first improvement in two decades over the $0.7666$ ratio of Swamy (2004). On the hardness side, we show that the integrality ratio of the standard SDP relaxation is at most $0.802$, and that approximation beyond that is Unique Games-hard. We also improve the best known ratios for the variant with at most $K$ clusters, MaxAgree$[K]$ (e.g., from $0.77$ to $0.8151$ for $K=3$). For Max $K$-Cut we resolve, via a structural property of the Hermite expansion, a conjecture of de Klerk et al. (2004) characterizing the Frieze--Jerrum approximation ratio for every $K\ge3$; we show that this ratio is tight, and determine it to within $10^{-6}$ accuracy for $K\le16$. Finally, we reduce the additive approximation error for modularity maximization from $0.42084$ (Kawase et al., 2021) to $0.3790$.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

David García-Soriano, Atsushi Miyauchi. 2026-09-28. Hermite Brings a Laptop: Analyzing Frieze-Jerrum Rounding Yields Improved Approximations for Clustering Problems. https://arxiv.org/abs/2609.35453

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Min-Sum Set Cover on Parallel Machines

We consider a generalization of the Min-Sum Set Cover to the setup with $m$ set-sequences, or in scheduling terminology, $m$ parallel machines. We call this problem Parallel Min-Sum Set Cover. To obtain approximation algorithms for its numerous variants we use a crucial sub-problem called Parallel Densest Subfamily. We prove that an $α$-approximation algorithm for this task gives a $4\cdotα$-approximation for the Parallel Min-Sum Set Cover, which yields $\frac{4\cdot e}{e-1}+ε$ and $4\cdot \frac{e}{e-1}^2+ε$-approximation ratios for identical and unrelated machines, respectively. To obtain the latter result we give a new $\frac{e}{e-1}^2+ε$-approximation algorithm for the Maximum Coverage Multiple Knapsacks problem which is of independent interest. If the sets are precedence-constrained, for unit cost sets we give an $\mathcal{O}(k^{2/3})$ approximation ($k$ is the number of sets). For the case of out-forest precedence constraints we improve this bound to $\mathcal{O}(\log k)$ via a reduction to the Group Steiner Orienteering problem, and show this is tight, unless $NP\subseteq ZTIME(n^{\mathcal{O}(\text{poly}(\log n))})$.

cs.DS↗

Learning Latent Algebraic Structure from Ambiguous Set Observations

We study when statistically learnable latent structure can also be recovered efficiently, and how membership queries change the answer. An unknown support $A\subseteq\mathbb F_2^n$ has small additive doubling and is observed through a fixed set $B$ satisfying $|A\triangle B|\leη|A|$. We seek one linear subspace $V$ such that every compatible support $A$ is covered by few $V$-cosets and satisfies $|V|\le|A|$. For every $η<1$, polynomially many uniform samples suffice statistically, with cost polynomial in the doubling constant and proportional to $(1-η)^{-1}$; this radius dependence is sharp. Under a specified hardness assumption for learning parities with noise (search-LPN), however, no polynomial-time sample-only learner achieves even constant covering cost, including when the latent support is unique. At fixed structural parameters and the same constant covering budget, adding exact membership queries to $B$ permits polynomial-time recovery. The general query learner constructs a short structural list and uses fresh samples to select one common output through a majority-coverage rule. Persistent structured cores make this candidate construction possible. At doubling one, a complementary distinction appears at $η=1/3$: coarse recovery remains polynomial time, while exact recovery requires exponentially many accesses in the worst case when latent cardinality is unknown.

cs.DS↗

Testing the Binary Rank with Polynomial Query Complexity

We design an adaptive two-sided error testing algorithm for the binary rank of a $0,1$ matrix $M$ with query complexity $O(d^3\log(d+1)/ε^2)$, where $d$ is the tested binary rank bound and $ε$ is the distance parameter. This answers an open question posed by Parnas, Ron and Shraibman~\cite{parnas2021property}, who asked if the binary rank can be tested with query complexity polynomial in $d$ and $1/ε$. Furthermore, our testing algorithm can be used to find an approximate binary decomposition of $M$ with an additional $d(n+m)$ queries. That is, under the promise that the binary rank of $M$ is at most $d$, we show how to find, with probability at least $5/6$, two $0,1$ matrices $A',B'$ such that $M' = A' \cdot B'$ is a $0,1$ matrix which differs from $M$ on at most an $O(ε)$ fraction of its entries. Our results also imply a testing algorithm with polynomial query complexity for the equivalent problem of testing if the edges of a bipartite graph can be partitioned into at most $d$ bicliques.

cs.DS↗