Search arXiv⌕ Search

arXiv · 2610.11741

Efficient Recovery of Latent Coordinate Structure from Sparse Observations of the Hypercube

Abstract

Recovering latent geometric structure from graph observations is a well-studied problem in statistical inference. Kapralov, Trevisan, and Wrzos-Kaminska (2026) introduced the problem of recovering the coordinate structure of the Boolean hypercube from a small random sample of its edges. More specifically, there are $n=2^d$ vertices, each corresponding to a distinct "feature" vector in $\{\pm 1\}^d$. Between each pair whose feature vectors are at Hamming distance one, an edge is observed independently with probability $p$. As long as the expected degree $pd$ is $\gtrsim \log d = \log \log n$, we give a polynomial-time algorithm that, given only the graph of observed edges, correctly recovers the entire feature vector of all but a vanishing fraction of vertices. This matches the information-theoretic guarantee of Kapralov, Trevisan, and Wrzos-Kaminska in polynomial rather than exponential time, resolving the algorithmic question left open by their work. Our algorithm combines a degree-$4$ sum-of-squares certificate for the structure of balanced near-minimum cuts with a rounding scheme originally developed for tensor decomposition by Ma, Shi, and Steurer (2016). The analysis relies on two novel ingredients: a sum-of-squares version of the Friedgut-Kalai-Naor theorem in Boolean Fourier analysis and a spectral concentration result for the observed subgraph of the hypercube. Finally, we provide a justification for why higher-degree sum-of-squares might be needed by showing a limitation of the basic SDP relaxation of this problem.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rares-Darius Buhai, Davide Mazzali, Weronika Wrzos-Kaminska. 2026-10-08. Efficient Recovery of Latent Coordinate Structure from Sparse Observations of the Hypercube. https://arxiv.org/abs/2610.11741

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Compression with wildcards: All, or all maximum, anticlques of a graph

By definition an anticlique is an independent set of vertices of a graph $G$. By duality all results obtained for anticliques carry over to cliques. (It is for technical reasons that we stick with anticliques throughout.) We display the set $Acl(G)$ of all anticliques of $G$ in a compressed format that uses wildcards. Likewise (albeit less compressed) for the subfamily $MACL(G)\s Acl(G)$ of all maximum-cardinality members. The second task works particularly well for bipartite graphs (in fact for the broader class of König-Egarváry graphs). In this scenario Boolean functions (of type 2-CNF) will be important. Dilworth's lattice of all maximum antichains of a poset also features prominently.

cs.DS↗

Matroid Base Packings: Improved Dynamic Matroid Density and Combinatorics of Tree Packings

Greedy minimum-weight spanning tree packings are an important tool in graph connectivity algorithms. We study the corresponding process of greedy base packing in matroids, following the work of de Vos and Grilnberger. Using a modified version of matroid base packings, we give a fully dynamic $(1 \pm \varepsilon)$-approximation to the matroid density using $O((ρ_{\max}^2\varepsilon^{-2}+ρ_{\max}\varepsilon^{-4})\log^3m_{\max})$ worst-case rank queries per update, where $ρ_{\max}$ upper-bounds the density and $m_{\max}$ upper-bounds the ground set size. Sampling yields a $(1 \pm \varepsilon)$-approximation with high probability against an oblivious adversary using $O(\varepsilon^{-6}\log^6m_{\max})$ worst-case rank queries per update. For graphic matroids, we strengthen the lower bound on the convergence rate of relative edge loads to ideal loads, closing the gap between the lower and upper bounds up to a logarithmic factor. We also show that a packing of $O(λ^5\log m)$ trees contains a tree crossing some minimum cut once, improving the bound $O(λ^7\log^3m)$ of Thorup. In the appendix, we consider a specialization of the greedy base packings to bicircular matroids, which yields a dynamic approximation of the graph density. For this, we develop a dynamic data structure that maintains a minimum-weight maximal pseudoforest.

cs.DS↗

Improved Online Hitting Set Algorithms for Structured and Geometric Set Systems

In the online hitting set problem, sets arrive over time, and the algorithm has to maintain a subset of elements that hit all the sets seen so far. Alon, Awerbuch, Azar, Buchbinder, and Naor (SICOMP 2009) gave an algorithm with competitive ratio $O(\log n \log m)$ for the (general) online hitting set and set cover problems for $m$ sets and $n$ elements; this is known to be tight for efficient online algorithms. Given this barrier for general set systems, we ask: can we break this double-logarithmic phenomenon for online hitting set/set cover on structured and geometric set systems? We provide an $O(\log n \log\log n)$-competitive algorithm for the weighted online hitting set problem on set systems with linear shallow-cell complexity, replacing the double-logarithmic factor in the general result by effectively a single logarithmic term. As a consequence of our results we obtain the first bounds for weighted online hitting set for natural geometric set families, thereby answering open questions regarding the gap between general and geometric weighted online hitting set problems.

cs.DS↗