Search arXiv⌕ Search

arXiv · 2610.08634

Spectral Recovery of Point Clouds from Noisy Geometric Graphs

Abstract

We study the problem of recovering low-dimensional latent geometry from a random geometric graph generated by noisy, high-dimensional data. Specifically, we analyze the performance of a spectral embedding algorithm on the Signal+Noise Graph Model, in which vertices are associated to points perturbed by Gaussian noise, and edges are included for pairs whose inner product exceeds a specified alignment threshold. In the high-dimensional regime where the number $n$ of points and the ambient dimension $d$ both tend to infinity, we show that under a spectral gap condition, the top eigenvectors and eigenvalues of the graph's adjacency matrix can be used to approximately recover the point cloud up to an orthogonal transformation. We illustrate our results on point clouds sampled from nested spheres and high-dimensional sinusoid curves.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tatiana Brailovskaya, Nicholas A. Cook, Sofia Poinelli. 2026-10-06. Spectral Recovery of Point Clouds from Noisy Geometric Graphs. https://arxiv.org/abs/2610.08634

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A resolution of the Borel-Kolmogorov Paradox via the Maximum Entropy Principle

Bayesian updating routinely conditions on events of prior probability zero, such as an exact observation of a continuous parameter or a parameter confined to a submanifold. Two equally natural parametrizations of the same event can then return different posteriors, which is the Borel--Kolmogorov paradox. We add a metric to the probabilistic model and define the posterior given a closed null set as the limit of Maximum Entropy solutions under a vanishing constraint on the distance to that set. We give a sufficient condition for existence and prove that, given this metric, the posterior is unique and invariant under measure-preserving isometries. It recovers the textbook Bayes formulas where these apply unambiguously, on positive-measure sets and in Euclidean models. Each stage of the limit has an explicit density that depends only on the distance to the conditioning set, so the posterior can be computed with standard tools for Kullback--Leibler minimization. The construction does not remove the choice of a metric but makes it an explicit modelling input. In a controlled experiment where the geometry of the measurement is known, we show that a Bayesian analysis ignoring this geometry can yield uncalibrated posteriors. On real palaeomagnetic data, where the geometry is disputed, we show how candidate metrics can be compared within the same Bayesian framework.

math.ST↗

Kernel Estimation Of Chatterjee's Dependence Coefficient

A dependence coefficient suggested by Chatterjee (2021) has attracted considerable attention in recent years. It takes values between 0 and 1, vanishes if and only if the corresponding random variables $X$ and $Y$ are independent, and equals 1 if and only if $Y$ is almost surely a measurable function of $X$. The coefficient coincides, for distributions with continuous marginals, with a dependence coefficient introduced in Dette, Siburg, and Stoimenov (2013). Several estimators of this measure have been proposed, including the kernel-type estimator of Dette, Siburg, and Stoimenov (2013), Chatterjee's rank correlation, and its nearest-neighbour modification by Lin and Han (2023). In this paper, we investigate the asymptotic properties and local power of the kernel-type estimator. We first establish a non-degenerate central limit theorem under independence: after suitable centring, the estimator is asymptotically normal at the rate $\sqrt{n/h_1}$, where $h_1\to0$ is a bandwidth. We then consider Gaussian-copula alternatives with correlation parameter $ρ_n\to0$. By deriving a second-order expansion of the mean, establishing variance stability and proving a central limit theorem under local alternatives, we obtain an explicit limiting power function and identify $(h_1/n)^{1/4}$ as the detection boundary. Under suitable bandwidth conditions, the detection boundary can attain rates arbitrarily close to $n^{-3/8}$. The exponent $3/8$ exceeds the exponent $1/4$ for the test based on Chatterjee's original rank correlation and the exponent approaching $5/16$ for the modified rank correlation of Lin and Han (2023) within the range of their central limit theorem. The proofs combine refined results for double-indexed permutation statistics with Gaussian rank-score expansions and Hermite-chaos projections. A simulation study illustrates the normal approximation and the resulting power properties.

math.ST↗

A Two-Sample Test on Weighted Persistence Intensity Functions in Topological Data Analysis

Persistence intensity functions provide interpretable and informative first-order summaries of random persistence diagram distributions. We study two-sample testing for equality of persistence intensity functions, allowing the underlying diagram distributions to differ under the null. We construct a weighted-kernel statistic as an unbiased estimator of the squared reproducing-kernel Hilbert space distance between the corresponding weighted intensity embeddings, and calibrate it by studentization. For shrinking bandwidths, we establish uniform asymptotic normality under the null and thereby obtain asymptotic Type I error control. Its power is characterized in terms of the $L^2$ discrepancy between the weighted intensity functions. To accommodate persistence diagrams with possibly unbounded cardinality, we introduce regularity conditions that control the effect of cardinality variation and yield the desired moment bounds for the test statistic. We further show that every probability density on $\{(x,y)\in\mathbb{R}^2:y>x > 0\}$ can be realized as the persistence intensity function of a random diagram. Using these results, we establish that the proposed test attains minimax-optimal separation rates over anisotropic Sobolev balls. Lastly, since the optimal bandwidth is not directly accessible in practice, we adapt a bandwidth aggregation framework.

math.ST↗