Search arXiv⌕ Search

arXiv · 2511.21074

Inference for Similarity and Alignability between Noisy High-Dimensional Datasets

Abstract

The rapid growth of high-dimensional datasets across a wide range of scientific domains has created an urgent need for new statistical methods to compare distributions with underlying low-dimensional structure. Assessing similarity between high-dimensional datasets whose observations concentrate near low-dimensional manifolds is particularly challenging due to the nontrivial effects of noise in high dimensions. We propose a principled framework for statistical inference on the similarity and alignability of high-dimensional datasets with low-dimensional smooth signals under heterogeneous noise. The key idea is to link the spectral properties of the observed data matrices to the geometry of their underlying signal distributions. Under a manifold signal-plus-noise model, we build on the principal variances associated to the underlying signals and develop a scale- and rotation-invariant dissimilarity measure between two datasets that may differ in sample size and noise structure. We further construct an estimator of the dissimilarity and a statistical test of dataset alignability, namely, whether the dissimilarity vanishes. The proposed methodology and its theoretical guarantees under high-dimensional asymptotic settings draw on recent advances in random matrix theory (RMT). The proposed framework accommodates heterogeneous noise across datasets and provides a fast, theoretically grounded approach to comparing high-dimensional datasets with low-dimensional structures. Through extensive simulations and analyses of multiple single-cell datasets, we demonstrate that the proposed method substantially outperforms existing approaches.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hongrui Chen, Rong Ma. 2026-08-15. Inference for Similarity and Alignability between Noisy High-Dimensional Datasets. https://arxiv.org/abs/2511.21074

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Transitional Conditional Independence

Statistical models contain variables that are not random: parameters, treatments, environments, design points. Ordinary conditional independence cannot express relations involving such variables. To apply it one must first put a distribution on them, and that changes the meaning of the statement. This paper introduces transitional conditional independence. It relates three variables on a Markov kernel $K(W|T)$ with non-stochastic input $T$, and is defined by a single factorization: \[ X\perp\!\!\perp_{K(W|T)} Y |Z \quad :\iff \quad \exists\, Q(X|Z):\; K(X,Y,Z|T) = Q(X|Z)\otimes K(Y,Z|T).\] The relation asserts a Markov kernel $Q(X|Z)$ that is the same for every input $t$. It therefore yields a factorization rather than an almost-sure identity between conditional expectations, and it needs no distribution on the input space. The relation is asymmetric. We show that the asymmetry is essential: symmetrizing it destroys the statements it was built to make. We prove left and right versions of all separoid rules except Symmetry. Ten of them hold on arbitrary measurable spaces, the remaining ones under one condition on the spaces involved, and we give criteria for when Symmetry itself holds. We axiomatize the resulting structure and show that it arises from any symmetric separoid by a shift. We give several applications. Ancillarity, sufficiency and adequacy become factorizations that hold pointwise in the parameter, without a prior and without null sets; the theorems of Fisher--Neyman and of Basu take this form. The invariance hypothesis of invariant prediction, $Y \perp\!\!\perp E | X_S$, receives its intended meaning: one kernel predicts $Y$ from $X_S$ in every environment $E$. And Bayesian networks with non-stochastic input nodes satisfy a directed global Markov property whose graphical id-separation criterion returns a factorization of Markov kernels, on arbitrary input spaces.

math.ST↗

Choosing the penalty in nonparametric regression: short and long-range dependence

In this work, we study the one-dimensional regression problem under random design and Gaussian errors. Our framework is very general: we make no prior assumptions about the design (which may be nonstationary and exhibit short or long-range dependence), nor do we assume that the errors are homoscedastic. We examine in detail the cases where the error process exhibits short or long-range dependence. We adopt a least-squares penalized strategy using piecewise polynomials to estimate the regression function, following the framework of Baron, Birg{é} and Massart [1999]. We derive explicit penalties, up to calibration constants, to obtain adaptive estimators for which we establish risk bounds. Since these penalties depend on the dependence properties of the error process, which are unknown in practice, we propose several adaptations of the dimension jump calibration algorithm to make our procedures fully data-driven.

math.ST↗

Information geometry of emergent symmetry quotients and their weak unfoldings

A regular observed statistical model may converge to a limit in which a previously identifiable signed parameter becomes identifiable only modulo a reflection. We study the local information geometry of this transition. For a twice differentiable Hellinger embedding with an exact limiting reflection, the observed displacement is forced into the two-jet form \(\varepsilonλJ_-+λ^2J_+/2\), up to higher-order terms. The mixed jet restores the sign away from the symmetric face, whereas the even jet is the first tangent inherited by the quotient. After nuisance elimination, a positive Gram determinant yields a nondegenerate cross-cap two-jet. The associated local asymptotic theory has three regimes governed by \(τ_n=\sqrt n\,\varepsilon_n^2\): regular signed LAN, a critical curved Gaussian subexperiment, and a quotient regime with the \(n^{-1/4}\) signed scale. We prove that the same parabolic critical experiment persists for predictive likelihoods along a single stationary dependent trajectory. The limiting quotient has a regular Fisher metric in the invariant coordinate, while its pullback degenerates in the signed coordinate. For a solvable CIR--OU benchmark motivated by coherent sea-clutter observations, we derive the quotient Fisher metric and curvature explicitly and show that the curvature is strictly negative. We also determine the restricted holonomy of the full Amari family: \(\operatorname{Hol}_0(\nabla^{(a)})=SO(2)\) for \(a=0\), whereas \(\operatorname{Hol}_0(\nabla^{(a)})=GL^+(2,\mathbb R)\) for \(a\neq0\). The results separate the intrinsic geometry of the limiting quotient from the transverse geometry of its weak unfolding.

math.ST↗