Search arXiv⌕ Search

arXiv · 2609.34315

Augmented James--Stein estimation for leading eigenvectors and eigenspaces in high dimensions

Abstract

Building on the James--Stein approach to leading eigenvector estimation (Goldberg and Kercheval, Proc. Natl. Acad. Sci. USA 120, e2207046120, 2023), we develop a data-adaptive augmented James--Stein shrinkage framework for estimating leading eigenvectors and eigenspaces under a generalized spiked population model, in the high-dimensional regime where the dimension $p$ and sample size $n$ grow proportionally. For each spiked eigenvector, we construct an augmented target subspace that combines auxiliary information, either from domain knowledge or prior information, with the remaining sample spiked eigenvectors. This augmentation allows information shared across the sample spiked components to be exploited while retaining a fully data-driven shrinkage rule. We show that the resulting eigenvector estimator strictly improves upon standard PCA whenever the target subspace contains nonvanishing information about the population eigenvector, while asymptotically reverting to PCA when the target is uninformative. The individual estimators further yield a nested sequence of estimators for all leading spiked eigenspaces, with analogous dominance properties. The proposed estimator also strictly improves upon the existing HDLSS-motivated shrinkage estimator in the proportional high-dimensional regime. Simulation studies demonstrate substantial finite-sample gains and robustness to misspecification of the number of spikes.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Giheon Seong, Seungki Hong, Sungkyu Jung. 2026-09-28. Augmented James--Stein estimation for leading eigenvectors and eigenspaces in high dimensions. https://arxiv.org/abs/2609.34315

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Intrinsic-dimension empirical Bernstein inequalities for bounded self-adjoint operators

Operator-valued concentration inequalities are foundational to the analysis of modern high-dimensional statistics and randomized algorithms. However, standard oracle bounds are frequently limited in practice: they require explicit a priori knowledge of the true variance, and often explicitly scale with the ambient dimension, rendering them vacuous for infinite-dimensional or heavily structured operators. Motivated by these challenges, we establish the first empirical Bennett and Bernstein inequalities for sums of independent, bounded, self-adjoint Hilbert-Schmidt operators. Our fully data-driven bounds replace the unknown variance with an empirical estimate and rely strictly on the intrinsic dimension rather than the ambient dimension. This structural shift yields computable, dimension-free guarantees with a sharper first-order asymptotic radius for non-isotropic random matrices and seamlessly extends to infinite-dimensional Hilbert spaces. We demonstrate that our empirical bounds achieve asymptotic sharpness with the best known oracle rates. Finally, as an independent byproduct, we derive novel empirical concentration guarantees for the intrinsic dimension itself.

math.ST↗

Bentkus-type asymptotic e-values

Asymptotic e-values are emerging as a powerful alternative to asymptotic p-values, particularly in post-hoc inference and multiple testing, where significance levels may be data-dependent. Existing asymptotic e-values, however, suffer from the ``missing factor,'' a scaling inefficiency resulting in overly conservative inference. Drawing on the framework of near-optimal concentration inequalities developed by Bentkus in the 2000s, we introduce Bentkus-type asymptotic e-values and prove that they successfully eliminate the missing factor. We also demonstrate both theoretically and empirically that Bentkus-type e-values consistently deliver sharper inference than existing alternatives, leading to tighter post-hoc confidence intervals and higher rejection rates in multiple testing procedures.

math.ST↗

The Spectra of the Henze-Zirkler and Henze-Wagner Operators for BHEP Tests

The Baringhaus-Henze-Epps-Pulley (BHEP) tests for multivariate normality are affine-invariant goodness-of-fit tests based on a Gaussian-weighted $L^2$ distance between empirical and Gaussian characteristic functions. In 1990, Henze and Zirkler expressed the limiting null distribution through the eigenvalues of an integral operator on the standard Gaussian space. In 1997, Henze and Wagner obtained a simpler covariance kernel and raised the problem of calculating the eigenvalues of the resulting operator on a Gaussian-weighted space. Although subsequent work treated the univariate case and numerical approximations in a few low dimensions, the complete all-dimensional spectral problem remained open. This paper determines both complete spectra for every dimension $d \in \mathbb{N}$ and every smoothing parameter $β> 0$. The two operators are shown to have the forms $\mathcal{X}_{β,d}^*\mathcal{X}_{β,d}$ and $\mathcal{X}_{β,d}\mathcal{X}_{β,d}^*$ for the same Hilbert-Schmidt operator $\mathcal{X}_{β,d}$. Consequently, their nonzero eigenvalues agree, including multiplicities, while the null space of the Henze-Zirkler operator is identified exactly. The Gaussian integral operator in the Henze-Wagner decomposition is diagonalized by Mehler's formula, and rotational symmetry confines the finite-rank correction to the sectors associated with spherical harmonics of degrees $0$, $1$, and $2$. The degree-$1$ and degree-$2$ eigenvalues are characterized by scalar transcendental equations, and the radial eigenvalues by an explicit pole-safe Fredholm determinant. The paper establishes nonnegativity, multiplicities, eigenfunction reconstruction, completeness, the trace identity, and a complete characterization of all exceptional pole cases.

math.ST↗