Search arXivSearch

arXiv · 2412.08475

Rethinking Mean Square Error: Information, Generalized Estimation, and the James-Stein Paradox

Abstract

The James-Stein estimator's dominance over maximum likelihood in mean square error has been called a paradox because maximum likelihood is known to be superior in many other respects. One response, due to Efron, is to question maximum likelihood. Another is to question MSE. We pursue the second and compare MSE with $Λ$-information (Vos and Wu, 2025) as criteria for assessing estimators. The comparison rests on two distinctions: between point estimators and generalized estimators -- functions of the sample and parameter jointly, with the score as archetype -- as inferential objects, and between pointwise and family-aware assessment criteria. An elementary lemma shows that no pointwise criterion, MSE or any other risk built from a loss function, admits a uniformly optimal estimator; $Λ$-information, which is family-aware and parameter-invariant, is uniformly maximized by the score. A point estimator is assessed through the generalized estimators it induces, and under the score map its $Λ$-efficiency is the fraction of Fisher information the statistic retains, placing the criterion in Fisher's information-loss tradition. On unbiased estimators, $Λ$-efficiency coincides with variance-based efficiency. Returning to James-Stein, the paradox dissolves: maximum likelihood is fully efficient because it is sufficient, while the James-Stein statistic is exactly two-to-one in the sample, and the information it destroys -- computed exactly -- is concentrated precisely where its MSE advantage is greatest. MSE retains its proper domain under genuine squared-error loss.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Paul W. Vos. 2026-07-15. Rethinking Mean Square Error: Information, Generalized Estimation, and the James-Stein Paradox. https://arxiv.org/abs/2412.08475

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Statistical models as natural transformations: meaningfulness, coherence and priors as states in Markov categories

We show that a statistical model in the sense of McCullagh, in the form given by Brøns, is a natural transformation between two functors from the category of designs to the Kleisli category Stoch of the Giry monad, provided that its components are measurable in the parameter. The condition is empty for finite models. A design-indexed quantity is a family of morphisms of Stoch defined on the parameter objects, called meaningful if it is natural. We prove that Tjur's criterion, imposed on parameter functions indexed by finite samples with multiplicities, forces the indexing by the support and then coincides with naturality over the insertions. For finite designs we show that a quantity can be corrected to a natural one within a given class of corrections if and only if a class vanishes in the first cohomology group of a Baues-Wirsching complex relative to that class, while its image in the absolute group is always zero. In the one-way layout, marginal dispersion is not meaningful, and within-group dispersion is the unique correction that leaves the merged design unchanged. A prior is a family of states on the parameter objects, called coherent over a class of design morphisms if it is natural over that class. We show that coherence at a merge confines the prior to the image of the corresponding parameter map, that coherence over the insertions is Kolmogorov consistency, and that coherence over the injections adds the exchangeability assumed by the categorical de Finetti theorem. In the finite one-way scheme, the coherent priors form polytopes of known dimension. The analogue of Jeffreys' general rule is not coherent, while the analogue for location-scale families is. Finally, we show that ridge regression is the Bayesian inversion of the Gaussian linear model with respect to a Gaussian prior, which is coherent over the insertions and never over the injections.

math.ST

Sample complexity and weak limits of nonsmooth multimarginal Schrödinger system with application to optimal transport barycenter

Multimarginal optimal transport (MOT) has emerged as a useful framework for many applied problems. However, compared to the well-studied classical two-marginal optimal transport theory, analysis of MOT is far more challenging and remains much less developed. In this paper, we study the statistical estimation and inference problems for the entropic MOT (EMOT), whose optimal solution is characterized by the multimarginal Schrödinger system. Assuming only boundedness of the cost function, we derive sharp sample complexity for estimating several key quantities pertaining to EMOT (cost functional and Schrödinger coupling) from point clouds that are randomly sampled from the input marginal distributions. Moreover, with substantially weaker smoothness assumption on the cost function than the existing literature, we derive distributional limits and bootstrap validity of various key EMOT objects. As an application, we propose the multimarginal Schrödinger barycenter as a new and natural way to regularize the exact Wasserstein barycenter and demonstrate its statistical optimality.

math.ST

Nonparametric spectral density estimation using interactive mechanisms under local differential privacy

We study the problem of estimating the spectral density of a centered stationary Gaussian time series under local differential privacy constraints. Specifically, we propose new interactive privacy mechanisms for three tasks: recovering a single covariance coefficient, recovering the spectral density at a fixed frequency, and global recovery. Our approach achieves faster rates through a two-stage process: we first apply the Laplace mechanism to the truncated value, and then use the resulting privatized sample to learn about the dependence mechanism in the time series. For spectral densities belonging to Hölder and Sobolev smoothness classes, we demonstrate that our algorithms improve upon the non-interactive mechanism of Kroll (2024) for small privacy parameter $α$, since the pointwise rates depend on $nα^2$ instead of $nα^4$. Moreover, we show that the rate $(nα^4)^{-1}$ is optimal for estimating a covariance coefficient with non-interactive mechanisms. However, the $L_2$ rate of our interactive estimator is slower than the pointwise rate. We show how to use these procedures to provide a bona fide locally differentially private estimator of the entire covariance matrix. A simulation study validates our findings.

math.ST