Search arXivSearch

arXiv · 2008.11840

Out-of-sample error estimate for robust M-estimators with convex penalty

Abstract

A generic out-of-sample error estimate is proposed for robust $M$-estimators regularized with a convex penalty in high-dimensional linear regression where $(X,y)$ is observed and $p,n$ are of the same order. If $ψ$ is the derivative of the robust data-fitting loss $ρ$, the estimate depends on the observed data only through the quantities $\hatψ= ψ(y-X\hatβ)$, $X^\top \hatψ$ and the derivatives $(\partial/\partial y) \hatψ$ and $(\partial/\partial y) X\hatβ$ for fixed $X$. The out-of-sample error estimate enjoys a relative error of order $n^{-1/2}$ in a linear model with Gaussian covariates and independent noise, either non-asymptotically when $p/n\le γ$ or asymptotically in the high-dimensional asymptotic regime $p/n\toγ'\in(0,\infty)$. General differentiable loss functions $ρ$ are allowed provided that $ψ=ρ'$ is 1-Lipschitz. The validity of the out-of-sample error estimate holds either under a strong convexity assumption, or for the $\ell_1$-penalized Huber M-estimator if the number of corrupted observations and sparsity of the true $β$ are bounded from above by $s_*n$ for some small enough constant $s_*\in(0,1)$ independent of $n,p$. For the square loss and in the absence of corruption in the response, the results additionally yield $n^{-1/2}$-consistent estimates of the noise variance and of the generalization error. This generalizes, to arbitrary convex penalty, estimates that were previously known for the Lasso.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pierre C Bellec. 2023-03-30. Out-of-sample error estimate for robust M-estimators with convex penalty. https://arxiv.org/abs/2008.11840

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A global spectral gap for Metropolis-adjusted Langevin algorithm with a uniformly randomized step size

Let $π(\mathrm{d} x)\propto e^{-U(x)}\,\mathrm{d} x$ on $\mathbb{R}^d$, where $U$ is continuously differentiable and $m$-strongly convex with a globally $L$-Lipschitz gradient, $0<m\leq L<\infty$, and $κ=L/m$. It is known that, under warm-start assumptions, fixed-step Metropolis-adjusted Langevin algorithm (MALA) with properly tuned step size has mixing time of order $κ\sqrt{d}$ up to logarithmic factors. By contrast, when the condition number is bounded away from one, no single fixed step size yields a matching spectral-gap lower bound of order $(κ\sqrt{d})^{-1}$ uniformly over this target class. We establish a global spectral-gap lower bound for MALA with a uniformly randomized step size. At each iteration, the algorithm draws $h$ uniformly from $(0,H)$ and performs one ordinary MALA transition. Choosing $H$ of order $$\frac{1}{L\sqrt{d[1+\log(d+1)+\logκ]}}$$ yields a right spectral-gap lower bound of order $$\frac{1}{κ\sqrt{d[1+\log(d+1)+\logκ]}},$$ uniformly over the target class. A key ingredient in the proof is a Cheeger-type inequality for aggregating estimates of the one-step flow of MALA out of measurable sets at various step-size scales. It allows the scale used to control the flow to depend on the set and avoids the additional loss that would result from first estimating the conductance and then applying the standard Cheeger inequality. This work was developed with substantial assistance from ChatGPT, which suggested the uniformly randomized-step approach, developed the principal proof arguments, and generated the simulation and Lean 4 code. The human author checked and verified the mathematical content and takes full responsibility for the results.

math.ST

Evidence thresholds of collapsed structured block models

A stochastic block model explains a network by splitting its nodes into communities and a matrix of tie probabilities between them. Collapsing those probabilities out scores each split by one number, its marginal likelihood, and every decision the model takes, how many communities, whether a weak one survives, what generates the ties inside it, is a comparison of two such numbers. The comparisons fail in known ways: weak structure is dropped, small communities are merged, a group of hubs is returned in place of the real groups. Each failure has a threshold, and the thresholds have not been computed. We compute them. We prove that the best single partition keeps a planted structure only above the first-moment bound, which sits above the Kesten-Stigum threshold for up to ten communities, while the evidence summed over all partitions switches at Kesten-Stigum itself: between the two, the posterior favours structure that every partition a sampler can report rejects. We prove three further thresholds, one for merging small communities, one for the density at which the representation alone decides, one for the degree prior that cuts a community along its degrees. The empirical finding confirms the theory: on planted partitions and on thirteen labelled networks, from the karate club to citation and co-purchase networks of up to 19,717 nodes, each threshold falls where it is predicted, against likelihood, spectral, penalised and nonparametric alternatives.

math.ST

On the consistency of the posterior distribution for nonlinear PDE parameter identification

In this work, we investigate the estimation of a parameter in PDEs using Bayesian procedures, and focus on posterior distributions constructed using Gaussian process priors, and its variational approximation. We establish contraction rates for the posterior distribution and the variational approximation in the regime of low-regularity parameters. Specifically, the ground truth is only assumed to belong to the support space of the prior rather than to its associated reproducing kernel Hilbert space. Also we derive contraction rates for the posterior distribution constructed using the randomly truncated Gaussian process prior. The analysis relies on a delicate approximation argument that approximates low-regularity ground truths by suitable elements in the reproducing kernel Hilbert space and balances various error sources. We illustrate the general theory on three nonlinear inverse problems for PDEs.

math.ST