Search arXiv⌕ Search

arXiv · 2204.05003

Local convergence rates of the nonparametric least squares estimator with applications to transfer learning

Abstract

Convergence properties of empirical risk minimizers can be conveniently expressed in terms of the associated population risk. To derive bounds for the performance of the estimator under covariate shift, however, pointwise convergence rates are required. Under weak assumptions on the design distribution, it is shown that least squares estimators (LSE) over 1-Lipschitz functions are also minimax rate optimal with respect to a weighted uniform norm, where the weighting accounts in a natural way for the non-uniformity of the design distribution. This implies that although least squares is a global criterion, the LSE adapts locally to the size of the design density. We develop a new indirect proof technique that establishes the local convergence behavior based on a carefully chosen local perturbation of the LSE. The obtained local rates are then applied to analyze the LSE for transfer learning under covariate shift.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Johannes Schmidt-Hieber, Petr Zamolodtchikov. 2023-12-29. Local convergence rates of the nonparametric least squares estimator with applications to transfer learning. https://arxiv.org/abs/2204.05003

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Optimal Transport Based Testing in Factorial Designs

We introduce a general framework for testing statistical hypotheses in factorial designs for probability measures supported on discrete spaces. The suggested methodology is based on the pairwise comparison of measures using optimal transport (OT). The formulation of hypotheses is intuitive: It is a direct extension of those underlying the analysis of variance (ANOVA) and its nonparametric counterparts to test for linear relationships between (discrete) probability measures in factorial designs. To this end, means or cumulative distribution functions simply will be replaced by measures. We derive under the null hypotheses and under (local) alternatives the asymptotic distribution of the corresponding empirical OT test statistic, which is the optimal value of a linear program with random objective function. It turns out that this requires to extend existing techniques from probability measures to signed measures, and we show directional Hadamard differentiability and the validity of the functional delta method. We discuss computational issues, permutation and bootstrap tests, and back up our findings with simulations. We illustrate our methodology on datasets from cellular biophysics and from biometric fingerprint identification.

math.ST↗

Stable Central Limit Theorems for Discrete-Time Lag Martingale Difference Arrays: Applications to Dynamic Causal Inference

Recent work in dynamic causal inference introduced a class of discrete-time stochastic processes that generalize martingale difference sequences and arrays as follows: the random variates in each sequence have expectation zero given certain lagged filtrations but not given the natural filtration. We formalize this class of stochastic processes and prove stable central limit theorems (CLTs) via martingale-coboundary decomposition, leveraging the classical martingale CLT. We develop a variety of sufficient conditions, including conditions under which the limiting variance has a simple form that depends on variances and covariances of neighboring variates. We demonstrate the application of these results to inference for time-averaged treatment effects in switchback designs and present a simulation study supporting their validity. The CLTs enable various extensions to existing methodology for design-based approaches to dynamic causal inference, including time-lagged effects, random limiting variances, cross-unit dependence, and vector-valued estimands.

math.ST↗

Exact Likelihood-Coin Poisson Sampling for Bayesian Inverse Problems with Sharp Complexity Bounds

We develop an exact posterior-sampling framework for Bayesian inverse problems when selected bounded forward observables can be accessed through Bernoulli events. A Bernstein--Poisson construction converts these forward coins into scaled Gaussian likelihood coins, and thinning an inflated prior Poisson point process yields posterior atoms that are iid conditional on their number. For independent Gaussian observations, we derive an exact mean-work identity for the implemented early-stopped factory and sharp small-noise complexity laws governed by local prior-predictive mass near the exact-fit set; a factorial-moment construction extends the likelihood factory to correlated Gaussian errors. The posterior algorithm is model-agnostic once Bernoulli access is available. As one continuum realization, we use Feynman--Kac sampling, Poisson killing, and lazy random-series evaluation for a bounded elliptic resolvent problem with a function-valued coefficient. Numerical experiments validate the forward and likelihood coins, support the predicted work regimes, and demonstrate posterior sampling without deterministic spatial discretization or fixed parameter truncation in the target. A matched-accuracy benchmark against finite-difference prior rejection illustrates how deterministic discretization bias changes the posterior accuracy--cost balance.

math.ST↗