Search arXiv⌕ Search

arXiv · 2610.03666

Bayesian Operator Learning: Posterior Existence and Convergence of Point Estimates for Gaussian Priors

Abstract

We develop a Bayesian framework for learning nonlinear operators between infinite-dimensional spaces. Given a map $g_0:\mathcal{X}\to\mathcal{Y}$ between separable Hilbert spaces, we study the recovery of $g_0$ from $n\in\mathbb{N}$ noisy input-output pairs $(\boldsymbol{X},\boldsymbol{Z})=(X_i,Z_i)_{i=1}^n$ with $Z_i= g_0 (X_i ) + E_i$. Here the $X_i\in\mathcal{X}$ are randomly drawn 'design' points in a compact subset of $\mathcal X$, and the $E_i$ are assumed to be i.i.d. draws from a Gaussian white noise process indexed by $\mathcal{Y}$. For any 'operator-valued' prior $\mathbb{P}_G$ supported on the space of continuous operators, we show existence of the posterior $\mathbb P_{G|(\boldsymbol{X},\boldsymbol{Z})}$ as a regular conditional distribution, and provide a characterization of its Radon-Nikodym derivative. For Gaussian priors, we establish algebraic (in the sample size $n$) convergence rates for the posterior mean towards the ground truth; this corresponds to a ridge regularized kernel estimator. Moreover, we show that the posterior mean is minimax optimal (up to logarithmic factors) over hyperrectangles when the smoothness of the prior matches that of the ground truth. To illustrate the applicability of our analysis, we derive explicit learning rates for the Darcy flow solution operator.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Niklas Reinhardt, Jakob Zech. 2026-10-02. Bayesian Operator Learning: Posterior Existence and Convergence of Point Estimates for Gaussian Priors. https://arxiv.org/abs/2610.03666

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Structural Causal Models for Extremes: an Approach Based on Exponent Measures

We introduce a new formulation of structural causal models for extremes, called the extremal structural causal model (eSCM). Unlike conventional structural causal models, where randomness is governed by a probability distribution, eSCMs use an exponent measure, an infinite-mass law that naturally arises in the analysis of multivariate extremes. Central to this framework are activation variables, which abstract the single-big-jump principle, along with additional randomization that enriches the class of eSCM laws. The eSCM provides a common foundation for the two main existing approaches to extremal causal modeling, the max-linear and the sum-linear structural causal models, recovering both as special cases within a single asymptotic formulation. More broadly, it encompasses all possible laws of directed graphical models under the recently introduced notion of extremal conditional independence. We also identify an inherent asymmetry in eSCMs under natural assumptions, enabling the identifiability of causal directions, a central challenge in causal inference. Finally, we propose a method that exploits this causal asymmetry and demonstrate its effectiveness on both simulated and real datasets.

math.ST↗

Consistency of penalized maximum likelihood estimation for multidimensional change-in-velocity detection

We establish statistical consistency for the penalized maximum-likelihood estimator underlying CPLASS, a method for detecting changes in velocity in $d$-dimensional time-series data. The signal is modeled as a continuous piecewise-linear trajectory observed with independent Gaussian noise. Unlike classical change-in-mean models, continuity across adjacent segments couples their parameters and prevents direct application of standard segmentation arguments. Under compact parameter spaces, minimum segment-length and velocity-jump conditions, and a strengthened Schwarz information criterion penalty $ρ_k(\log n)^γ$ with $γ>1$, we prove joint consistency of the estimated number of segments and all changepoint locations in the high-frequency regime $n\to\infty$. The maximum changepoint-location error is $O_{\mathbb{P}}\{(\log n/n)^{1/2}\}$. The proof first uses empirical-process theory to establish convergence of the fitted signal, variance, and likelihood for each fixed model size. It then combines overfitting and underfitting arguments with a two-stage geometric localization analysis, yielding an initial $(\log n/n)^{1/3}$ rate that is sharpened by exploiting the local two-segment continuous piecewise-linear structure.

math.ST↗

Optimal Community Recovery by Spectrally Initialized Variational EM in General Stochastic Block Models

We prove a Chernoff-exponent guarantee for the output of spectrally initialized batch variational EM after a prescribed iteration budget, without assuming global maximization of its variational objective. The iteration repeatedly estimates the entire block probability matrix and community proportions from the same sparse graph. We consider a fixed number of communities with proportions bounded away from zero and fixed, positive, distinct connectivity profiles; neither assortativity nor full rank is required. The algorithm uses simultaneous softmax updates without sample splitting or posterior thresholding. Its analysis must control the feedback between estimated parameters, soft labels and reused edges at an exponentially small risk scale. We establish a uniform one-step bound over data-dependent soft assignments whose random remainder has exponentially small expectation. Combined with an exponentially reliable regularized spectral initializer, this yields an unconditional expected misclassification rate bounded by $\exp\{-(1-o(1))J_n\}$ throughout the sparse, diverging-degree regime, where $J_n$ is the minimum nodewise Chernoff information. Matching lower bounds establish first-order logarithmic minimax optimality on local parameter spaces allowing unknown connectivity and varying community counts. The algorithm also attains the sharp first-order exact-recovery threshold. Numerical experiments illustrate refinement gains and sensitivity to initialization and imbalance.

math.ST↗