Search arXiv⌕ Search

arXiv · 2204.04386

Efficient Derivative-free Bayesian Inference for Large-Scale Inverse Problems

Abstract

We consider Bayesian inference for large scale inverse problems, where computational challenges arise from the need for repeated evaluations of an expensive forward model. This renders most Markov chain Monte Carlo approaches infeasible, since they typically require $O(10^4)$ model runs, or more. Moreover, the forward model is often given as a black box or is impractical to differentiate. Therefore derivative-free algorithms are highly desirable. We propose a framework, which is built on Kalman methodology, to efficiently perform Bayesian inference in such inverse problems. The basic method is based on an approximation of the filtering distribution of a novel mean-field dynamical system into which the inverse problem is embedded as an observation operator. Theoretical properties of the mean-field model are established for linear inverse problems, demonstrating that the desired Bayesian posterior is given by the steady state of the law of the filtering distribution of the mean-field dynamical system, and proving exponential convergence to it. This suggests that, for nonlinear problems which are close to Gaussian, sequentially computing this law provides the basis for efficient iterative methods to approximate the Bayesian posterior. Ensemble methods are applied to obtain interacting particle system approximations of the filtering distribution of the mean-field model; and practical strategies to further reduce the computational and memory cost of the methodology are presented, including low-rank approximation and a bi-fidelity approach. The effectiveness of the framework is demonstrated in several numerical experiments, including proof-of-concept linear/nonlinear examples and two large-scale applications: learning of permeability parameters in subsurface flow; and learning subgrid-scale parameters in a global climate model from time-averaged statistics.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Daniel Zhengyu Huang, Jiaoyang Huang, Sebastian Reich, Andrew M. Stuart. 2022-08-11. Efficient Derivative-free Bayesian Inference for Large-Scale Inverse Problems. https://arxiv.org/abs/2204.04386

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Error Estimates for the Arnoldi Approximation of a Matrix Square Root

The Arnoldi process provides an efficient framework for approximating functions of a matrix applied to a vector, i.e., of the form $f(M)\bm{b}$, by repeated matrix-vector multiplications. In this paper, we derive error estimates for approximating the action of a matrix square root using the Arnoldi process, where the integral representation of the error is reformulated in terms of the error for solving the linear system $M\bm{x}=\bm{b}$. The results extend the error analysis of the Lanczos method for Hermitian matrices in [Chen et al., SIAM J. Matrix Anal. Appl., 2022] to non-Hermitian cases and provide an improved bound for the Hermitian case. Furthermore, in practical settings, the matrix may only be available via approximate or structured representations. Motivated by this, we extend the analysis and establish a generalized error bound for perturbed matrices. The numerical results on matrices with different structures demonstrate that our theoretical analysis yields a reliable upper bound. Finally, simulations on large-scale matrices arising in particulate suspensions, represented in hierarchical matrix form, validate the effectiveness and practicality of the approach.

math.NA↗

Study of a TPFA scheme for the stochastic Allen-Cahn problem with constraint through numerical experiments

We study the properties of a fully discrete numerical scheme for the stochastic Allen-Cahn problem with constraint on a bounded polygonal domain in two or three dimensions with homogeneous Neumann boundary condition. The scheme under consideration is of Two Point Flux Approximation (TPFA) type with respect to space and of semi-implicit Euler-Maruyama type with respect to time. The constraint is implemented by a multivalued, subdifferential operator which makes the problem a differential inclusion. This operator is incorporated into the scheme via its Yosida approximation, at the expense of introducing an additional regularization parameter. From previous studies it is known that the time discretization of the Yosida approximation has to be implicit in order to obtain a convergent scheme. Consequently, a non-linear and non-smooth equation has to be solved at each time step and this is challenging for the computation of approximate solutions. In this contribution, we introduce a splitting method to compute the solutions of the TPFA approximations. This method allows us to solve a linear equation in a first step. Then, a piecewise affine projection operator coming from the Yosida approximation can be computed explicitly in a second step. We quantify the error coming from the splitting method and show that the splitting method is accurate. We study properties of the scheme and provide convergence rates through numerical experiments.

math.NA↗

Spectral density estimation for normal matrices

The spectral density estimation problem asks for an algorithm that, given an $n\times n$ matrix $A$, outputs a probability measure that is a good approximation to the uniform distribution on the eigenvalues of $A$, called the spectral density of $A$. This paper considers the setting where $A$ is a large normal matrix that is accessible only through matrix-vector product queries. We provide an algorithm that makes just $m$ matrix-vector queries to $A$ and returns, with high probability, a measure within earth mover's distance $O(1/m+\log m/{\sqrt n})$ of the true spectral density of $A$. We provide a complementary lower bound that any algorithm producing an $\varepsilon$-approximation to the true spectral density for large matrices must make $Ω(1/\varepsilon)$ matrix-vector queries. The lower bound holds even for the more restricted case of real symmetric input matrices. In combination with our upper bound, it shows that spectral density estimation is essentially no harder for complex normal matrices than for real symmetric matrices.

math.NA↗