Search arXivSearch

arXiv · 2008.11778

Anderson Acceleration for Seismic Inversion

Abstract

The state-of-art seismic imaging techniques treat inversion tasks such as FWI and LSRTM as PDE-constrained optimization problems. Due to the large-scale nature, gradient-based optimization algorithms are preferred in practice to update the model iteratively. Higher-order methods converge in fewer iterations but often require higher computational costs, more line search steps, and bigger memory storage. A balance among these aspects has to be considered. We propose using Anderson acceleration (AA), a popular strategy to speed up the convergence of fixed-point iterations, to accelerate the steepest descent algorithm, which we innovatively treat as a fixed-point iteration. Independent of the dimensionality of the unknown parameters, the computational cost of implementing the method can be reduced to an extremely low-dimensional least-squares problem. The cost can be further reduced by a low-rank update. We discuss the theoretical connections and the differences between AA and other well-known optimization methods such as L-BFGS and the restarted GMRES and compare their computational cost and memory demand. Numerical examples of FWI and LSRTM applied to the Marmousi benchmark demonstrate the acceleration effects of AA. Compared with the steepest descent method, AA can achieve fast convergence and provide competitive results with some quasi-Newton methods, making it an attractive optimization strategy for seismic inversion.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yunan Yang. 2020-11-13. Anderson Acceleration for Seismic Inversion. https://arxiv.org/abs/2008.11778

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Cost-Accuracy Trade-offs: Neural Operator vs Classical Numerical Solver

Neural operators are data-driven models that learn mappings from inputs that parameterize partial differential equations, such as spatially varying coefficients, initial conditions, forcing terms, boundary conditions, or geometries, to solution fields or quantities of interest. Once trained, they can serve as surrogates for classical numerical solvers in many-query settings that require repeated evaluations for varying inputs. We address the question of when, and then why, neural operator surrogates outperform classical numerical solvers, in terms of cost for a given accuracy. We focus on the post-training, many-query limit, in which data-acquisition and training costs are treated as fixed and fully amortized. Even in this deliberately favorable regime for neural operators, there are regimes in which classical solvers outperform the surrogate models. We compare the cost-accuracy performance of neural operator surrogates and classical numerical solvers through a reproducible benchmark study comparing neural operators with problem-matched classical solvers on representative problems in computational science and engineering, focusing on prediction error, per-query floating-point cost, and wall-clock runtime. Neural operators are most competitive at low-to-moderate accuracy requirements. Their floating-point cost advantage depends strongly on the problem structure, arising when they avoid temporal or nonlinear iterations or predict a reduced quantity of interest rather than a full solution field. Additional wall-clock speedups result from dense tensor operations that are well suited to modern hardware. As the target accuracy is tightened, achieving the required accuracy with neural operators becomes increasingly challenging, and classical solvers outperform surrogates in this regime; thus classical solvers will remain important for verification and high-accuracy computation.

math.NA

Learning to Integrate

We address uncertainty quantification for a generic input distribution to some resource-intensive simulation, e.g., requiring the solution of a partial differential equation. While efficient numerical methods exist to compute integrals for high-dimensional Gaussian and other separable distributions based on sparse grids (SG), input data arising in practice often does not fall into this class. We therefore employ transport maps to transform complex distributions to multivariate standard normals. In generative learning, a number of neural network architectures have been introduced that accomplish this task approximately. Examples are affine coupling flows (ACF) and ordinary differential equation-based networks such as conditional flow matching (CFM). To compute the expectation of a quantity of interest, we numerically integrate the composition of the inverse of the learned transport map with the simulation code output. As this map is integrated over a multivariate Gaussian distribution, SG techniques can be applied. Viewing the images of the SG quadrature nodes as learned quadrature nodes for a given complex distribution motivates our title. We demonstrate our method for monomials of total degrees for which the unmapped SG rules are exact. We also apply our approach to the stationary diffusion equation with coefficients modeled by exponentiated Lévy random fields, using modal expansions of Karhunen--Loève type with 9 and 25 modes. In a series of numerical experiments, we investigate errors due to learning accuracy, quadrature, statistical estimation, truncation of the modal series of the input random field, and training data size for three normalizing flows (ACF, conditional flow matching and optimal transport flow matching). We discuss the mathematical assumptions on which our approach is based and demonstrate its shortcomings when these are violated.

math.NA

Variational estimates for Schur complements of symmetric saddle point problems

We present sharp variational estimates for the extremal eigenvalues of Schur complements arising in symmetric saddle point problems. These estimates make explicit the role of right inverses with controlled energy in lower bounds and of null space corrections in upper bounds. The same argument applies to projected Schur complements when the primal operator is semidefinite. Applications to the augmented Lagrangian method, mixed finite element methods, and nonoverlapping domain decomposition methods yield concise proofs of established convergence estimates.

math.NA