Search arXiv⌕ Search

arXiv · 2609.35716

When does least-squares residual minimization solve a differential equation?

Abstract

Least-squares residual minimization approximates a differential equation through an optimization problem. This paper investigates the conditions required to transition from stationary points and minimizing sequences of the residual functional to the exact solution. Starting from the classical orthogonality relation between the residual and the range of the linearized operator, we study how boundary conditions and parameterized variations determine what can be inferred about the residual at a stationary point. Smooth spurious critical points can occur even for well-posed problems; for linear equations, a compatibility condition yields an exact projection characterization. For a scalar conservation law with transonic boundary data, every minimizing sequence converges to a continuous function that does not solve the equation, even when a stationary entropy shock exists. We give conditions under which approaching this limit forces the parameter norms to diverge. A Hamilton--Jacobi example shows that a function can have zero residual without being the viscosity solution. Favorable results for a viscous model and uniformly convex Monge--Ampère critical points identify conditions that restore the corresponding implications. For piecewise-smooth trial families, the regularity needed to define the residual functional does not by itself justify the sampled gradient used in training. Moving residual jumps can contribute terms absent from that gradient; an explicit example has identically vanishing sampled gradients and a nonzero continuous derivative. Together, these results clarify what stationarity and sampled gradients imply about the residual, and under which conditions minimizing sequences lead to the intended solution.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Carlos Esteve-Yagüe, Richard Tsai, Stanley Osher. 2026-09-28. When does least-squares residual minimization solve a differential equation?. https://arxiv.org/abs/2609.35716

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Accelerating Natural Gradient Descent for PINNs with Randomized Numerical Linear Algebra

Natural Gradient Descent (NGD) has emerged as a promising optimization algorithm for training neural network-based solvers for partial differential equations (PDEs), such as Physics-Informed Neural Networks (PINNs). However, its practical use is often limited by the high computational cost of solving linear systems involving the Gramian matrix. While matrix-free NGD methods based on the conjugate gradient (CG) method avoid explicit matrix inversion, the ill-conditioning of the Gramian significantly slows the convergence of CG. In this work, we extend matrix-free NGD to broader classes of problems than previously considered and propose the use of Randomized Numerical Linear Algebra (RandNLA) techniques for efficient preconditioning of the inner CG solver. The resulting algorithms demonstrate substantial performance improvements over existing NGD-based methods on a range of PDE problems discretized using neural networks, and offer competitive results compared to other state-of-the-art optimizers.

math.NA↗

Helmholtz boundary integral methods and the pollution effect

This paper is concerned with solving the Helmholtz exterior Dirichlet and Neumann problems with large wavenumber $k$ and smooth obstacles using the standard second-kind boundary integral equations. We consider Galerkin and collocation methods -- with subspaces consisting of $\textit{either}$ piecewise polynomials (in 2-d for collocation, in any dimension for Galerkin) $\textit{or}$ trigonometric polynomials (in 2-d) -- as well as a fully discrete quadrature (Nyström) method based on trigonometric polynomials (in 2-d). For each of these methods, we prove -- in many cases for the first time -- rigorous results about the fundamental question: how quickly must the number of degrees of freedom (the dimension of the approximation space) grow with $k$ to maintain accuracy of the computed solution? Importantly, we determine which of these methods suffer from $\textit{the pollution effect}$. That is, we address the question: must the number of points per wavelength $\to \infty$ to maintain accuracy as $k\to\infty$?

math.NA↗

DE-Sinc approximation for unilateral rapidly decreasing functions and its computable error bound

The Sinc approximation is highly effective for functions that decay rapidly at both ends of the real axis. For unilateral rapidly decreasing functions, which decay algebraically as $t\to-\infty$ and exponentially as $t\to\infty$, an appropriate variable transformation is required. Existing single-exponential transformations for this class of functions yield only root-exponential convergence, even when improved transformations are employed. This paper develops a double-exponential (DE)-Sinc approximation based on the transformation $t=2\sinh(\log(\log(1+\exp(π\sinh x))))$, which was previously introduced for numerical integration. The main contribution is a rigorous and computable error bound of order $\operatorname{O}(\exp(-cn/\log n))$ with a constant explicitly expressed in terms of the problem parameters. Under the stated assumptions, the resulting approximation achieves almost exponential convergence and is suitable for computation with guaranteed accuracy. Numerical examples satisfying these assumptions confirm the predicted convergence behavior and the validity of the derived error bound.

math.NA↗