Search arXiv⌕ Search

arXiv · 2511.20322

Modified Equations for Stochastic Optimization

Abstract

In this thesis, we extend the recently introduced theory of stochastic modified equations (SMEs) for stochastic gradient optimization algorithms. In Ch. 3 we study time-inhomogeneous SDEs driven by Brownian motion. For certain SDEs we prove a 1st and 2nd-order weak approximation properties, and we compute their linear error terms explicitly, under certain regularity conditions. In Ch. 4 we instantiate our results for SGD, working out the example of linear regression explicitly. We use this example to compare the linear error terms of gradient flow and two commonly used 1st-order SMEs for SGD in Ch. 5. In the second part of the thesis we introduce and study a novel diffusion approximation for SGD without replacement (SGDo) in the finite-data setting. In Ch. 6 we motivate and define the notion of an epoched Brownian motion (EBM). We argue that Young differential equations (YDEs) driven by EBMs serve as continuous-time models for SGDo for any shuffling scheme whose induced permutations converge to a det. permuton. Further, we prove a.s. convergence for these YDEs in the strongly convex setting. Moreover, we compute an upper asymptotic bound on the convergence rate which is as sharp as, or better than previous results for SGDo. In Ch. 7 we study scaling limits of families of random walks (RW) that share the same increments up to a random permutation. We show weak convergence under the assumption that the sequence of permutations converges to a det. (higher-dimensional) permuton. This permuton determines the covariance function of the limiting Gaussian process. Conversely, we show that every Gaussian process with a covariance function determined by a permuton in this way arises as a weak scaling limit of families of RW with shared increments. Finally, we apply our weak convergence theory to show that EBMs arise as scaling limits of RW with finitely many distinct increments.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Stefan Perko. 2025-11-25. Modified Equations for Stochastic Optimization. https://arxiv.org/abs/2511.20322

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Selection in the Zero-Noise Limit for High-Dimensional Diffusions with Measurable Drift

This paper investigates the zero-noise limit of high-dimensional small-noise diffusion processes governed by the stochastic differential equation (SDE) \begin{equation*} \mathrm{d}X_{t}^{\varepsilon }=b(X_{t}^{\varepsilon })\,\mathrm{d}t+\varepsilon \,\mathrm{d}W_{t},\quad X_{0}^{\varepsilon }=0,\quad \varepsilon >0, \end{equation*} where the drift coefficient $b$ is assumed to be measurable and bounded. Under this condition, the associated deterministic ordinary differential equation (ODE) $\dot{x}_{t}=b(x_{t})$ may admit multiple Filippov solutions, i.e., solutions in the sense of differential inclusions, due to the lack of Lipschitz continuity or uniqueness criteria. However, the introduction of non-degenerate additive noise restores well-posedness: the perturbed system admits a unique strong (pathwise) solution for each $\varepsilon >0$. Our analysis is based on the radial-spherical decomposition. We prove that, under certain assumptions, the radius dominates a one-dimensional comparison process and escapes to infinity at a linear rate, and the extrinsic angular martingale converges. Under an explicit separable gradient structure, expressed in terms of a radial potential $V(r,ω)=\langle b(rω),ω\rangle $, the angle converges, for every fixed noise, almost surely to the set of global angular maximizers of $V$, and its distance from that set tends to zero in probability as $\varepsilon\rightarrow0$. The limiting measure is compact, carried by immediate-departure Filippov paths, and, under a Lipschitz radial-graph condition, purely singular with Hausdorff dimension at most $d-1$. A continuous-mapping criterion for exit-time convergence is also established.

math.PR↗

Asymptotic optimality of dynamic first-fit packing on the half-axis

We revisit a classical problem in dynamic storage allocation. Items arrive in a linear storage medium, modeled as a half-axis, at a Poisson rate $r$ and depart after an independent exponentially distributed unit mean service time. The arriving item sizes (lengths) are assumed to be independent and identically distributed (i.i.d.) from a common distribution $H$. A widely employed algorithm for allocating the items is the "first-fit" discipline, namely, each arriving item is placed in the left-most vacant interval large enough to accommodate it. In a seminal 1985 paper, Coffman, Kadota, and Shepp ([6]) proved that in the special case of unit length items (i.e. degenerate $H$), as $r$ tends towards infinity, the first-fit algorithm is asymptotically optimal in the following sense: the steady-state ratio of expected "empty space" (gaps between items) to expected occupied space tends towards $0$. In a sequel to [6], Coffman, Kadota, and Shepp ([5]) conjectured that the first-fit discipline is also asymptotically optimal for non-degenerate $H$. In this paper we provide the first proof of first-fit asymptotic optimality for non-degenerate distributions $H$ of item sizes. Our main result is for the case when $H$ is concentrated on countably many positive real sizes forming an increasing sequence that is either finite or goes to infinity, with the average item size being finite. We prove that under the first-fit discipline, as $r$ tends towards infinity, the steady-state packing configuration (scaled down by $r$) converges in distribution to the limiting packing configuration with smaller items on the left, larger items on the right, and with no gaps between. In particular, this proves asymptotic optimality of first-fit in the following sense: if $P$ is the expected occupied space, then in steady-state the empty space (scaled down by $r$) in $[0,P]$ vanishes.

math.PR↗

Mind the jumps: well-posedness of semi-martingale 2BSDEs

We develop a well-posedness theory for second-order backward stochastic differential equations (2BSDEs) with jumps. This work covers two complementary notions of solution: an extrinsic and an intrinsic one. We also discuss the obstruction to aggregating jump integrands. The chosen framework allows for controlled diffusions with jumps, pure-jump processes, and discrete-time processes in a unified setting.

math.PR↗