Search arXivSearch

arXiv · 2206.01246

Stochastic gradient descent introduces an effective landscape-dependent regularization favoring flat solutions

Abstract

Generalization is one of the most important problems in deep learning (DL). In the overparameterized regime in neural networks, there exist many low-loss solutions that fit the training data equally well. The key question is which solution is more generalizable. Empirical studies showed a strong correlation between flatness of the loss landscape at a solution and its generalizability, and stochastic gradient descent (SGD) is crucial in finding the flat solutions. To understand how SGD drives the learning system to flat solutions, we construct a simple model whose loss landscape has a continuous set of degenerate (or near degenerate) minima. By solving the Fokker-Planck equation of the underlying stochastic learning dynamics, we show that due to its strong anisotropy the SGD noise introduces an additional effective loss term that decreases with flatness and has an overall strength that increases with the learning rate and batch-to-batch variation. We find that the additional landscape-dependent SGD-loss breaks the degeneracy and serves as an effective regularization for finding flat solutions. Furthermore, a stronger SGD noise shortens the convergence time to the flat solutions. However, we identify an upper bound for the SGD noise beyond which the system fails to converge. Our results not only elucidate the role of SGD for generalization they may also have important implications for hyperparameter selection for learning efficiently without divergence.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ning Yang, Chao Tang, Yuhai Tu. 2022-06-02. Stochastic gradient descent introduces an effective landscape-dependent regularization favoring flat solutions. https://doi.org/10.1103/physrevlett.130.237101

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

R-transforms for non-Hermitian matrices: a spherical integral approach

In this paper, we establish a connection between the formalism of $\mathcal{R}$-transforms for non-Hermitian random matrices and the framework of spherical integrals, using the replica method. This connection was previously proved in the Hermitian setting and in the case of bi-invariant random matrices. We show that the $\mathcal{R}$-transforms used in the non-Hermitian context in fact originate from a single scalar function of two variables. This provides a new and transparent way to compute $\mathcal{R}$-transforms, which until now had been known only in restricted cases such as bi-invariant, Hermitian, or elliptic ensembles.

cond-mat.dis-nn

Spectral boundaries of deterministic matrices deformed by rotationally invariant random non-Hermitian ensembles

One of the great miracles of random matrix theory is that, in the $N \to \infty$ limit, many otherwise intractable matrix problems with horrendously complicated finite-$N$ expressions admit remarkably simple and elegant asymptotic solutions. In this paper, we illustrate this phenomenon in the context of spectral boundaries (or spectral edges) for deformed random matrices. Specifically, we consider matrices of the form $\mathbf{A} + \mathbf{B}$, where $\mathbf{A}$ is a deterministic $N\times N$ matrix (not necessarily Hermitian) and $\mathbf{B}$ is a rotationally invariant random matrix. In the large-$N$ limit, we show that the complex eigenvalue distribution of $\mathbf{A} + \mathbf{B}$ satisfies remarkably simple boundary equations that depend on the $\mathcal{R}_1$ and $\mathcal{R}_2$ transforms of $\mathbf{B}$. We illustrate our results on several explicit random matrix ensembles and support them with numerical simulations.

cond-mat.dis-nn

Electrical conductivity of crack-template-based transparent conducting films: mean-field approximation, effective-medium theory, and simulation

In this work, crack-template-based transparent conducting films were modeled as networks corresponding to the edges of a two-dimensional Poisson--Voronoi diagram. Two types of networks were considered: the original one, in which the conductance of each edge was inversely proportional to its length, and the effective one, in which all edges had the same conductance obtained from the effective-medium theory. The mean-field approximation was used for analytical evaluation of the electrical conductivity. Direct numerical calculations for the Poisson--Voronoi diagram showed that the mean-field approximation overestimated the effective conductivity of the original network by approximately 13\%, and of the effective network by 79\%. In addition, a honeycomb network with an edge conductance distribution corresponding to the Poisson--Voronoi diagram was studied: for it, the predictions of the effective-medium theory turned out to be more accurate than for the Poisson--Voronoi diagram, which was explained by the greater structural homogeneity of the periodic honeycomb lattice. The results indicate that, when modeling crack-template-based transparent conducting films, the application of the mean-field approximation may lead to significant errors if the resistance of individual conductors is not simply proportional to their length. This possibility is discussed as a motivation for future studies of hierarchical cracks with variable width, which are not directly investigated here.

cond-mat.dis-nn