Search arXivSearch

arXiv · 1806.05294

Hessian spectrum at the global minimum of high-dimensional random landscapes

Abstract

Using the replica method we calculate the mean spectral density of the Hessian matrix at the global minimum of a random $N \gg 1$ dimensional isotropic, translationally invariant Gaussian random landscape confined by a parabolic potential with fixed curvature $μ>0$. Simple landscapes with generically a single minimum are typical for $μ>μ_{c}$, and we show that the Hessian at the global minimum is always {\it gapped}, with the low spectral edge being strictly positive. When approaching from above the transitional point $μ= μ_{c}$ separating simple landscapes from 'glassy' ones, with exponentially abundant minima, the spectral gap vanishes as $(μ-μ_c)^2$. For $μ<μ_c$ the Hessian spectrum is qualitatively different for 'moderately complex' and 'genuinely complex' landscapes. The former are typical for short-range correlated random potentials and correspond to 1-step replica-symmetry breaking mechanism. Their Hessian spectra turn out to be again gapped, with the gap vanishing on approaching $μ_c$ from below with a larger critical exponent, as $(μ_c-μ)^4$. At the same time in the 'most complex' landscapes with long-ranged power-law correlations the replica symmetry is completely broken. We show that in that case the Hessian remains gapless for all values of $μ<μ_c$, indicating the presence of 'marginally stable' spatial directions. Finally, the potentials with {\it logarithmic} correlations share both 1RSB nature and gapless spectrum. The spectral density of the Hessian always takes the semi-circular form, up to a shift and an amplitude that we explicitly calculate.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yan V Fyodorov, Pierre Le Doussal. 2018-06-16. Hessian spectrum at the global minimum of high-dimensional random landscapes. https://doi.org/10.1088/1751-8121%2Faae74f

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction

Using a statistical mechanics framework, we investigate the parameter space of transformer models trained on protein sequence data. We sample the loss landscape at varying temperatures using Langevin dynamics to characterize the low-loss manifold, and to understand the mechanisms underlying transformers' superior performance in protein structure prediction. We find that, at variance with networks not based on the attention mechanism, the lack of a first--order--like transition in the loss of the transformer produces a range of intermediate temperatures with good learning properties; this is true both for synthetic and natural protein sequences. We also show that the parameters of most layers are highly conserved at these temperatures if the dimension of the embedding is optimal, and we provide an operative way to find this dimension. Additionally, we show that the attention matrix is more predictive of the contact maps of the protein at higher temperatures and for higher dimensions of the embedding than those optimal for learning. Finally, we showed that the models sampled at intermediate temperatures can predict the free-energy variation upon mutation, better than models obtained through standard optimization techniques.

cond-mat.dis-nn

The Cross-Substrate Access Assay: What an Indicator Test Must Declare to Travel from Brain to Language Model

Testing an artificial system for a property linked to consciousness means applying a measurement developed on brains to a system that is not one. Such a transfer must re-examine five parts of the procedure: the competing statistical models, how they are fitted, the unit the inference generalizes over, the quantity the uncertainty interval is about, and the rule that turns a result into a verdict. The Cross-Substrate Access Assay declares all five. Because brain and model signals share no physical scale, every model is scored by the cross-entropy it assigns to held-out data, in nats per trial. The test case is the global neuronal workspace theory, which predicts that near threshold a stimulus either enters a capacity-limited workspace or does not, so that single-trial responses form a mixture of two states. A published test of this prediction on twenty people's electroencephalograms partly reproduces in a re-implementation: the first crossing and the broad ordering over time match, the window-by-window agreement does not. On 12,000 synthetic datasets generated with a single graded state, all of them members of the families the procedure fits and none within 0.0067 nat per trial of the decision boundary, the two models of that test carried over unchanged reported two states in 989 and the expanded families in none; on 600 datasets carrying a mixture the expanded procedure reported two states in 599. Its nominal 95% interval contained the procedure's mean result less often than the required 90% at six of twelve graded settings. No claim about experience is made.

cond-mat.dis-nn

Measure-zero delocalization in the complex plane: exact mobility arcs in a non-Hermitian off-diagonal quasiperiodic lattice

We investigate Anderson localization in a one-dimensional lattice with non-Hermitian off-diagonal quasiperiodic disorder, extending a recently studied Hermitian mosaic model to the non-Hermitian regime. Using Avila's global theory, we derive the exact Lyapunov exponent and the complete phase diagram in the complex energy plane. This work contains two central findings. First, we discover mobility arcs---open curved segments in the complex plane---as a new class of mobility edges and the generic form of open mobility edges, which coexist with closed mobility rings in a complementary parameter regime. These arcs share the same localization physics as the previously reported mobility lines: eigenstates are delocalized if and only if their energies lie exactly on these sets; any deviation yields localized states. This constitutes a striking measure-zero delocalization phenomenon: delocalized states occupy only zero-measure sets (arcs or lines) in the complex plane, in sharp contrast to the mobility rings, which enclose a finite-area region of delocalized states. Second, we reveal that mobility rings, arcs, and lines all share a common mathematical origin in the generalized Joukowski transformation $P(E) = \frac{1}{2}(u - w^2/u)$, rooted in the algebraic structure of the underlying polynomial: the preimage of the boundary of an elliptical region under the polynomial map $P(E)$ gives the rings, while the branch cut inside this ellipse gives rise to the mobility arcs and lines in the complementary parameter regime.

cond-mat.dis-nn