Search arXiv⌕ Search

arXiv · 2609.31101

A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine

Abstract

The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee machine, where both the ERM estimator and the Hessian spectrum are analytically tractable in the proportional high-dimensional limit. First, we use a zero-temperature Gibbs formulation to obtain predictions for the observables of the typical minimizers of the empirical loss. Secondly, we use Edwards-Jones formalism to derive the limiting Hessian resolvent around these typical minimizers. All predictions agree with finite-size gradient-descent simulations. Finally, we study three measures of flatness, namely the left and right edges and the spectral mean, and check if a decrease in generalization error as the dataset size is increased corresponds to an increase in flatness. We find that the answer strongly depends on the learning task and on the ratio of the number of parameters to the number of data points. In regression, the spectral mean and right edge correlate with the generalization error, while the left edge does so only in the overparametrized regime. In classification this correlation reliably holds only in the highly overparametrized phase, while for underparametrized networks it can even reverse.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Brandon Livio Annesi, Davide Straziota, Enrico Maria Malatesta. 2026-09-25. A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine. https://arxiv.org/abs/2609.31101

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

On the robustness of noisy solutions in non-convex neural networks

Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite being relatively rare. At zero temperature this picture has been formalized in binary perceptrons through the overlap gap property (OGP), which limits algorithmic access to configurations with zero training error above a critical constraint density $α_{\rm OGP}$. Here we extend this description to finite temperature, where a positive training error is allowed and statistically penalized. We first show that the frozen one-step replica-symmetry-breaking solution, dominating the zero temperature equilibrium measure, survives at any finite temperature. We furthermore derive a general criterion, based on the smoothness of the single-pattern Gibbs weight near the decision boundary, that determines when a finite-temperature relaxation of the loss removes freezing. We then extend the OGP construction to finite temperature and show that dense, algorithmically accessible regions of finite-energy configurations persist beyond $α_{\rm OGP}$, up to a threshold $α_{\rm OGP}(ε)$ that grows with the allowed training error $ε$. Finally, in the teacher-student setting, we show that these wide, finite-energy regions still retain good generalization. Using a finite energy message-passing algorithm, we demonstrate numerically that thermal noise enables effective generalization in the regime of constraint densities where both recovering the teacher and finding a zero temperature solution are computationally hard.

cond-mat.dis-nn↗

Connecting finite-size scaling and renormalization-group flows in Anderson localization on random graphs

We investigate the relation between two complementary descriptions of the Anderson transition on small-world random graphs: finite-size scaling of eigenfunction moments and renormalization-group (RG) flows of multifractal dimensions $D_q$. The RG flows recover the qualitative structure previously reported for the information dimension $D_1$, while their extension to other $D_q$ with smaller moment orders $q<1/2$ confirms the strongly multifractal nature of the localized phase characterized by a clear separation between multifractal and localized behaviors at strong disorder. On the other hand, we show that we can construct another beta function characterizing the flow of eigenfunction moments, whose collapse onto a single function across different disorder strengths provides a clear confirmation of single-parameter scaling property. These results indicate that the family of RG trajectories observed for the information dimension $D_1$ need not imply two-parameter scaling and Kosterlitz-Thouless like critical behavior. To characterize these properties, one should consider different observables, such as large $q>1/2$ and small $q<1/2$ eigenfunction moments, which are associated with distinct critical behaviors, in particular different critical exponents.

cond-mat.dis-nn↗

Transition path sampling in Ising models on heterogeneous graphs

Activated transitions have rates that are often exponentially small in system size. Extracting the associated activation barriers is challenging in practice, especially in the deeply metastable regimes and in the presence of disorder. Here, we use transition path sampling to evaluate transition probabilities between ferromagnetic states in the Ising model on finite sparse random graphs, which are perhaps the simplest example of a disordered system with metastable states. To interpret the transient onset of the transition probability curve, we introduce a minimal three-state kinetic description that highlights the role of intermediate configurations. We validate the method on the heterogeneous Zachary Karate Club network, where distinct dynamical regimes emerge as temperature varies. We then apply the method to random regular graphs and Erdős-Rényi graphs, showing that sample-to-sample fluctuations are weak in the former but that quenched topological disorder induces sizable instance variability in the latter. For Erdős-Rényi graphs, we introduce an instance-dependent temperature rescaling that restores a consistent finite-size scaling of dynamical rates and enables a direct comparison with the corresponding static free-energy barrier.

cond-mat.dis-nn↗