arXiv · 2401.11004
Predictive power of a Bayesian effective action for fully-connected one hidden layer neural networks in the proportional limit
Abstract
We perform accurate numerical experiments with fully-connected (FC) one-hidden layer neural networks trained with a discretized Langevin dynamics on the MNIST and CIFAR10 datasets. Our goal is to empirically determine the regimes of validity of a recently-derived Bayesian effective action for shallow architectures in the proportional limit. We explore the predictive power of the theory as a function of the parameters (the temperature $T$, the magnitude of the Gaussian priors $\lambda_1$, $\lambda_0$, the size of the hidden layer $N_1$ and the size of the training set $P$) by comparing the experimental and predicted generalization error. The very good agreement between the effective theory and the experiments represents an indication that global rescaling of the infinite-width kernel is a main physical mechanism for kernel renormalization in FC Bayesian standard-scaled shallow networks.
Explore related subjects
Keep this discovery
P. Baglioni, R. Pacelli, R. Aiudi, F. Di Renzo, A. Vezzani, R. Burioni, P. Rotondo. 2024-01-19. Predictive power of a Bayesian effective action for fully-connected one hidden layer neural networks in the proportional limit. https://arxiv.org/abs/2401.11004
Cite the original work for its findings. Save a collection to share your selection of sources.