Search arXivSearch

arXiv · 2109.00454

The emergence of a concept in shallow neural networks

Abstract

We consider restricted Boltzmann machine (RBMs) trained over an unstructured dataset made of blurred copies of definite but unavailable ``archetypes'' and we show that there exists a critical sample size beyond which the RBM can learn archetypes, namely the machine can successfully play as a generative model or as a classifier, according to the operational routine. In general, assessing a critical sample size (possibly in relation to the quality of the dataset) is still an open problem in machine learning. Here, restricting to the random theory, where shallow networks suffice and the grand-mother cell scenario is correct, we leverage the formal equivalence between RBMs and Hopfield networks, to obtain a phase diagram for both the neural architectures which highlights regions, in the space of the control parameters (i.e., number of archetypes, number of neurons, size and quality of the training set), where learning can be accomplished. Our investigations are led by analytical methods based on the statistical-mechanics of disordered systems and results are further corroborated by extensive Monte Carlo simulations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Elena Agliari, Francesco Alemanno, Adriano Barra, Giordano De Marzo. 2021-09-01. The emergence of a concept in shallow neural networks. https://arxiv.org/abs/2109.00454

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

R-transforms for non-Hermitian matrices: a spherical integral approach

In this paper, we establish a connection between the formalism of $\mathcal{R}$-transforms for non-Hermitian random matrices and the framework of spherical integrals, using the replica method. This connection was previously proved in the Hermitian setting and in the case of bi-invariant random matrices. We show that the $\mathcal{R}$-transforms used in the non-Hermitian context in fact originate from a single scalar function of two variables. This provides a new and transparent way to compute $\mathcal{R}$-transforms, which until now had been known only in restricted cases such as bi-invariant, Hermitian, or elliptic ensembles.

cond-mat.dis-nn

Spectral boundaries of deterministic matrices deformed by rotationally invariant random non-Hermitian ensembles

One of the great miracles of random matrix theory is that, in the $N \to \infty$ limit, many otherwise intractable matrix problems with horrendously complicated finite-$N$ expressions admit remarkably simple and elegant asymptotic solutions. In this paper, we illustrate this phenomenon in the context of spectral boundaries (or spectral edges) for deformed random matrices. Specifically, we consider matrices of the form $\mathbf{A} + \mathbf{B}$, where $\mathbf{A}$ is a deterministic $N\times N$ matrix (not necessarily Hermitian) and $\mathbf{B}$ is a rotationally invariant random matrix. In the large-$N$ limit, we show that the complex eigenvalue distribution of $\mathbf{A} + \mathbf{B}$ satisfies remarkably simple boundary equations that depend on the $\mathcal{R}_1$ and $\mathcal{R}_2$ transforms of $\mathbf{B}$. We illustrate our results on several explicit random matrix ensembles and support them with numerical simulations.

cond-mat.dis-nn

Electrical conductivity of crack-template-based transparent conducting films: mean-field approximation, effective-medium theory, and simulation

In this work, crack-template-based transparent conducting films were modeled as networks corresponding to the edges of a two-dimensional Poisson--Voronoi diagram. Two types of networks were considered: the original one, in which the conductance of each edge was inversely proportional to its length, and the effective one, in which all edges had the same conductance obtained from the effective-medium theory. The mean-field approximation was used for analytical evaluation of the electrical conductivity. Direct numerical calculations for the Poisson--Voronoi diagram showed that the mean-field approximation overestimated the effective conductivity of the original network by approximately 13\%, and of the effective network by 79\%. In addition, a honeycomb network with an edge conductance distribution corresponding to the Poisson--Voronoi diagram was studied: for it, the predictions of the effective-medium theory turned out to be more accurate than for the Poisson--Voronoi diagram, which was explained by the greater structural homogeneity of the periodic honeycomb lattice. The results indicate that, when modeling crack-template-based transparent conducting films, the application of the mean-field approximation may lead to significant errors if the resistance of individual conductors is not simply proportional to their length. This possibility is discussed as a motivation for future studies of hierarchical cracks with variable width, which are not directly investigated here.

cond-mat.dis-nn