arXiv · 2410.01243
An Information Theory of Compute-Optimal Size Scaling, Emergence, and Plateaus in Language Models
Abstract
Recent empirical studies show three phenomena with increasing size of language models: compute-optimal size scaling, emergent capabilities, and performance plateauing. We present a simple unified mathematical framework to explain all of these language model scaling phenomena, building on recent skill-text bipartite graph frameworks for semantic learning. Modeling the learning of concepts from texts as an iterative process yields an analogy to iterative decoding of low-density parity check (LDPC) codes in information theory. Thence, drawing on finite-size scaling characterizations of LDPC decoding, we derive the compute-optimal size scaling (Chinchilla rule) for language models. Further, using tools from random network theory, we provide a simple explanation for both emergence of complex skills and plateauing of performance as the size of language models scale. We see multiple plateaus.
Explore related subjects
Keep this discovery
Anuj K. Nayak, Lav R. Varshney. 2024-10-02. An Information Theory of Compute-Optimal Size Scaling, Emergence, and Plateaus in Language Models. https://arxiv.org/abs/2410.01243
Cite the original work for its findings. Save a collection to share your selection of sources.