arXiv · 2610.00620
Misalignment of Low-Loss Regions Causes Grokking
Abstract
Grokking refers to the delayed emergence of validation-set generalization after a model has already overfit the training set. Although first observed in small algorithmic tasks trained with transformers, its underlying mechanism remains unsettled. In this work, we develop an analysis framework based on mode connectivity and the geometry of low-loss regions. The framework predicts that the standard modular-arithmetic setting does not always produce grokking: under a symmetry-preserving train/validation split, we observe a stable anti-grokking case in which validation performance does not recover. This counterexample challenges several existing correlational explanations of grokking. More broadly, our analysis framework and results further suggest that grokking arises when the low-loss regions induced by the training and validation partitions are misaligned. Once these regions become well aligned, training hyperparameters alone cannot produce grokking and the observed dynamics collapse to either trainable or non-trainable behavior.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yongding Tian, Zaid Al-Ars, Maksim Kitsak, Peter Hofstee. 2026-09-30. Misalignment of Low-Loss Regions Causes Grokking. https://arxiv.org/abs/2610.00620
Cite the original work for its findings. Save a collection to share your selection of sources.