arXiv · 2610.08577
How Learning Governs Unlearning across the Memorization-Generalization Spectrum
Abstract
While unlearning seeks to negate undesired capabilities acquired through learning, little research has examined how the way models learn shapes their subsequent unlearning. In this paper, we investigate this connection from the perspectives of memorization and generalization, the two most representative yet competing strategies that models employ during training. We first classify memorization- and generalization-heavy models using grokking in modular addition and compare their responses to unlearning, showing that the latter suffer greater retain damage, i.e., a larger performance drop on the retain set. Furthermore, we conduct a finer-grained analysis by introducing bucketed modular addition, in which the respective contributions of the two strategies can be explicitly controlled across the memorization-generalization spectrum. In this setup, we reaffirm that the same trend persists and is nearly monotonic. We further demonstrate that this relationship also holds in LLM unlearning across verbatim and factual recall settings. Finally, we provide two practical insights for developing better unlearning methods, highlighting the importance of accounting for learning dynamics in unlearning.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hwiyeong Lee, Hyelim Lim, Ingyu Bang, Hoki Kim, Taeuk Kim. 2026-10-06. How Learning Governs Unlearning across the Memorization-Generalization Spectrum. https://arxiv.org/abs/2610.08577
Cite the original work for its findings. Save a collection to share your selection of sources.