Search arXivSearch

arXiv · 2602.05375

Erase at the Core: Representation Unlearning for Machine Unlearning

Abstract

Many approximate machine unlearning methods demonstrate strong logit-level forgetting -- such as near-zero accuracy on the forget set -- yet continue to preserve substantial information within their internal feature representations. We refer to this discrepancy as superficial forgetting. Recent studies indicate that most existing unlearning approaches primarily alter the final classifier, leaving intermediate representations largely unchanged and highly similar to those of the original model. To address this limitation, we introduce the Erase at the Core (EC), a framework designed to enforce forgetting throughout the entire network hierarchy. EC integrates multi-layer contrastive unlearning on the forget set with retain set preservation through deeply supervised learning. Concretely, EC attaches auxiliary modules to intermediate layers and applies both contrastive unlearning and cross-entropy losses at each supervision point, with layer-wise weighted losses. Experimental results show that EC not only achieves effective logit-level forgetting, but also substantially reduces representational similarity to the original model across intermediate layers. Furthermore, EC is model-agnostic and can be incorporated as a plug-in module into existing unlearning methods, improving representation-level forgetting while maintaining performance on the retain set.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jaewon Lee, Yongwoo Kim, Donghyun Kim. 2026-02-27. Erase at the Core: Representation Unlearning for Machine Unlearning. https://arxiv.org/abs/2602.05375

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Online Regularized Statistical Learning in Reproducing Kernel Hilbert Space With Non-Stationary Data

We study recursive regularized learning algorithms in the reproducing kernel Hilbert space (RKHS) with non-stationary online data streams. We introduce the concept of a random Tikhonov regularization path and decompose the tracking error of the algorithm's output for the regularization path into random difference equations in RKHS. We show that the tracking error vanishes in mean square and almost surely if the regularization path is slowly time-varying. Then, leveraging the monotonicity of inverse operators and the spectral decomposition of compact operators, and introducing the RKHS persistence of excitation condition, we develop a dominated convergence method to prove the mean square and almost sure consistency between the regularization path and the unknown function to be learned. Especially, for independent and non-identically distributed data streams, the mean square and almost sure consistency between the algorithm's output and the unknown function is achieved if the input data's marginal probability measures are slowly time-varying and the average measure over each fixed-length time period is uniformly above a strictly positive finite Borel measure.

cs.LG

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models

Direct Preference Optimization (DPO) improves relative preference by increasing the margin between chosen and rejected responses, but this objective does not specify how probability mass should be redistributed to achieve that margin. Consequently, preference separation can improve even as probability mass collapses away from both preferred and unrelated responses, a phenomenon associated with the squeezing effect. We show that this pathology is localized: continued negative updates become destructive when rejected responses enter low-probability valleys, where further suppression induces harmful redistribution through the coupled softmax geometry. We introduce Gradient-Gated Preference Optimization (Gate-DPO), which uses the current probability geometry of each rejected response to selectively attenuate its contribution, limiting excessive rejected suppression while delaying preference-loss saturation. Gate-DPO is objective-agnostic and composes with existing preference methods. Across four architectures, two preference datasets, and multiple objectives, gating consistently reduces squeezing and improves chosen-response likelihood, while exhibiting robust reductions in rejected-response collapse across architectures and objectives. Mass-dynamics analysis further shows that large preference margins can conceal substantial suppression of unrelated responses, while gating achieves healthier redistribution. Together, our results show that preference margin alone is insufficient to characterize optimization health and that selectively controlling probability dynamics provides a simple, general mechanism for stabilizing preference optimization.

cs.LG

Noise-Debiased Thermodynamic Variance for Local Learning Coefficient Probes

Local learning coefficient (LLC) probes offer a singularity-aware view of neural-network training, but mean-energy methods require a local loss baseline that is ambiguous at transient checkpoints. Thermodynamic variance avoids this input; under mini-batch evaluation, however, direct variance mixes cross-state loss fluctuations with same-state noise. We operationalize this route with the \emph{Shift-Invariant Variance Estimator} (SIVE), which estimates and subtracts the latter component using repeated evaluations. Conditional on any fixed retained path, unclipped SIVE is unbiased for noiseless path variance without requiring MCMC stationarity. The finite-scale diagnostic remains indexed by localization scale $h$---even a locally linear loss has tether-dependent variance---while interpretation as a Real Log Canonical Threshold (RLCT) requires additional stationary low-temperature conditions. Toy experiments recover calibrated finite-scale targets. At the primary localization scale, all five MNIST MLP trajectories exhibit a mid-training trough followed by a rebound in SIVE, while Raw Variance decreases from Epoch 40 to 100 in every trajectory. Across four localization scales, the joint early-drop/late-rise criterion is met in 19 of 20 trajectory--scale pairs. At Epoch 40, the estimated observation-noise correction accounts for $77.5\%$ of Raw Variance. Same-state debiasing thus reveals a reproducible turning structure masked by time-varying observation noise.

cs.LG