Search arXivSearch

arXiv · 2602.09061

Coherent information deletion: Bayes' theorem and generalized Bayesian unlearning

Abstract

Bayes' theorem admits an information-processing interpretation due to Zellner (1988): under the Shannon-information criterion, the posterior is the unique rule that processes prior and data information without information loss. We revisit these ideas, but from the perspective of information deletion. Given a posterior based on a complete dataset, what distribution should replace it when a subset of the data is removed? We define information deletion using the same information conservation principle as Zellner (1988), and show that the optimalpost-deletion distribution is exactly the leave-data-out posterior. We then extend the framework beyond likelihood-based inference from Bayes to the generalized Bayesian updating of Bissiri et al. (2016) based on loss functions. We introduce a sequential coherence requirement for deletion, under which, removing two pieces of information jointly is equivalent to removing them successively. The resulting coherent deletion rule exactly recovers the generalized Bayesian posterior based only on the retained data. Restricting these optimization problems to variational families yields corresponding formulations of variational Bayesian and generalized Bayesian unlearning.

Explore related subjects

Keep this discovery

BibTeXRIS

Hans Montcho, Håvard Rue. 2026-08-27. Coherent information deletion: Bayes' theorem and generalized Bayesian unlearning. https://arxiv.org/abs/2602.09061

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

A simple derivation of the Kalman filter

In this lecture note, we present a concise and self-contained derivation of the discrete-time Kalman filter equations that requires only a basic understanding of least squares estimation. The treatment is designed to minimize mathematical overhead while preserving both rigor and generality.

math.OC

Probabilistic Symbolic Regression for Equation Discovery via Operator-induced and Regularized Symbolic Forests

Symbolic regression has emerged as a powerful tool for artificial intelligence-driven scientific discovery by learning interpretable analytical expressions that reveal governing relationships directly from data. Existing methods, however, often rely on heuristic search, struggle to balance predictive accuracy with expression complexity in noisy settings, and offer limited characterization of symbolic uncertainty. Probabilistic approaches that address these challenges in a unified manner remain underexplored. We introduce a probabilistic symbolic regression framework that represents mathematical expressions as ensembles of symbolic trees. A regularizing prior over tree topology controls expression complexity, while an Occam's window-based posterior summary captures uncertainty across multiple plausible symbolic models. Given the limited existing theoretical treatment of symbolic regression, we develop posterior concentration guarantees when symbolic expressions approximate the underlying relationship arbitrarily well, with a near-parametric rate when an exact finite formula exists. Additionally, we establish a sharp oracle concentration result under symbolic misspecification. Comparisons of our proposed framework with state-of-the-art competitors demonstrate superior predictive accuracy, optimal symbolic complexity, and stable structural recovery when learning benchmark scientific equations, together with the identification of scientifically interpretable descriptor formulas in a challenging materials discovery application.

stat.ME

Estimating systematic errors in Bayesian inversion using transport maps

In indirect measurements, the sought parameters have to be determined by solving an inverse problem, typically in a Bayesian framework. Often, the accurate numerical simulation of the measuring process is computationally demanding, making it necessary to rely on approximate models. These surrogates, however, introduce an additional model error and thus may distort the resulting parameter distribution. Moreover, even with the additional speed granted by the surrogate, posterior determination through conventional means such as Markov chain Monte Carlo might be cost intensive, specifically for complicated posterior shapes. In this paper, we propose a unified framework that combines Bayesian inference, model error correction and a transport-based sampling scheme to address these issues. To train the transport scheme, we investigate two different losses: one equivalent to the Kullback-Leibler divergence associated to the transport problem and one based on an upper bound of this loss, generally known as the evidence lower bound. We demonstrate that training the transport based on the latter changes the optimisation landscape drastically, potentially introducing an undesired bias in approximating the target posterior. We compare the computational cost of our approach with established methods and underline the theoretical results with numerical examples.

stat.ME