arXiv · 2610.01788
SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples
Abstract
As machine learning increasingly relies on public, untrusted data sources, data poisoning attacks, which inject malicious examples into training data to induce misclassification of a chosen target, pose a growing threat. Existing defenses either assume zero ground-truth information about which examples are poisoned, or they assume access to a large set of examples verified to be clean. Satisfying the latter assumption incurs significant cost since reliable verification can be very resource- or labor-intensive. This cost is particularly high for clean-label attacks, where poisoned examples are visually indistinguishable from clean data. Since requiring a large set of verified examples is impractical, we propose relying on a small set of verified examples including both clean and poisoned ones, i.e., each example verified either to be clean or poisoned through inspection by a forensic expert. The challenge is then to detect poisons based on a set of verified examples that is so small that most classification models would overfit. To address this challenge, we propose Similarity-based Approach for Ground-truth-driven Exclusion (SAGE), which trains a generic feature extractor on a separate dataset and then flags poisoned training examples using a non-parametric, similarity-weighted prediction based on the verified set. On standard benchmarks against seven clean-label attack methods, we demonstrate that having access to even a handful of verified poisoned examples provides a substantial advantage. We also find that the distribution of verified clean examples across classes matters more than the number of verified examples.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chaeeun Han, Soodeh Atefi, Yevgeniy Vorobeychik, Aron Laszka. 2026-10-01. SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples. https://arxiv.org/abs/2610.01788
Cite the original work for its findings. Save a collection to share your selection of sources.