arXiv · 2609.37914
The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment
Abstract
Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training data attribution to quantitatively estimate how much each harmful example contributes to EM. We benchmark the quality of the attribution via retraining -- a sound attribution score should enable us to enhance or attenuate EM by filtering data on that score. Score-based filtering can substantially enhance or attenuate EM; we find that both data-attribution scores and a black-box harmfulness score can identify consequential examples. All models we test become misaligned when trained on the same dataset, and influence scores perform best when filtering data from the same model that computed them. We find cross-model generalization of influence scores from scores derived from the three model families we tested, but this generalization does not recover same model filtering performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gonçalo Paulo, Louis Jaburi, Nora Belrose, Lucia Quirke, Stella Biderman. 2026-09-29. The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment. https://arxiv.org/abs/2609.37914
Cite the original work for its findings. Save a collection to share your selection of sources.