Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
Malicious finetuning attacks pose a major safety threat against open-weight large language models (LLMs). However, existing alignment-stage defenses provide limited protection against strong attacks that use full-parameter finetuning. To address this issue, we propose Patcher, a novel and efficient adversarial training algorithm that alternates between an attack stage and a defense stage: the attack stage simulates many-step malicious finetuning and computes the resulting parameter displacement as an "attack vector", while the defense stage aims to preserve safe behavior under such perturbations. Compared with other adversarial training methods, a key feature of Patcher is that the attack vector can be reused across multiple defense updates, substantially reducing training costs. Theoretically, we characterize how the relative gradient estimation error is influenced by the number of attack steps during training and reuse-duration of attack vectors. Empirically, experiment results show that Patcher reduces Attack Success Rate by 67.5%, 53.4% and 54.8% on three benchmarks compared to the second-best method while preserving the model's utility. Moreover, Patcher remains effective under diverse model sizes and attack scenarios, and generalizes to prompt-based jailbreak attacks. Code is available at https://github.com/haomingwen/patcher