Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
Fine-tuning on benign data is known to degrade safety alignment in text and vision LLMs, but whether distinct input properties drive this vulnerability differently remains unclear. Audio introduces a richer problem where benign samples can neighbor harmful content through what is said or how it sounds. We present the first systematic study of benign fine-tuning safety in Audio LLMs, evaluating three state-of-the-art models with a proximity-based framework that decomposes embedding-space distance into semantic, acoustic, and mixed axes. We find that the dominant vulnerability axis is architecture-conditioned, determined by how each model's encoder and projector transform audio into the backbone LLM's input space. Across three models, benign fine-tuning elevates Jailbreak Success Rate (JSR) from single digits to as high as 87%, with the most damaging axis shifting from semantic to acoustic proximity depending on encoder design. Mechanistically, fine-tuning selectively suppresses late-layer refusal circuits while frozen encoders preserve upstream representations: the model still detects harmful content but stops refusing, a recognition-refusal dissociation. Two practical defenses, filtering training data to maximize distance from harmful embeddings and a textual system prompt at inference, reduce JSR to near-zero without architectural modification. These findings show that safety evaluation should account for modality and architecture, while highlighting Audio LLMs as a useful testbed for understanding alignment fragility.