DuplexJail: Spoken Interruption Attacks on Full-Duplex Speech Models
Full-duplex speech models accept user speech while generating responses, making input timing a potential safety concern. We introduce DuplexJail, which delivers fixed, request-independent spoken jailbreak prompts through the user audio channel. Across four open-source models and 720 harmful requests from AdvBench and HarmBench, we compare fixed-delay and refusal-triggered interruption with request-end and post-response controls. On AdvBench, Guided Completion at a 1.0 s delay raises whole-response attack success rates to 40.3% for PersonaPlex and 48.7% for PersonaPlex-RL, increases of 33.8 and 39.3 percentage points over baseline. On HarmBench, which was not used for prompt selection, the same prompt at a 0.5 s delay increases ASR by 14.7 and 12.3 points, respectively. Effects vary across models: selected conditions increase FLM-Audio's harmfulness, while BayLing-Duplex shows decreases. These results show that spoken-jailbreak effectiveness depends on delivery timing and motivate safety evaluation across stages of full-duplex interaction.