arXiv · 2609.16458
Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection
Abstract
Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Minu Kim, Ji Sub Um, Hoirin Kim. 2026-09-15. Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection. https://arxiv.org/abs/2609.16458
Cite the original work for its findings. Save a collection to share your selection of sources.