Search arXiv⌕ Search

arXiv subjects

Runqiu Xu

Publications and source records attributed to Runqiu Xu.

2 recordsLinked to original sources

Silence-the-Mimic: Accelerating Imperceptible Perturbation Generation Against Voice Cloning

Deep neural network-based Voice Conversion (VC) and Text-to-Speech (TTS) models have rapidly advanced, enabling realistic voice cloning with minimal input data. Such capabilities raise serious concerns over unauthorized cloning of speaker identities and the associated privacy and security risks. Current imperceptible adversarial protection methods rely on quality control losses that are highly sensitive to hyperparameter tuning and computationally expensive due to lengthy optimization. To address these limitations, we propose a fast protection method that generates perceptually constrained perturbations in the frequency domain under a psychoacoustic masking-based constraint. Our approach strictly enforces perceptibility bounds during adversarial training, eliminating the need for iterative quality balancing and significantly reducing computational cost. Experiments on multiple state-of-the-art VC and TTS models show that STM achieves competitive or superior protection performance with substantially better perceptual quality and up to $45.3\times$ speedup over existing white-box baselines. These results demonstrate the effectiveness of frequency-domain perturbations with perceptual constraints as a practical paradigm for protecting against voice cloning.

eess.AS↗

Voices as Handles: Reasoning about Speaker Identity with Frozen Text LLMs

Multi-user voice agents must track who said what across dialogue sessions. Text LLMs are attractive backbones for such agents, but transcripts alone do not expose acoustic speaker identity, leaving the model without a persistent reference for linking information to speakers across sessions. We address this gap by introducing Speaker Handles, soft-token representations that expose acoustic speaker identity to a frozen text LLM for cross-session speaker-dependent reasoning. A three-stage curriculum trains a lightweight projector, with fewer than 0.1% of the backbone's parameters, to map speaker embeddings into these handles. Establishing whether the resulting handles truly support cross-session speaker-dependent reasoning is challenging with existing benchmarks because textual cues can partially reveal fact ownership. We therefore present SpeakerBind, a controlled shared-agent benchmark in which overlapping facts across users require correct cross-session speaker attribution. Speaker Handles achieve 97.40-98.36% accuracy on VoxCeleb1 and 70.40% on SpeakerBind, close to the 71.88% topline. These results show that the proposed Speaker Handles provide an efficient way to integrate acoustic speaker identity into frozen text LLMs for speaker-content reasoning.

eess.AS↗