Search arXiv⌕ Search

arXiv · 2610.05310

Transferable Adversarial Robustness for Speech Foundation Models via Hierarchical Stabilization

Abstract

Frozen speech foundation models (SFMs) make downstream adaptation efficient: the backbone can stay fixed while a task learns layer fusion and a lightweight classifier. Full adversarial fine-tuning is a standard route to robustness, but generating adversarial examples and updating the backbone for every task sacrifices that efficiency. We ask whether robustness can instead be learned before future tasks are known. For a frozen backbone and linear classifier, robustness can be understood through the interaction between representation stability and decision-boundary margin. This leads directly to our design: we stabilize representations across the hidden layers, rather than only the final layer, while preserving clean representations; after clean adaptation selects the layer mixture, we keep it fixed and enlarge only the classifier margin, without downstream adversarial examples. We evaluate Wav2Vec2, HuBERT, and WavLM Large on four tasks under adaptive 30 dB attacks. Across 12 backbone-task pairs, hierarchical robustification improves robust accuracy by 46.4 pp, while margin refinement adds 4.0 pp for 1.1 pp of clean accuracy. Code and configurations are available at https://github.com/arefmousavi/hierarchical-robust-sfm.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aref Mousavi, Shahab Sherafat, Kiarash Kiani Feriz, Amirparsa Safari, Raoof Zare Moayedi, Mohammad Hossein Rohban, Mohammad Sabokrou. 2026-10-04. Transferable Adversarial Robustness for Speech Foundation Models via Hierarchical Stabilization. https://arxiv.org/abs/2610.05310

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Personal VAD: Speaker-Conditioned Voice Activity Detection

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only triggers for the target user, which helps reduce the computational cost and battery consumption, especially in scenarios where a keyword detector is unpreferable. We achieve this by training a VAD-alike neural network that is conditioned on the target speaker embedding or the speaker verification score. For each frame, personal VAD outputs the probabilities for three classes: non-speech, target speaker speech, and non-target speaker speech. Under our optimal setup, we are able to train a model with only 130K parameters that outperforms a baseline system where individually trained standard VAD and speaker recognition networks are combined to perform the same task.

eess.AS↗

Textual Echo Cancellation

In this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo from overlapping speech recordings. Such a system can largely improve speech recognition performance and user experience for intelligent devices such as smart speakers, as the user can talk to the device while the device is still playing the TTS signal responding to the previous query. We implement this system by using a novel sequence-to-sequence model with multi-source attention that takes both the microphone mixture signal and source text of the TTS playback as inputs, and predicts the enhanced audio. Experiments show that the textual information of the TTS playback is critical to enhancement performance. Besides, the text sequence is much smaller in size compared with the raw acoustic signal of the TTS playback, and can be immediately transmitted to the device or ASR server even before the playback is synthesized. Therefore, our proposed approach effectively reduces Internet communication and latency compared with alternative approaches such as acoustic echo cancellation (AEC).

eess.AS↗

Multiplexing Neural Audio Watermarks with Adaptive Routing

Audio watermarking supports speech authenticity verification. We study whether combining heterogeneous watermarks can retain at least one detectable provenance signal when their failure modes differ. This Any-survival objective concerns complementary evidence retention, not simultaneous survival of all constituent marks or recovery of all payloads. To our knowledge, this is the first systematic study of neural audio watermark multiplexing, with a scoped benchmark covering five released systems, parallel and sequential baselines, and 14 evaluation conditions. The benchmark shows complementary failure modes, but naive composition does not reliably turn them into system-level robustness under the Any-survival objective. We therefore formulate multiplexing as watermark allocation and study perceptual-adaptive time-frequency multiplexing (PA-TFM), a training-free routing method, and MaskNet, a learned time-domain router for separately trained systems with native detectors. We use the five-system scoped benchmark to characterize multiplexing behavior and the AudioSeal-PerTh pair as a representative heterogeneous pair for adaptive watermark routing. Compared with direct parallel composition, MaskNet improves average TPR@1%FPR from 0.75 to 0.88 and SNR from 15.20 dB to 25.36 dB; compared with PA-TFM, it gives a clear system-level robustness gain under the Any-survival objective while retaining similar high-fidelity behavior. A SpeechTokenizer-aware case study further improves average TPR@1%FPR to 0.91 and SpeechTokenizer robustness from 0.20 to 0.60, showing that channel-adapted constituents can enter the same routing framework.

eess.AS↗