arXiv · 2610.10121
DuRe-ST: Dual-Relation Spectro-Temporal Modeling for Speech Deepfake Detection
Abstract
Previous speech deepfake detectors can adaptively capture spectro-temporal dependencies through graph attention, yet they largely overlook the co-variation between spectral and temporal representations. To address this gap, we construct a normalized affinity graph from their joint covariance and apply polynomial graph filtering to capture higher-order covariance-induced dependencies. We first develop Cov-ST to isolate the contribution of covariance-based relational modeling. Although it improves detection performance, its sensitivity to the polynomial order suggests limited robustness when covariance relations are modeled alone. We therefore propose DuRe-ST, which jointly exploits covariance-induced and graph-attention-induced relations to capture complementary second-order co-variation and adaptive spectro-temporal dependencies. Experiments show that DuRe-ST achieves an average relative EER reduction of 25.9% over XLSR-AASIST on the ASVspoof benchmarks and 28.4% across four cross-dataset benchmarks with only 4-8k additional trainable back-end parameters. It further outperforms the strongest publicly available comparison models by 2.2-13.2% in relative EER on four benchmarks, while remaining smaller than the publicly available models considered.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shaole Li, Siqing Qin, Youzhi Tu, Kong Aik Lee. 2026-10-07. DuRe-ST: Dual-Relation Spectro-Temporal Modeling for Speech Deepfake Detection. https://arxiv.org/abs/2610.10121
Cite the original work for its findings. Save a collection to share your selection of sources.