Search arXiv⌕ Search

arXiv · 2610.05737

Revisiting Frame-Wise Saliency for Audio Moment Retrieval

Abstract

This paper revisits frame-wise saliency for audio moment retrieval (AMR). We show that the frame-wise saliency sequence, conventionally used only as an auxiliary output in DETR-based AMR models, can itself serve as an effective source of moment predictions. We convert the saliency sequence into ranked moments using a simple SED-inspired segmentation rule with no learned parameters, enabling moment retrieval directly from frame-wise temporal information. On the CASTELLA dataset, saliency-based prediction consistently outperforms decoder-based prediction from the same model across all 18 runs of QD-DETR and CG-DETR. For QD-DETR, simply replacing the inference output improves R1@0.7 from 21.0 to 36.1. The advantage remains 7-15 points when the two outputs are evaluated at their independently selected best epochs. The same tendency extends to TaskWeave and UVCOM, whereas TR-DETR shows the opposite behavior, suggesting that how saliency construction may matter. The performance gap is especially pronounced for short moments: for queries whose annotated moments average at most 2 s, R1@0.7 improves from 6.3 to 27.8 with QD-DETR. Decoder supervision nevertheless benefits saliency-based prediction, indicating that its role during training differs from the utility of its inference output.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tatsuya Munakata, Hokuto Munakata. 2026-10-05. Revisiting Frame-Wise Saliency for Audio Moment Retrieval. https://arxiv.org/abs/2610.05737

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Personal VAD: Speaker-Conditioned Voice Activity Detection

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only triggers for the target user, which helps reduce the computational cost and battery consumption, especially in scenarios where a keyword detector is unpreferable. We achieve this by training a VAD-alike neural network that is conditioned on the target speaker embedding or the speaker verification score. For each frame, personal VAD outputs the probabilities for three classes: non-speech, target speaker speech, and non-target speaker speech. Under our optimal setup, we are able to train a model with only 130K parameters that outperforms a baseline system where individually trained standard VAD and speaker recognition networks are combined to perform the same task.

eess.AS↗

Textual Echo Cancellation

In this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo from overlapping speech recordings. Such a system can largely improve speech recognition performance and user experience for intelligent devices such as smart speakers, as the user can talk to the device while the device is still playing the TTS signal responding to the previous query. We implement this system by using a novel sequence-to-sequence model with multi-source attention that takes both the microphone mixture signal and source text of the TTS playback as inputs, and predicts the enhanced audio. Experiments show that the textual information of the TTS playback is critical to enhancement performance. Besides, the text sequence is much smaller in size compared with the raw acoustic signal of the TTS playback, and can be immediately transmitted to the device or ASR server even before the playback is synthesized. Therefore, our proposed approach effectively reduces Internet communication and latency compared with alternative approaches such as acoustic echo cancellation (AEC).

eess.AS↗

Multiplexing Neural Audio Watermarks with Adaptive Routing

Audio watermarking supports speech authenticity verification. We study whether combining heterogeneous watermarks can retain at least one detectable provenance signal when their failure modes differ. This Any-survival objective concerns complementary evidence retention, not simultaneous survival of all constituent marks or recovery of all payloads. To our knowledge, this is the first systematic study of neural audio watermark multiplexing, with a scoped benchmark covering five released systems, parallel and sequential baselines, and 14 evaluation conditions. The benchmark shows complementary failure modes, but naive composition does not reliably turn them into system-level robustness under the Any-survival objective. We therefore formulate multiplexing as watermark allocation and study perceptual-adaptive time-frequency multiplexing (PA-TFM), a training-free routing method, and MaskNet, a learned time-domain router for separately trained systems with native detectors. We use the five-system scoped benchmark to characterize multiplexing behavior and the AudioSeal-PerTh pair as a representative heterogeneous pair for adaptive watermark routing. Compared with direct parallel composition, MaskNet improves average TPR@1%FPR from 0.75 to 0.88 and SNR from 15.20 dB to 25.36 dB; compared with PA-TFM, it gives a clear system-level robustness gain under the Any-survival objective while retaining similar high-fidelity behavior. A SpeechTokenizer-aware case study further improves average TPR@1%FPR to 0.91 and SpeechTokenizer robustness from 0.20 to 0.60, showing that channel-adapted constituents can enter the same routing framework.

eess.AS↗