Search arXiv⌕ Search

arXiv · 2609.31785

Cross-Modal Knowledge Distillation for Acoustic Pedestrian Detection

Abstract

Audio-only pedestrian detection is attractive for urban sensing but limited by weak acoustic cues. An appealing strategy is cross-modal knowledge distillation, in which a video teacher supervises the audio student during training so that the deployed model runs on audio alone. Under this task's severe class imbalance and wide video-audio modality gap, however, what such distillation contributes is unclear. We introduce Trust-Filtered Distillation (TFD), which selectively suppresses teacher supervision on pedestrian samples, and interpret its logit formulation as conditional label smoothing under a shared temperature. Across ten distillation configurations under five-fold cross-session validation on ASPED, most methods yield modest gains in macro accuracy accompanied by small changes in PR-AUC. The main effect is higher no-pedestrian accuracy at the cost of lower pedestrian recall. Adding TFD to logit distillation strengthens this trade-off without improving mean PR-AUC. These findings clarify the benefits and limitations of selective cross-modal supervision by distinguishing operating-point shifts from discrimination gains in imbalanced acoustic detection.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yonghyun Kim, Chaeyeon Han, Sancho Gatungay, Subhrajit Guhathakurta, Alexander Lerch. 2026-09-24. Cross-Modal Knowledge Distillation for Acoustic Pedestrian Detection. https://arxiv.org/abs/2609.31785

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios

Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora for training. Synthetic visual data have been shown to be an effective augmentation strategy for addressing AV data scarcity. However, a more challenging scenario arises for languages such as Catalan, where no real audiovisual data are available for training. In this study, we investigate whether AVSR can be bootstrapped in such a zero-AV-resource setting, using synthetic visual data as the sole source of visual supervision. We synthesize over 700 hours of talking-head video and fine-tune a pre-trained AV-HuBERT model. On a manually annotated Catalan benchmark, our model achieves near state-of-the-art (SOTA) performance with much fewer parameters and training data than SOTA ASR systems such as Whisper-large-v3, outperforms an identically trained audio-only baseline, and preserves multimodal advantages under acoustic degradation. Scalable synthetic video thus offers a viable substitute for real recordings in zero-AV-resource AVSR.

eess.AS↗

Teacher-Free Self-Distilled Consistency Trajectory Learning for Fast Speech Enhancement

Consistency trajectory models offer a route to fast, high- quality speech enhancement, collapsing the many reverse steps of diffusion-based enhancers into a handful. When instantiated on a Schrödinger bridge (SB), which pins the generative process to fixed clean and noisy endpoints, exist- ing consistency-trajectory enhancers (SBCTMs) still require a pretrained teacher to supply trajectory supervision, which raises training cost and ties the final quality to that of the teacher. We propose a teacher-free, self-distilled consistency- trajectory framework that removes the external teacher result- ing in a 5X reduction in per epoch training time. Our model is trained with a three-stage curriculum of clean speech pre- diction, a self-distilled shortcut objective, and perceptual fine-tuning with a multi-resolution short-time Fourier trans- form (MR-STFT) loss. Using the same NCSN++ backbone as SBCTM, our model attains a wide-band PESQ of 3.01, ES- TOI 0.87, and SI-SDR 19.07 dB on VoiceBank+DEMAND compared to 3.57, 0.87 and 12.8 dB for the teacher based model. Further, we find that a geometric schedule at low reverse step count maximizes perceptual quality, while a higher-step uniform schedule favors signal fidelity.

eess.AS↗

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1$\times$ with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10$\times$ faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.

eess.AS↗