arXiv · 2609.38780
RAST: Resolution-Aware Privileged Structure Transfer for Low-Resolution Audio Activity Recognition
Abstract
Audio is increasingly used for human activity recognition (HAR) because it captures object interactions, environmental events, and contextual cues in everyday environments. High-resolution (HR) audio provides rich acoustic information for model development but incurs substantial energy and storage costs and may expose sensitive speech content. Low-resolution (LR) audio offers a more privacy-preserving and resource-efficient alternative for deployment, but reduced sampling rates can remove acoustic cues essential for activity recognition, leading to significant performance degradation. We formulate this training-deployment mismatch as sensor-resolution privileged learning, in which HR audio is available during training, while inference relies exclusively on LR audio. We propose RAST, a resolution-aware transfer framework that compresses HR teacher representations by preserving token-level information and neighborhood structure before performing localized HR-LR alignment. Experiments on the SAMoSA and AudioIMU datasets show that RAST consistently outperforms LR-only training and direct teacher-transfer baselines, improving LR-only recognition by up to approximately 7.8% while requiring only LR audio at inference.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ji Hwan Park, Gautham Krishna Gudur, Yufei Shen, Dawei Liang, Edison Thomaz. 2026-09-30. RAST: Resolution-Aware Privileged Structure Transfer for Low-Resolution Audio Activity Recognition. https://arxiv.org/abs/2609.38780
Cite the original work for its findings. Save a collection to share your selection of sources.