arXiv · 2609.28683
RESTORE: REal-time Steerable Music resTORation and bandwidth Extension via stem disentanglement
Abstract
Neural methods for audio restoration are typically framed as rigid mappings from degraded inputs to single clean outputs, enforcing decisions about what audio content is removed, and potentially adding unwanted content to the restored signal. Because what constitutes a restored audio signal is subjective, we introduce RESTORE, a framework that formulates audio restoration as a six-source semantic decomposition to allow for real-time interactive user control over the process. By expanding a pretrained HTDemucs backbone, a single forward pass disentangles a degraded mixture into vocals, music, broadband hiss, impulsive transients, and an unmodeled residual, while jointly synthesizing a high-frequency extension. Users may steer the restoration by adjusting stem gains, ensuring generative content remains isolated and auditable. RESTORE improves audio quality on diverse historical recordings compared to baselines,lowering Frechet Audio Distance (FAD) (12.13 VGGish; 0.92 CLAP) and delivering aesthetic steerability (Spearman rho greather than 0.91) at 50x real-time on a single GPU. Code and audio samples are available at https://melissachen15.github.io/restore-audio-demo.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Meiying Chen, Benjamin R. Thompson, Michael C. Heilemann. 2026-09-23. RESTORE: REal-time Steerable Music resTORation and bandwidth Extension via stem disentanglement. https://arxiv.org/abs/2609.28683
Cite the original work for its findings. Save a collection to share your selection of sources.