arXiv · 2609.23453
LiteCASS: A Lightweight End-to-End Network for Real-Time Stereo Cinematic Audio Source Separation
Abstract
Cinematic audio source separation (CASS) decomposes a soundtrack into dialogue, music, and sound-effects (SFX) stems. Existing CASS methods, however, suffer from two critical limitations: they rely on heavily parameterized network architectures and GPU-class hardware, limiting their use in real-time and resource-constrained scenarios, and they are overwhelmingly designed for monaural signals, leaving the stereo scenario largely unexplored. We present LiteCASS, to our knowledge the first lightweight end-to-end network for real-time stereo CASS. LiteCASS combines deterministic STFT subband rearrangement with two jointly trained compact U-Nets: the first extracts dialogue, and the second separates music and SFX from the predicted non-speech component. A multi-task waveform-domain L1 loss supervises all stems. On a spatialized stereo extension of DnR v3, LiteCASS-K8 uses only 1.06M parameters and 0.72G MACs per second, while achieving the highest averaged SI-SDR among the compared CASS baselines.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuanxin Guo, Qiang Ji, Mengmei Liu, Yuhan Lv, Ningning Pan, Gongping Huang. 2026-09-20. LiteCASS: A Lightweight End-to-End Network for Real-Time Stereo Cinematic Audio Source Separation. https://arxiv.org/abs/2609.23453
Cite the original work for its findings. Save a collection to share your selection of sources.