arXiv · 2610.03602
Evaluating Inference-time Algorithms for Semantic Sound Scene Segmentation
Abstract
Spatial semantic segmentation of sound scenes (S5) involves the detection and extraction of target sound events in audio files. However, in typical detection-extraction pipelines, errors in the detection module propagate to the extraction module, degrading overall performance. In this work, we develop a multichannel detection-extraction model and evaluate inference-time algorithms to improve the class detection and target sound extraction in a multi-stage setup. We evaluate two fitness functions: binary cross entropy and mixture consistency; two relabeling strategies: naive and conditional confusion matrix relabeling; and three source estimate evaluation models: a multi-label classifier, a fine-tuned single-source classifier, and an off-the-shelf audio judge. We thoroughly evaluate the impact of these design choices on the DCASE 2025 Task 4 dataset, and validate entropy-based gating of inference-time updates on the held-out evaluation set, laying the groundwork for inference-time updates of S5 systems in real-world scenarios.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sripathi Sridhar, Gordon Wichern, Yoshiki Masuyama, Mark Cartwright, Jonathan Le Roux. 2026-10-02. Evaluating Inference-time Algorithms for Semantic Sound Scene Segmentation. https://arxiv.org/abs/2610.03602
Cite the original work for its findings. Save a collection to share your selection of sources.