Evaluating Inference-time Algorithms for Semantic Sound Scene Segmentation
Spatial semantic segmentation of sound scenes (S5) involves the detection and extraction of target sound events in audio files. However, in typical detection-extraction pipelines, errors in the detection module propagate to the extraction module, degrading overall performance. In this work, we develop a multichannel detection-extraction model and evaluate inference-time algorithms to improve the class detection and target sound extraction in a multi-stage setup. We evaluate two fitness functions: binary cross entropy and mixture consistency; two relabeling strategies: naive and conditional confusion matrix relabeling; and three source estimate evaluation models: a multi-label classifier, a fine-tuned single-source classifier, and an off-the-shelf audio judge. We thoroughly evaluate the impact of these design choices on the DCASE 2025 Task 4 dataset, and validate entropy-based gating of inference-time updates on the held-out evaluation set, laying the groundwork for inference-time updates of S5 systems in real-world scenarios.