The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge: A Benchmark with Natural Mixtures and Degraded Video
Audio-visual speech enhancement (AVSE) uses a target speaker's visible articulation to recover that speaker's voice from overlapped speech. Yet most evaluation protocols rely on synthetic mixtures and reliable video. The ISCSLP 2026 Real-World AVSE Challenge addresses both gaps. Track~1 combines naturally recorded two-talker mixtures, which lack a clean reference, with reference-available remixes of the same speakers; Track~2 additionally degrades the target video in five ways and adds 3-m far-field recordings. Sixteen and twelve teams were ranked on speaker-disjoint test data by rank averaging over waveform fidelity, predicted quality, transcription accuracy, and speaker similarity. On Track~1 remixes, the best system reaches 12.7~dB SI-SDR and 0.85 STOI, but natural recordings remain harder: even the lowest CER rises from 9.0\% to 14.7\%. The leading systems use video mainly for speaker attribution rather than signal reconstruction, and the top two lose under 0.5~dB SI-SDR on Track~2. UTMOS and DNSMOS rank systems differently from the other metrics, so no single metric captures target-speech recovery. We release the baselines, evaluator, and official results.