Deep Filter Estimation from Inter-Frame Correlations for Monaural Speech Dereverberation
Speech dereverberation with a distant microphone is challenging because reverberation is correlated with the target speech, and models trained on simulated data often generalize poorly to real recordings. We propose IF-CorrNet, a correlation-to-filter architecture for monaural dereverberation. Instead of feeding raw complex STFT coefficients to the network, IF-CorrNet computes inter-frame correlations among neighboring frames at each time-frequency bin and estimates multi-frame deep filters from these features with a dual-path Transformer backbone. This design makes inter-frame dependencies explicit at the network input while retaining a multi-frame filtering output, a pairing motivated by the normal equation of linear multi-frame filtering. On the REVERB Challenge corpus, IF-CorrNet achieves the best CD, LLR, SNRfw, and PESQ among the compared dereverberation baselines on SimData, and the highest SRMR among the compared systems on RealData. The ablation shows higher RealData SRMR with correlation inputs for both filtering and masking, with filtering adding a further gain.