arXiv · 2610.02823
What Actually Makes Correlation-Based SSL Distillation Noise-Robust? A Mechanistic Correction
Abstract
Self-supervised learning (SSL) speech models are accurate but large. Knowledge distillation compresses them, but the student loses the teacher's noise robustness. Correlation-based distillation addresses this with two terms: cross-correlation aligning student and teacher's representations, and self-correlation decorrelating the student's features. Both the original method and De'HuBERT credited the self-correlation term without isolating it. We show the opposite. On LibriSpeech-100 distillation with held-out CHiME-3 noise at 10\,dB, an unbiased full-dimensional probe, a Pearson-variance decomposition, a same-noise causal control, and a per-dimension analysis identify the cross-correlation diagonal as the mechanism that encourages noise invariance, lowering noise-classification accuracy from 76.98\% to 55.02\%, whereas adding the self-correlation term back leaves it at 55.55\%. The self-correlation term instead reorganises the feature space and improves accuracy on the clean downstream tasks, but removes essentially no noise. Across nine speech and music tasks this yields a concrete design rule for distillation: weight the cross-correlation term for noise robustness, tune the self-correlation term for better downstream performance, and sample teacher and student noise independently.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Fabian Ritter-Gutierrez, Nancy F. Chen, Eng Siong Chng. 2026-10-02. What Actually Makes Correlation-Based SSL Distillation Noise-Robust? A Mechanistic Correction. https://arxiv.org/abs/2610.02823
Cite the original work for its findings. Save a collection to share your selection of sources.