Training-Free Uncertainty Estimation for Embedding Models
Embedding models, often obtained via self-supervised learning, extract general-purpose representations from data. Quantifying the reliability of these representations is crucial, as many downstream models rely on them as input for their own tasks. To this end, we introduce a formal definition of representation reliability: the representation for a given test point is considered to be reliable if the downstream models built on top of that representation can, on average, consistently generate accurate predictions for that test point across various downstream tasks. However, accessing the downstream data to quantify the representation reliability is often limited or restricted for various reasons. We propose training-free methods for estimating the representation reliability without access to the downstream data. Our method is based on the concept of neighborhood consistency (NC) across distinct pre-trained representation spaces. The key insight is to find shared neighboring points as anchors to align these representation spaces before comparing them. We provide theoretical justifications for NC and develop two practical approaches: (1) directly computing NC when multiple pre-trained models are available, and (2) a perturbation-based NC (PNC), which creates synthetic ensembles from a single model through isotropic Gaussian noise, avoiding the computational cost of training deep ensembles. We further propose PNC-spread tuning, which systematically determines the perturbation magnitude by maximizing the spread of the PNC scores on a reference set. We demonstrate through comprehensive numerical experiments that our methods effectively capture the representation reliability with a high degree of correlation, achieving robust and favorable performance compared with baseline methods.