S$^2$COPE: Self-Supervised Concept Discovery via Preference Learning
Self-supervised learning can learn useful representations from unlabeled images. Yet these representations remain largely opaque, making it difficult to understand what visual concepts they encode. Can self-supervision instead produce useful visual concepts that are directly understandable to people? We introduce S$^2$COPE, a framework that represents images as sets of concepts expressed in natural language. A vision-language model proposes candidate concepts, and a self-supervised objective evaluates how well each concept describes the source image while distinguishing it from other images. This feedback is used to construct preferences and train the model to propose better concepts. Across eight diverse visual datasets, our self-supervised concept representations transfer without adaptation and improve downstream recognition accuracy over prior interpretable approaches. In medical imaging, our discovered concepts reveal spurious visual cues and enable a more robust classifier. These results show that self-supervision can discover useful, human-readable visual representations from unlabeled images.