Search arXiv⌕ Search

arXiv subjects

Daniil Zverev

Publications and source records attributed to Daniil Zverev.

3 recordsLinked to original sources

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

The Platonic Representation Hypothesis posits that neural networks trained on different modalities (e.g., text and images) converge toward a shared representation of reality. If true, this has significant implications for whether modality choice matters at all. In this paper, we show that the evidence for this claim is substantially weaker than subsequent work suggests. The mutual $k$-nearest-neighbor metric used on 1024 text-image pairs in the original study captures only coarse structure. To keep the alignment from collapsing as one scales up the data, $k$ has to grow proportionally, undercutting the argument for fine-grained representational convergence. The reported increase in alignment with language model strength saturates for recent models. Moreover, the one-to-one text-image pairing favors alignment, while alignment decreases with non-bijective data. We further find that image and text representations indeed share coarse semantic structure, but neither stronger language models nor richer captions yield fine-grained alignment. Thus, multimodal representations share coarse structure without evidence of convergence to a shared representation -- arguably, full representational convergence would require fine-grained alignment.

cs.CV↗

VGGSounder: Audio-Visual Evaluations for Foundation Models

The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.

cs.MM↗

On the Dangers of Bootstrapping Generation for Continual Learning and Beyond

The use of synthetically generated data for training models is becoming a common practice. While generated data can augment the training data, repeated training on synthetic data raises concerns about distribution drift and degradation of performance due to contamination of the dataset. We investigate the consequences of this bootstrapping process through the lens of continual learning, drawing a connection to Generative Experience Replay (GER) methods. We present a statistical analysis showing that synthetic data introduces significant bias and variance into training objectives, weakening the reliability of maximum likelihood estimation. We provide empirical evidence showing that popular generative models collapse under repeated training with synthetic data. We quantify this degradation and show that state-of-the-art GER methods fail to maintain alignment in the latent space. Our findings raise critical concerns about the use of synthetic data in continual learning.

cs.LG↗