Search arXiv⌕ Search

arXiv · 2609.31456

Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis

Abstract

Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hypothesis is that models struggle to integrate multiple components, leading to training interventions to improve compositional binding. However, this assumption has never been directly quantified. Existing benchmarks evaluate captions only in their composed form, making it impossible to separate the cost of joint reasoning from the cost of recognizing individual components under increasing load. We introduce COMPASS (COMPositional Analysis of SkillS), a controlled evaluation framework designed to isolate and measure the distinct factors underlying compositional failure. By comparing performance on composed captions with their decomposed counterparts , we directly quantify the cost of compositional integration across 87K image-caption pairs. Across multiple VLMs, this gap is real but partial, accounting for only part of the observed degradation. This motivates a finer-grained investigation into what additional factors govern model behavior. We analyze performance at the level of individual skills: object detection, attribute binding, and relation reasoning, using skill-targeted perturbations across 274K image-caption pairs. We find a consistent skill-specific pattern: each skill degrades primarily with the count of its own primitive type (self-load), while cross-load effects are predominantly positive, suggesting that primitives of different types provide useful grounding context. This pattern holds across standard contrastive encoders, explicitly trained compositional reasoning models, and non-contrastive architectures. These findings show that compositional degradation reflects multiple separable factors that cannot be reduced to joint reasoning alone.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mona Gandhi, Cenk Merih Olcay, Kuan-Chieh Lo, Santiago Castro, Christopher W. Myers, Srinivasan Parthasarathy. 2026-09-25. Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis. https://arxiv.org/abs/2609.31456

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Long-Tailed 3D Detection via Multi-Modal Fusion

Contemporary autonomous vehicle (AV) benchmarks have significantly advanced multimodal (LiDAR+RGB) 3D detection. However, despite the naturally long-tailed distribution of object classes, existing benchmarking protocols primarily focus on frequent categories (e.g., pedestrian and car), largely overlooking rare but safety-critical classes such as stroller and emergency vehicle. In practice, reliable detection of both common and rare classes is essential for safe autonomous driving. We formalize this problem as Long-Tailed 3D Detection (LT3D), where evaluation encompasses all annotated classes, including rare ones. To address LT3D, we introduce hierarchical losses that promote feature sharing across classes, diagnostic metrics that assign partial credit to semantically reasonable mistakes with respect to the semantic hierarchy (e.g., confusing a child with an adult), and a multimodal late-fusion (MMLF) framework to fuse detections. In particular, we show that rare-class accuracy benefits substantially from MMLF of independently trained uni-modal LiDAR and RGB detectors. Because of the modular design, unlike prevailing end-to-end trained multi-modal detectors that require paired LiDAR-RGB data, MMLF enables the use of advanced unimodal detectors that are trained on large-scale uni-modal datasets with sufficient data for rare classes. Lastly, we examine three fundamental design choices in MMLF, including the RGB detector representation (2D vs. 3D), cross-modal association (3D vs. image plane), and fusion strategy. We find that 2D RGB detectors recognize rare classes more reliably than 3D RGB detectors, image-plane association is more robust to depth estimation errors, and probabilistic score-calibrated fusion consistently yields the best performance. Extensive experiments on nuScenes and Argoverse2 demonstrate substantial improvements of MMLF, establishing a new state of the art.

cs.CV↗

DiDA: Video Object Segmentation with Distillation Learning of Deformable Attention

Video object segmentation is a fundamental research problem in computer vision. Recent techniques have often applied attention mechanism to object representation learning from video sequences. However, due to temporal changes in the video data, attention maps may not well align with the objects of interest across video frames, causing accumulated errors in long-term video processing. In addition, existing techniques have utilised complex architectures, requiring highly computational complexity and hence limiting the ability to integrate video object segmentation into low-powered devices. To address these issues, we propose DiDA, a new method for video object segmentation based on Distillation Learning of Deformable Attention. Specifically, we devise a lightweight architecture for video object segmentation that is effectively adapted to temporal changes. This is enabled by deformable attention mechanism, where the keys and values capturing the memory of a video sequence in the attention module have flexible locations updated across frames. The learnt object representations are thus adaptive to both the spatial and temporal dimensions. We train the proposed architecture using a new knowledge distillation paradigm where deformable attention maps are integrated into the distillation loss. We qualitatively and quantitatively evaluate our method and compare it with existing methods on benchmark datasets including DAVIS 2016/2017 and YouTube-VOS 2018/2019. Experimental results verify the superiority of our method via its achieved state-of-the-art performance on YouTube-VOS18 dataset and optimal memory usage. Project page and code: https://github.com/quangtrungtruong/DiDA.

cs.CV↗

Towards Formal Verification of Deep Neural Networks for Object Detection

Deep neural networks (DNNs) are widely used in real-world computer vision applications, yet they remain vulnerable to errors and adversarial attacks. Formal verification offers a systematic approach to identify and mitigate these vulnerabilities, enhancing model robustness and reliability. While most existing verification methods focus on image classification models, this work extends formal verification to the more complex domain of object detection models. We propose a formulation for verifying the robustness of such models and demonstrate how state-of-the-art verification tools, originally developed for classification, can be adapted for this purpose. Through a comprehensive evaluation, we highlight the ability of formal verification to uncover vulnerabilities in object detection models, and derive formal robustness guarantees, underscoring the potential and need to further extend verification efforts in this domain. This work lays the foundation for further research into formal verification of object detection models across a broader range of computer vision applications. Our source code is publicly available online.

cs.CV↗