arXiv · 2509.22737
CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
Abstract
Visual comparison reasoning is a fundamental capability of vision-language models (VLMs), covering judgments of object quantity, geometric dimensions, spatial relations, and temporal order. Yet existing benchmarks rarely isolate comparison as a reasoning axis, leaving it unclear whether models can reliably perform comparative visual judgments. We introduce a benchmark suite organized around three top-level resources: TallyBench, a 2,000-image object counting benchmark; OmniCaps, a 716-image caption and tag resource; and CompareBench, a 1,200-QA visual comparison benchmark. CompareBench contains four sub-benchmarks spanning quantity, geometric, spatial, and temporal comparison, with the temporal component unifying historical scenes, landmarks, and public figures. Evaluating nine closed-source model routes from Anthropic, Google, and OpenAI on TallyBench and CompareBench reveals strong overall performance but persistent failures in counting, spatial reasoning, geometric comparison, and temporal ordering. These results show that visual comparison remains a systematic weakness of current VLMs and establish CompareBench as a focused benchmark for multimodal reasoning evaluation. All data, code, and prompts will be released at https://github.com/caijie0620/CompareBench.
Explore related subjects
Keep this discovery
Jie Cai. 2026-08-27. CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models. https://arxiv.org/abs/2509.22737
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.