arXiv · 2609.27395
PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models
Abstract
Compact vision-language models (VLMs) now power a growing share of multimodal applications. The benchmarks used to compare them, however, inherit a frontier-centric design: each model is reduced to a single accuracy number, narrowing the inter-model gap on saturated suites and pressing models into low-score bands on harder ones. We introduce PRISM-VLM, a multi-axis discriminative benchmark that scores every item along seven axes covering the recurring failure modes (task quality, behavioral robustness, and capability bottlenecks) and combines them into a single PScore, with items recycled from fifteen public benchmarks. Across compact VLMs from the past two years, PScore separates model pairs more reliably than prior single-axis benchmarks under an item-level paired bootstrap, and surfaces behavioral differences these benchmarks average away. Even models with statistically indistinguishable PScores diverge sharply along the per-axis profile, particularly on sycophancy, which is nearly orthogonal to single-prompt accuracy. We will release the full pipeline, prompts, and per-item annotations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sanghee Park, Kee-Eung Kim. 2026-09-23. PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models. https://arxiv.org/abs/2609.27395
Cite the original work for its findings. Save a collection to share your selection of sources.