arXiv · 2609.37287
VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics
Abstract
Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to comprehensively assess this capability. To systematically evaluate this capability, we introduce VISTA-Bench, covering 22 languages and 10 domains, and develop an image-specific rubric evaluation protocol. The benchmark combines sampling for language and scenario coverage with model-assisted, human-verified annotations that group related text into coherent semantic units and provide multilingual reference translations. The rubrics specify essential content, semantic relations, and acceptable translation variants, yielding separate output-based scores for translation quality and the preservation of visual and knowledge-dependent information. We conduct extensive evaluations of 16 mainstream models, including 12 multimodal models and four text-input models, and provide systematic analyses across languages, domains, and evaluation dimensions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bo Lv, Mao Zheng, Zheng Li, Fangxu Liu, Mingrui Sun, Tao Chen. 2026-09-29. VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics. https://arxiv.org/abs/2609.37287
Cite the original work for its findings. Save a collection to share your selection of sources.