arXiv · 2609.32188
Presence Is Not Faithfulness: Figurative Vehicle Intrusion in Text-to-Image Generation
Abstract
Text-to-image (TTI) models increasingly generate high-quality images from natural-language prompts, yet figurative language exposes a failure: a vehicle that should guide the depiction of a tenor may instead be rendered as a visible object. We call this failure Figurative Vehicle Intrusion: the intruding content is textually licensed, but it is assigned the wrong visual role, showing that visual presence is not always faithfulness and that presence-oriented evaluation can miss such errors. To study it systematically, we introduce Vehicle Intrusion and Semantic Tenor Assessment (VISTA), a multilingual benchmark of figurative prompts organized by Figurative Form and Mapping Mechanism. We further propose V-Score, a diagnostic question-answering metric that evaluates role-aware figurative faithfulness in generated images. Evaluations on recent high-performing TTI models show that vehicle intrusion persists across languages and figurative categories. As a lightweight mitigation, we introduce VISTA-Guard, which partially reduces vehicle intrusion and suggests a practical path toward more figuratively faithful TTI generation. All resources will be released publicly.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiaoyu Ma, Chen Yang, Hao Chen. 2026-09-26. Presence Is Not Faithfulness: Figurative Vehicle Intrusion in Text-to-Image Generation. https://arxiv.org/abs/2609.32188
Cite the original work for its findings. Save a collection to share your selection of sources.