arXiv · 2610.01180
Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models
Abstract
Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently struggle to understand negation and produce incorrect answers when questions involve negated clauses. To address this limitation, we propose Skeleton-and-Strategy Prompting (\textbf{SSP}), a training-free, in-context learning method that improves VLM negation understanding capabilities without any parameter updates. Given a negation question, our method first abstracts the underlying question structure into a skeleton, retrieves a small set of same-skeleton questions from a lightweight question pool, then prompts the VLM to analyze their shared negation pattern and synthesize a single-sentence answering strategy. The skeleton and strategy are prepended to the test sample to guide the model correctly tackle the negation problems. Experiments on multiple negation VQA benchmarks show that SSP achieves state-of-the-art performance on negation-focused VQA tasks while remaining computationally efficient.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuliang Cai, Mohammad Rostami, Jesse Thomason. 2026-10-01. Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models. https://arxiv.org/abs/2610.01180
Cite the original work for its findings. Save a collection to share your selection of sources.