arXiv · 2605.27311
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
Abstract
Chart question-answering (QA) benchmarks aim to pose questions that require visual reasoning to correctly answer, but vision-language models (VLMs) can often reach solutions through shortcuts or prior familiarity with a chart or question. To strictly evaluate visual reasoning, we propose counterfactual charts where the chart-question task remains fixed, but the underlying data and the corresponding answer are varied. We introduce Chartographer, a framework to reverse engineer charts into executable code, validate reconstruction fidelity, generate counterfactual variants, and derive new answers from executable QA logic. We apply this framework to existing chart QA datasets and evaluate proprietary and open-source VLMs, measuring variant sensitivity and generalizability. Counterfactual charts reveal failures hidden by single-chart performance: VLMs often fail to generalize after answering the original chart correctly. We find that failures are most prevalent when updated charts require novel visual reasoning pathways.
Explore related subjects
Keep this discovery
Yifan Jiang, Dae Yon Hwang, Jesse C. Cresswell, Freda Shi. 2026-05-26. Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models. https://arxiv.org/abs/2605.27311
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.