PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides
Coding agents are increasingly moving beyond text-based software tasks to reconstruct visual targets through code. This capability, visual coding, requires agents to translate their understanding of visual targets into executable code. Slides provide a natural testbed for this capability, combining rich visual structure with objects that can be programmatically created and edited. To measure this capability, we introduce PPTBench, a benchmark for reconstructing scientific flow diagrams as editable PowerPoint slides. PPTBench covers 500 scientific flow-diagram tasks across 10 presentation domains, drawn from real research papers, and is evaluated with a four-stage agentic judge covering artifact validity, process and connector fidelity, rendering quality, and visual fidelity. Evaluation of ten models across 36 model--harness--effort configurations shows a substantial gap in reliable visual coding: the best configuration achieves 77.34, while the median across configurations is 24.38. Fine-grained analysis shows that agents can generally produce valid slide files, but struggle to produce high-quality reconstructions that faithfully recover the semantics and visual structure of the target. Further analysis shows that increasing reasoning effort primarily improves hard-gate passage rather than mean detail quality on each configuration's passing tasks, while configurations with more inspection tend to achieve higher overall scores. PPTBench establishes a measurable testbed for studying visual coding and advancing agents toward more reliable visual creation.