GUI-GenBench: Evaluating Image Generation Models as Interactive GUI Environments
An image generation model used as a graphical user interface (GUI) environment must produce not only a plausible screen but also the correct response to an action. Existing visual-generation benchmarks do not directly assess this combination of functional correctness and temporal consistency in GUIs. We introduce GUI-GenBench, a benchmark of 700 curated samples across five task suites: single-step transitions, multi-step planning, fictional-app generation, real-app trajectories, and coordinate-based grounding. We also introduce GUI-Score, a VLM-based metric that assesses Goal Achievement, Interaction Logic, Content Consistency, UI Plausibility, and Visual Quality. Evaluating 12 image generation models on Chinese and English subsets, we find that strong single-step performance does not consistently extend to coherent multi-step trajectories or accurate grounding. The highest aggregate GUI-Scores are 69.62 on the Chinese subset and 63.16 on the English subset. Qualitative analysis further identifies failures in text rendering, icon interpretation, and coordinate-conditioned transitions. Together, these results distinguish visual plausibility from functional reliability and identify limitations that must be addressed before generated GUIs can serve as dependable interaction environments.