arXiv · 2610.04301
EnvDreamer: Large-Scale Multimodal-to-Environment Generation for Embodied AI
Abstract
Large datasets and high capacity models have accelerated progress in vision and language. This work introduces a platform aimed at bringing comparable gains to embodied learning, world models, and robotics. We present EnvDreamer, a framework that uses large language and vision language models to generate Unreal Engine 5 environments for embodied AI and robot training. EnvDreamer enables sampling of large, diverse, interactive, customizable, and validator passed virtual environments for training and evaluation across navigation, interaction, and manipulation. We illustrate the platform with a large set of generated scenes and simple baselines. Policies trained on EnvDreamer generated environments, without explicit mapping or human task supervision, achieve competitive results on multiple embodied benchmarks spanning navigation, rearrangement, and manipulation. EnvDreamer also supports image-conditioned reconstruction for real-to-sim studies. Finally, we release EnvDreamer-20k, a dataset of 20,000 validator passed environments with task programs, scene graphs, trajectories, and metadata to support reproducible benchmarking.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kabir Swain, Sijie Han, Antonio Torralba. 2026-10-03. EnvDreamer: Large-Scale Multimodal-to-Environment Generation for Embodied AI. https://arxiv.org/abs/2610.04301
Cite the original work for its findings. Save a collection to share your selection of sources.