arXiv · 2610.04139
From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models
Abstract
Predicting future spatial states supports collision avoidance and timely decision-making in dynamic environments. However, existing vision-language models (VLMs) and benchmarks for spatial reasoning primarily focus on observed scenes, leaving predictive spatial reasoning beyond the observed interval underexplored. To this end, we introduce SpatialMind, a metric-scale VLM for spatial reasoning and future prediction. Its metric depth adapter anchors spatial reasoning to real-world scale, while its progressive state chain establishes current spatial states and observed dynamics as the foundation for future prediction. Given a video prefix, SpatialMind predicts distances, motion directions, and spatial relations in both observed and unseen future frames. For training and evaluation, we build a scalable data engine that grounds entity descriptions in metric geometry to generate question-answer pairs and state supervision. Using this engine, we construct the SpatialMind-30K dataset and the SpatialMind-2K benchmark, both covering driving and everyday egocentric scenes. The benchmark spans eight tasks across three levels: current-state understanding, observed-dynamics understanding, and future prediction. Experiments show that SpatialMind substantially outperforms both general and spatially specialized models on our benchmark while achieving competitive zero-shot performance on VSI-Bench, OSI-Bench, and VLM4D.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Feiran Wang, Xiaoqi Wang, Ziwei Li, Wenbin He, Yan Yan, Liu Ren. 2026-10-02. From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models. https://arxiv.org/abs/2610.04139
Cite the original work for its findings. Save a collection to share your selection of sources.