arXiv · 2610.05526
Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in Egocentric Assistants
Abstract
Egocentric assistants must connect what users say with what they see across long interaction histories. We formalize this challenge as Spatially grounded Conversational Reasoning (SpaCR): cross-scene, recall-oriented, and counterfactual spatial queries that combine user-stated facts with geometric evidence. Direct vision-language models incur high inference costs and context limits as histories grow, while keyframe selection and retrieval can omit objects or evidence needed for complete recall. We propose Spatially grounded Conversational Memory (SpaC-MEM), an object-centric working memory that uses 3D reconstruction and segmentation to ground conversational information in persistent physical objects. It compresses multimodal histories while preserving spatial evidence and allowing object-specific facts to be updated through dialogue. We also introduce Ego-SpaCR, a benchmark comprising 620 ScanNet video sessions augmented with 95 task-oriented conversations and 3,100 evaluation queries. SpaC-MEM achieves the highest overall answer accuracy among the evaluated methods and improves object recall while requiring substantially fewer reference input tokens than native-video baselines. Removing 3D spatial information substantially degrades performance, highlighting the importance of preserving spatial and conversational evidence together.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiazhou Liang, Liam Gallagher, Kiko Chen, David Guo, Armin Toroghi, Yifan Simon Liu, Scott Sanner. 2026-10-04. Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in Egocentric Assistants. https://arxiv.org/abs/2610.05526
Cite the original work for its findings. Save a collection to share your selection of sources.