arXiv · 2508.06291
VLEM: Real-Time 3D Vision-Language Embedding Mapping
Abstract
Semantic scene understanding in robotics requires representations that are both metric-accurate and queryable via natural language in real-time. While recent Vision-Language Models enable powerful 2D image-text alignment, their integration into real-time 3D mapping systems remains challenging due to their requirements on ground truth poses, computational cost, and memory constraints. We present VLEM (Vision-Language Embedding Mapping), a real-time framework for integrating pixel-aligned 2D vision-language embeddings from various backends into a globally consistent, metric-accurate 3D representation, requiring only a raw RGB-D stream. Compared to ConceptFusion, Open-Fusion, and RayFronts, VLEM provides better open-set segmentation performance and a more compact representation. We further demonstrate VLEM's versatility in interactive real-time robotic manipulation tasks and mobile mapping scenarios.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Christian Rauch, Björn Ellensohn, Linus Nwankwo, Vedant Dave, Elmar Rueckert. 2026-09-16. VLEM: Real-Time 3D Vision-Language Embedding Mapping. https://arxiv.org/abs/2508.06291
Cite the original work for its findings. Save a collection to share your selection of sources.