arXiv · 2609.24189
A Topological Representation with Object-Path Graphs for Open-Vocabulary Instance Navigation
Abstract
Vision-language navigation requires embodied agents to navigate environments using natural language instructions and visual observations. Existing approaches typically decompose navigation into sequential language-guided decisions or rely on online exploration without prior environmental knowledge. Scene graph representations offer compact semantic memory but remain decoupled from downstream navigation, which still depends on dense metric maps. To close this gap, we propose an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation. The proposed representation jointly supports semantic grounding, graph-based localization, and navigation within a single lightweight topological framework. Building on this graph, we introduce a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly over the graph without dense metric reconstruction. Experiments on HM3D and Replica demonstrate competitive performance in open-vocabulary object grounding through the proposed hierarchical graph structure, while achieving effective navigation performance. Real-world robot experiments further validate the practicality of the proposed framework.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Linwei Zheng, Daojie Peng, Bingtao Wang, Haoang Li, Jun Ma. 2026-09-21. A Topological Representation with Object-Path Graphs for Open-Vocabulary Instance Navigation. https://arxiv.org/abs/2609.24189
Cite the original work for its findings. Save a collection to share your selection of sources.