arXiv · 2609.32595
RECAST: Recasting Vision-Language Semantics into an Actionable Cost Map for Robot Navigation
Abstract
Safe and robust robot navigation across diverse environments requires a high-level understanding of complex scenes and the ability to carry it into stable motion. Recent works tackle this with learning-based models trained at scale and with approaches built on vision-language models (VLMs). However, learning-based models break down outside their training distribution, while VLM-based approaches bring that understanding but rarely ground it in the scene or align the action with it. To address these limitations, we present RECAST, a robot navigation framework that combines the reasoning of a VLM with the spatial grounding of vision foundation models to build an Actionable Cost map. Given the robot's front view and the user's instruction, we first decompose the scene with the VLM, judging which surfaces are traversable, which objects pose a risk, which heading to prefer, and which gaps are passable. Vision foundation models then ground these surfaces and objects in the image, and all four judgments are spatially recast into one compact cost map. This map both conditions the trajectory decoders and scores their proposals to select the one to execute. As the VLM's answers trail the live scene, both steps draw on cost maps from two points in time: the pivot frame the VLM judged, which carries all four judgments, and the current frame, whose terrain and collision costs are rebuilt from the current image. RECAST improves success over the strongest prior method by 13.3 points in simulation and 31.4 points on a real quadruped, and reduces the collision rate relative to it by 9.6 and 14.3 points, reaching the lowest collision rate among all methods. The project page is available at https://recast-nav.github.io/
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Incheol Cho, Jintae Park, Jinkyu Kim, Jungbeom Lee, Jaegul Choo, Seokha Moon. 2026-09-26. RECAST: Recasting Vision-Language Semantics into an Actionable Cost Map for Robot Navigation. https://arxiv.org/abs/2609.32595
Cite the original work for its findings. Save a collection to share your selection of sources.