Search arXiv⌕ Search

arXiv · 2607.23797

Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map

Abstract

A robot carrying a persistent, behavior-annotated map faces two very different planning questions, and its memory answers only one of them well. The spatial-navigation question - how to walk around a room - we address first, and report a negative: building on Vision-Language-Motion Maps (VLMM), a behavior-aware planner cost cuts a planning-time objective by ~35% over 28 AI2-THOR scenes, but under closed-loop execution the benefit nearly vanishes (~4%) and an on-demand vision-language model (VLM) does as well. The resource-allocation question is different: under a limited perception budget, what should the robot attend to right now to keep its own map fresh? Framing re-perception as this attention decision, we show a persistent map's memory - change-history, or even just recency of last sighting - yields the best of the re-perception schedules we test (held-out), while the memoryless category prior is the weakest of them under the sqrt-law schedule -- though under a Whittle index it leads memory until heterogeneity is real. The gain grows with per-instance heterogeneity as a Cauchy-Schwarz bound predicts, tracking Var(sqrt(lambda)), the variance of root-volatility, and reallocates budget toward the important objects the schedule protects; against a real CLIP movability prior it is +21-26%, of which roughly half survives once that prior's saturated scale is calibrated (+7-13%). The map's full combination earns its keep when the task is language-conditioned: told what to keep track of, VLMM grounds the relevant objects (open-vocabulary) and tracks their change (memory), beating a relevance-weighted recency baseline (+2.9% over 26 queries, at full heterogeneity; the ordering reverses when instance rates track category norms) - and a category prior (+9.2%). The map earns its keep not by telling the robot how to walk around a room, but by telling it what to pay attention to.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dibyendu Ghosh. 2026-09-12. Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map. https://arxiv.org/abs/2607.23797

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Physics-Guided Residual Reinforcement Learning for Humanoid Narrow-Path Traversal

Traversing narrow paths is challenging for humanoid robots due to the sparse and safety-critical footholds required. Purely template-based or end-to-end reinforcement learning-based methods suffer from such harsh terrains. This paper proposes a two stage training framework for such narrow path traversing tasks, coupling a template-based foothold planner with a low-level foothold tracker from Stage-I training and a lightweight perception aided foothold modifier from Stage-II training. With the curriculum setup from flat ground to narrow paths across stages, the resulted controller in turn learns to robustly track and safely modify foothold targets to ensure precise foot placement over narrow paths. This framework preserves the interpretability from the physics-based template and takes advantage of the generalization capability from reinforcement learning, resulting in easy sim-to-real transfer. The learned policies outperform purely template-based or reinforcement learning-based baselines in terms of success rate, centerline adherence and safety margins. Validation on a Unitree G1 humanoid robot yields successful traversal of a 0.2m wide and 3m long beam for 20 trials without any failure.

cs.RO↗

HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

End-to-end autonomous driving models increasingly benefit from large vision-language models for semantic understanding, yet safe and reliable planning under long-tail conditions remains challenging, particularly in mixed-traffic environments involving heterogeneous road users and rare safety-critical interactions. This paper proposes HERMES, a holistic risk-aware end-to-end multimodal driving framework that explicitly incorporates long-tail semantic knowledge into trajectory planning. HERMES employs a foundation-model-assisted annotation pipeline to construct structured Long-Tail Scene Context and Long-Tail Planning Context, capturing hazard-centric scene information, maneuver intent, and risk-aware planning guidance. A Tri-Modal Driving Module then integrates multi-view visual observations, historical ego-motion, and long-tail semantic instructions through intent- and risk-aware conditioning for trajectory generation. Extensive experiments on a large-scale real-world long-tail driving benchmark demonstrate consistent improvements over representative recent baselines in overall planning performance and across diverse safety-critical scenarios. Ablation studies further validate the effectiveness and complementary roles of the major components within HERMES.

cs.RO↗

StageCraft: Execution Aware Mitigation of Distractor and Obstruction Failures in VLA Models

Large scale pre-training on text and image data along with diverse robot demonstrations has helped Vision Language Action models (VLAs) to generalize to novel tasks, objects and scenes. However, these models are still susceptible to failure in the presence of execution-time impediments such as distractors and physical obstructions in the robot's workspace. Existing policy improvement methods finetune base VLAs to improve generalization, yet they still struggle in unseen distractor settings. To address this problem, we investigate whether internet-scale pretraining of large vision-language models (VLMs) can be leveraged to reason about these impediments and mitigate policy failures. To this end, we propose StageCraft, a training-free approach to improve pretrained VLA policy performance by manipulating the environment's initial state using VLM-based in-context reasoning. StageCraft takes policy rollout videos and success labels as input and leverages VLM's reasoning ability to infer which objects in the initial state need to be manipulated to avoid anticipated execution failures. StageCraft is an extensible plug-and-play module that does not introduce additional constraints on the underlying policy, and only requires a few policy rollouts to work. We evaluate performance of state-of-the-art VLA models with StageCraft and show an absolute 40% performance improvement across three real world task domains involving diverse distractors and obstructions. Our simulation experiments in RLBench empirically show that StageCraft tailors its extent of intervention based on the strength of the underlying policy and improves its performance with more in-context samples. Videos of StageCraft in effect can be found at https://stagecraft-decorator.github.io/stagecraft/ .

cs.RO↗