arXiv · 2609.36774
LexiconVLA: Learning Reusable Atomic Action Codebooks for Unseen Tasks
Abstract
Vision-language-action (VLA) models struggle to reuse recurring interactions in unseen tasks. Our diagnostic study reveals that reliable task completion does not imply consistent execution of constituent atomic actions across task contexts. We present LexiconVLA, a retrievable atomic-action lexicon for cross-task reuse. Global and detail codebooks capture shared interaction structure and fine-grained execution variation, respectively, preserving both reusable patterns and execution details. Visual-Atomic Action Alignment couples trajectory reconstruction from visual state changes with visual outcome prediction from action codes, grounding the lexicon in motion and its effects. We learn these codebooks with trajectory reconstruction and visual alignment on our AtomAction Dataset of 57,803 segments from 69 tasks. A planner and scene-aware adapter translate new goals into code-conditioned subtasks for a shared policy, without skill-specific experts or deployment-time parameter updates. Across five policy backbones on 26 RLBench tasks, LexiconVLA largely maintains performance on 18 seen tasks while improving success on 8 tasks held out from policy training. With BridgeVLA, unseen-task success rises from 16.67% to 34.17% (+17.50 percentage points), and overall success reaches 71.08%, the highest among methods with reported results. Real-robot experiments demonstrate stepwise execution and failure recovery.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zeming Wei, Jianheng Ye, Xinshuai Song, Sirui Chen, Yang Liu, Liang Lin. 2026-09-29. LexiconVLA: Learning Reusable Atomic Action Codebooks for Unseen Tasks. https://arxiv.org/abs/2609.36774
Cite the original work for its findings. Save a collection to share your selection of sources.