arXiv · 2609.26071
StrataVLA: Hierarchical and Efficient 3D Geometric Grounding for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models inherit strong semantic priors from large-scale vision-language pretraining, yet remain limited in robotic manipulation by insufficient 3D spatial awareness. Existing approaches either require explicit depth or point-cloud inputs, compress geometry into training-time supervision, or inject it only at the model input or action expert, leaving the vision-language backbone without persistent access to task-relevant spatial information. We introduce StrataVLA, a plug-and-play framework for hierarchical geometric grounding. A frozen geometry foundation model extracts shared geometric features from RGB observations, while sparse, layer-specific Geometry Adapters allow visual representations at selected backbone depths to retrieve relevant geometric evidence through cross-attention. To make inference-time geometry practical, StrataVLA further combines task-aware routing with an LRU feature cache that exploits temporal redundancy during task manipulation. Experiments on LIBERO, SimplerEnv, and real-world manipulation demonstrate consistent gains over strong VLA baselines. StrataVLA achieves 98.53% average success on LIBERO suites while reducing geometry-model invocations by up to 88%, establishing hierarchical geometry injection as an effective and efficient way to achieve spatially grounded robotic control.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jin Cui, Zhaoyu Pu, Botao Cai, Jun Ye, Xinyue Long, Boran Zhao, Pengju Ren. 2026-08-03. StrataVLA: Hierarchical and Efficient 3D Geometric Grounding for Vision-Language-Action Models. https://arxiv.org/abs/2609.26071
Cite the original work for its findings. Save a collection to share your selection of sources.