arXiv · 2603.07875
Foundation and Small Models Coordination for Visuomotor Policy Learning
Abstract
Visuomotor policy learning enables robots to perform a wide range of tasks, but small policy models often remain sensitive to changes in object and background appearance. In this work, we investigate the coordination of pretrained vision foundation models with small policy models to improve appearance generalization. We propose a framework in which a small policy model operates on task-relevant visual observations constructed through semantic repainting. A segmentation foundation model identifies the robot and target object, which are rendered with fixed role colors on a constant background. An alternative representation replaces the target's role color with normalized monocular depth predicted by a depth foundation model, providing additional geometric cues. The perception models are adapted using in-distribution data where needed and held fixed during policy training. This design combines the perceptual capabilities of foundation models with a small policy model trained on the resulting observations for action prediction. Evaluations with flow matching policies on simulation benchmarks, together with experiments on two real-world robotic tasks, demonstrate substantial improvements in task success under the evaluated appearance shifts.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Haoran Ding, Liang Ma, Yaxun Yang, Wen Yang, Tianyu Liu, Xiaodan Liang, Dezhen Song, Ivan Laptev, Yoshihiko Nakamura, Anqing Duan. 2026-09-15. Foundation and Small Models Coordination for Visuomotor Policy Learning. https://arxiv.org/abs/2603.07875
Cite the original work for its findings. Save a collection to share your selection of sources.