arXiv · 2609.25785
VisForce: Visual Grounding of Current and Desired Forces for Goal-Conditioned Dexterous Manipulation
Abstract
Vision-Language-Action (VLA) models have emerged as general-purpose robotic manipulation policies. However, in dexterous hand manipulation, contact forces are typically provided as separate states or force-specific representations, making it difficult to explicitly represent the spatial correspondence between force and their corresponding visual locations. In this work, we propose VisForce, which visually grounds the current and desired forces at their corresponding fingertip locations. VisForce renders current and desired visual force cues on the current wrist image and a task-specific goal image, and combines the two representations through goal-conditioned cross-attention to generate force-aware actions. We evaluate VisForce using a real UR10 robot equipped with an RH56F1 dexterous hand through force-conditioned grasping and three multi-stage manipulation tasks. In force-conditioned grasping experiments, VisForce exhibited a consistent grip-force response as the desired force increased, and achieved grasp-and-lift success rates of 70% and 80% for an egg and a toothpaste tube, respectively. It further achieved final success rates of 70%, 55%, and 40% on cup insertion/bottle pouring, tong-assisted bread transfer, and slip-modulated peg-in-hole, respectively. These results show that fingertip-aligned visual force representations can be effectively used for force-aware conditioning in VLA-based dexterous hand manipulation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jung-Woo Lee, Soo-Chul Lim. 2026-09-22. VisForce: Visual Grounding of Current and Desired Forces for Goal-Conditioned Dexterous Manipulation. https://arxiv.org/abs/2609.25785
Cite the original work for its findings. Save a collection to share your selection of sources.