Tac2Pix: Image-Space Visuo-Tactile Fusion for Dexterous Manipulation
Vision and touch are naturally complementary for dexterous manipulation. Yet they are represented in fundamentally different forms: visual observations are dense and spatially structured, whereas tactile measurements are sparse, heterogeneous, and robot-centric. This mismatch makes it challenging to establish explicit spatial correspondence between the two modalities, particularly when they are fused in latent space. However, each tactile measurement is associated with a known location on the robot's kinematic chain, offering a localization prior that could bridge this representational gap. This raises a natural question: can this localization prior be used to construct a shared representation that explicitly aligns touch with vision? We introduce Tac2Pix, an image-space visuo-tactile fusion framework that projects sparse tactile measurements onto the RGB image plane using forward kinematics and camera calibration, then renders them as force-aware saliency maps. By expressing touch in the same pixel space as vision, Tac2Pix directly leverages pretrained 2D visual representations. We further employ zero-initialized input adaptation to preserve the original visual pathway at the start of policy learning. To assess whether this geometry-based interface yields consistent policy gains, we evaluate Tac2Pix on three simulated and three real-world dexterous manipulation tasks. In simulation, Tac2Pix consistently improves visuo-tactile policy learning across tasks, policy architectures, and visual backbones. In real-world experiments, it improves average success under visual occlusion by 26.7 percentage points over the strongest latent visuo-tactile fusion baseline, with gains persisting under both random and physical occlusions. Together, these results support image space as an effective interface for fusing sparse touch with pretrained visual representations.