Hand-centric Human-to-Robot Trajectory Transfer from Video Demonstrations via Object Category-agnostic Temporal Localization
Human videos provide a scalable source of demonstrations for robot imitation learning. However, converting them into executable and semantically aligned robot trajectories requires determining \emph{when} task-relevant interactions occur, \emph{how} demonstrated human grasps should be retargeted across embodiments, and \emph{which} post-contact motions should be reproduced. In this paper, we present \textit{HOCALT} (\emph{H}and--\emph{O}bject \emph{C}ategory--\emph{A}gnostic \emph{L}ocalization and \emph{T}ransfer), a hand-centric retargeting framework that unifies 3D hand motion reconstruction, temporal contact localization, and cross-embodiment trajectory generation. Given synchronized stereo videos, \textit{HOCALT} identifies task-relevant contact intervals by jointly reasoning over reconstructed 3D hand articulation and category-agnostic object motion cues. These intervals serve as temporal anchors for initializing cross-embodiment transfer, where the demonstrated human grasps are retargeted into multi-modal robot grasp hypotheses. Within each interval, these hypotheses are then propagated along the demonstrated hand motion, yielding executable trajectories. Finally, we augment the transferred trajectories to generate diverse variants from a single demonstration. Across various tasks, \emph{HOCALT} outperforms VLM-based temporal localization baselines, achieving higher replay success than existing retargeting approaches.