arXiv · 2609.27227
Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints
Abstract
Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic reconstruction network for estimating instrument position, orientation and jaw angle from monocular video. Our visual representation combines global attention pooling of frozen DINOv3 features with local pooling at instrument landmarks from fine-tuned SAM 3.1 masks. Our shared Transformer encoder and temporal convolutional heads integrate this representation with mask geometry, monocular depth and visual state estimates from arm-specific multilayer regression networks. Our position branch predicts displacement magnitude and direction separately to preserve traveled distance. We fit trajectories to predicted state observations and motion increments by differentiable weighted least squares, expressing quaternion observations relative to cumulative predicted rotations to obtain a quadratic orientation objective. We evaluate reconstruction across 2,802 Open-H episodes. Compared with LiveMAE on the main Open-H benchmark, our method reduces path-length mean absolute error from 0.45 to 0.34\,cm and increases temporal mean average precision for motion segmentation from 44.54\% to 54.44\%.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mehmet Kerem Turkcan, Soham Samal, Zoran Kostic. 2026-09-23. Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints. https://arxiv.org/abs/2609.27227
Cite the original work for its findings. Save a collection to share your selection of sources.