StreamTTO: Efficient Online Test-Time Optimization for Video Depth Completion
Monocular depth foundation models generalize across diverse scenes, but recovering accurate metric depth consistent with a target sensor remains challenging under sensor variation and domain shift. Test-time optimization (TTO) guided by sparse depth addresses this limitation, yet independent per-frame adaptation incurs substantial computational cost. We propose StreamTTO, an online test-time optimization framework for video depth completion that reuses visual features and adaptation state. A frozen feature extractor with causal temporal attention incorporates past visual context into features cached for repeated decoder optimization. Causal Sliding-Window Optimization directly updates a shared depth decoder using cached features and sparse observations from recent frames, requiring only two passes over the window per incoming frame after initialization. Reusing recent observations and the adapted decoder substantially reduces optimization steps. We also introduce MS-Depth, comprising approximately 1.45M synchronized RGB--LiDAR frames from 40 hours of recordings. Its uninterrupted sequences capture transitions among indoor, underground, and outdoor environments under daytime and nighttime conditions, enabling evaluation of adaptation to changes in illumination, scene structure, and depth range. Experiments on KITTI, Bonn, NYUv2 Raw, TUM RGB-D, DDAD, and MS-Depth demonstrate competitive depth completion accuracy. On MS-Depth, StreamTTO achieves average processing speeds of approximately 15~FPS with VGGT and 25~FPS with MoGe-2 at an input resolution of 518 X 392.