arXiv · 2610.06594
VGGT-Bridge: Beyond Sequential Pose Graphs via Coarse-Stride Skip Edges
Abstract
Feed-forward visual geometry transformers such as VGGT reconstruct dense 3D structure from images in a single forward pass, simplifying multi-view 3D reconstruction. However, their quadratic attention complexity makes them difficult to scale to long sequences with thousands of frames. Chunk-and-align frameworks address this by splitting a long sequence into overlapping chunks and stitching their local reconstructions into a pose graph. Yet existing methods connect only sequentially adjacent chunks, so small per-frame errors accumulate along the chain into large-scale drift. To move beyond sequential edges, we propose VGGT-Bridge, which adds long-range skip edges that directly constrain non-adjacent chunks without retraining. By running VGGT on sparsely sampled coarse chunks, each coarse chunk bridges distant fine chunks into a single direct constraint. We further turn VGGT's first-frame scale bias into a drift correction by feeding selected coarse chunks in reverse, and a loop-aware policy keeps this reversal compatible with existing loop closures. VGGT-Bridge reduces ATE by 28.3% on KITTI Odometry, 18.8% on Virtual KITTI, and 10.0% on Waymo Open over the SwiftVGGT baseline, achieving the best performance among all chunk-and-align methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sungjae Choi, Hanna Bae, Sunghyun Baek, Junmo Kim. 2026-10-05. VGGT-Bridge: Beyond Sequential Pose Graphs via Coarse-Stride Skip Edges. https://arxiv.org/abs/2610.06594
Cite the original work for its findings. Save a collection to share your selection of sources.