arXiv · 2509.04276
PAOLI: Pose-free Articulated Object Learning from Sparse-view Images
Abstract
We present a method for modeling articulated objects from sparse images with unknown camera poses. Existing approaches require dense multi-view observations and ground-truth camera poses shared across articulation states, whereas our approach operates with as few as four unposed views per articulation state, without requiring ground-truth poses or shared cross-state calibration. The key challenge is that sparse-view reconstruction can recover each articulation state independently, but does not provide the cross-state correspondences required for part segmentation and kinematic reasoning. We address this by optimizing a dense deformation field that aligns independently reconstructed states and establishes dense correspondences across them. An iterative robust estimation procedure then separates static and moving parts and recovers their kinematic parameters. Finally, we refine geometry, appearance, and motion using self-supervised losses that enforce cross-view and cross-state consistency. Experiments on standard benchmarks and real-world examples show that our method produces accurate articulated object representations under substantially weaker input assumptions than prior work.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jianning Deng, Kartic Subr, Hakan Bilen. 2026-09-19. PAOLI: Pose-free Articulated Object Learning from Sparse-view Images. https://arxiv.org/abs/2509.04276
Cite the original work for its findings. Save a collection to share your selection of sources.