arXiv · 2610.03013
RYOPO: Bringing End-to-End Category-Level Object Pose Estimation into Real Time
Abstract
Category-level object pose estimation predicts the rotation, translation, and metric size of unseen instances within known categories. Many accurate RGB-D methods rely on external instance segmentation and crop-based pose estimation, introducing separate stages and object-dependent processing costs that hinder real-time inference. To bring accurate pose estimation into real time, we present \ours{}, an end-to-end trainable query-based RGB-D set predictor. It jointly detects and segments objects and estimates their \mbox{9-DoF} poses without explicit CAD-derived shape priors or a separately trained instance segmentor. Shared image and scene encoding avoids repeated per-object crop encoding. A query-conditioned geometry pathway associates observed 3D points and RGB features with object queries and incorporates shared scene context. Object-centric refinement uses the resulting point descriptors to update an explicit pose state through pose-conditioned cross-attention and recurrent residual corrections. On NOCS, \ours{} substantially improves on published RGB-D joint detection and pose estimation results. It achieves competitive performance compared with two-stage methods under all-object evaluation on REAL275 and HouseCat6D, while enabling real-time full-frame pose estimation at $31.8$ FPS on an RTX~A6000. Project page: https://yopo-series.github.io/RYOPO-project-page/.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hakjin Lee, Junghoon Seo, Jaehoon Sim. 2026-10-02. RYOPO: Bringing End-to-End Category-Level Object Pose Estimation into Real Time. https://arxiv.org/abs/2610.03013
Cite the original work for its findings. Save a collection to share your selection of sources.