Search arXiv⌕ Search

arXiv subjects

Qijun Ying

Publications and source records attributed to Qijun Ying.

3 recordsLinked to original sources

WATCH: World-aware Allied Trajectory and pose reConstruction for Camera and Human

Reconstructing global human motion from monocular video is fundamental to VR, graphics, and robotics, yet remains ill-posed due to depth ambiguity, motion ambiguity, and the entanglement of camera and human movements. Human-motion-centric methods achieve strong physical plausibility but leave two signals unused: camera orientation is processed through a fixed coordinate transformation with no independent supervision of its components, and camera velocity is discarded entirely despite being directly observable from SLAM. Camera-trajectory-centric methods use camera translation directly, but hard-decoding SLAM trajectories into human positions propagates depth errors and fails entirely under static cameras. We present WATCH (World-aware Allied Trajectory and pose reConstruction for Camera and Human). The key observation is that once camera orientation is made explicit, camera velocity becomes a natural additional input rather than an ambiguous one. We therefore decompose camera rotation into a network-estimated roll-pitch component and an analytically recoverable yaw, supervising each independently. This decomposition exposes a clean geometric interface through which camera velocity is incorporated as a learned spatial prior in the backbone, without the physically implausible artifacts that arise from hard-decoding. WATCH outperforms prior human-motion-centric methods on both static-camera (RICH) and dynamic-camera (EMDB) benchmarks in global trajectory accuracy, temporal smoothness, and physical plausibility, and remains robust when ground-truth camera is replaced with DPVO estimates.

cs.CV↗

From Camera to World: A Plug-and-Play Module for Human Mesh Transformation

Reconstructing accurate 3D human meshes in the world coordinate system from in-the-wild images remains challenging due to the lack of camera rotation information. While existing methods achieve promising results in the camera coordinate system by assuming zero camera rotation, this simplification leads to significant errors when transforming the reconstructed mesh to the world coordinate system. To address this challenge, we propose Mesh-Plug, a plug-and-play module that accurately transforms human meshes from camera coordinates to world coordinates. Our key innovation lies in a human-centered approach that leverages both RGB images and depth maps rendered from the initial mesh to estimate camera rotation parameters, eliminating the dependency on environmental cues. Specifically, we first train a camera rotation prediction module that focuses on the human body's spatial configuration to estimate camera pitch angle. Then, by integrating the predicted camera parameters with the initial mesh, we design a mesh adjustment module that simultaneously refines the root joint orientation and body pose. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods on the benchmark datasets SPEC-SYN and SPEC-MTP.

cs.CV↗

PI-HMR: Towards Robust In-bed Temporal Human Shape Reconstruction with Contact Pressure Sensing

Long-term in-bed monitoring benefits automatic and real-time health management within healthcare, and the advancement of human shape reconstruction technologies further enhances the representation and visualization of users' activity patterns. However, existing technologies are primarily based on visual cues, facing serious challenges in non-light-of-sight and privacy-sensitive in-bed scenes. Pressure-sensing bedsheets offer a promising solution for real-time motion reconstruction. Yet, limited exploration in model designs and data have hindered its further development. To tackle these issues, we propose a general framework that bridges gaps in data annotation and model design. Firstly, we introduce SMPLify-IB, an optimization method that overcomes the depth ambiguity issue in top-view scenarios through gravity constraints, enabling generating high-quality 3D human shape annotations for in-bed datasets. Then we present PI-HMR, a temporal-based human shape estimator to regress meshes from pressure sequences. By integrating multi-scale feature fusion with high-pressure distribution and spatial position priors, PI-HMR outperforms SOTA methods with 17.01mm Mean-Per-Joint-Error decrease. This work provides a whole

cs.CV↗