arXiv · 2601.14056
POCI-Diff: 3D-Layout Guided Diffusion for Controllable Synthetic Surveillance Data Generation
Abstract
Training robust visual surveillance models requires large-scale datasets with precise spatial annotations, yet collecting real surveillance data is costly, privacy-sensitive, and often legally constrained. Synthetic data generation offers a compelling alternative, but existing methods lack fine-grained 3D control over object placement and appearance, limiting geometric consistency across camera viewpoints. We introduce POCI-Diff (Positioning Objects Consistently and Interactively), a framework that generates annotated synthetic scenes from explicit 3D bounding-box layouts with per-object semantic control. By integrating Blended Latent Diffusion with depth-conditioned ControlNet, POCI-Diff synthesises complex multi-object scenes in a single forward pass, binding individual text descriptions to specific 3D locations. We further propose a warping-free editing pipeline supporting object insertion, removal, and transformation via regeneration, enabling efficient scene variation for data augmentation. Object identity across edits is preserved by conditioning on reference images via IP-Adapter, ensuring appearance consistency throughout interactive scene manipulation. Experiments show that POCI-Diff outperforms state-of-the-art 3D layout-guided generation methods in visual fidelity and layout adherence, while eliminating warping-induced geometric artifacts.
Explore related subjects
Keep this discovery
Andrea Rigo, Luca Stornaiuolo, Weijie Wang, Mauro Martino, Bruno Lepri, Nicu Sebe. 2026-08-29. POCI-Diff: 3D-Layout Guided Diffusion for Controllable Synthetic Surveillance Data Generation. https://arxiv.org/abs/2601.14056
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.