Search arXiv⌕ Search

arXiv · 2609.38077

MUGEN: Interactive Panoramic World Exploration via Camera Control

Abstract

Interactive panoramic video generation aims to synthesize immersive 360\textdegree{} videos that remain visually coherent while following user-specified camera trajectories during exploration. However, progress is limited by a coupled data-and-model gap: existing panoramic video datasets are often short, weakly annotated, or lack camera trajectories, while existing camera-controlled video generation models are designed for perspective videos and do not directly support panoramic geometry. In this paper, we introduce MUGEN and Wan360 to address these limitations. MUGEN is a large-scale real-world panoramic video dataset tailored to interactive 360-degree world exploration, comprising over 1,300 hours of at least 4K panoramic videos with rich semantic and geometric annotations. Built on MUGEN, we further present Wan360, a camera-controllable interactive panoramic video generation model. Panoramic videos are commonly represented by EquiRectangular Projection (ERP), which unfolds a spherical 360-degree view into a rectangular frame with cyclic longitude seams and pole distortions. To this end, Wan360 introduces three parameter-free ERP-aware components: periodic longitude RoPE for seam-consistent positional encoding, ERP-aware padding for reducing boundary artifacts, and random roll yaw for consistent learning. For camera control, Wan360 uses a panoramic Plücker embedding that represents camera motion with ERP rays rather than perspective pinhole rays. Experiments show that MUGEN serves as a data foundation for panoramic world exploration, and that Wan360 enables high-quality, temporally coherent, camera-controllable 360-degree video generation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jiaming Tan, Zhen Li, Shuwei Shi, Minggui Teng, Siqi Yang, Yuwei Wu, Bo Zheng, Chuanhao Li, Kaipeng Zhang. 2026-09-29. MUGEN: Interactive Panoramic World Exploration via Camera Control. https://arxiv.org/abs/2609.38077

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Long-Tailed 3D Detection via Multi-Modal Fusion

Contemporary autonomous vehicle (AV) benchmarks have significantly advanced multimodal (LiDAR+RGB) 3D detection. However, despite the naturally long-tailed distribution of object classes, existing benchmarking protocols primarily focus on frequent categories (e.g., pedestrian and car), largely overlooking rare but safety-critical classes such as stroller and emergency vehicle. In practice, reliable detection of both common and rare classes is essential for safe autonomous driving. We formalize this problem as Long-Tailed 3D Detection (LT3D), where evaluation encompasses all annotated classes, including rare ones. To address LT3D, we introduce hierarchical losses that promote feature sharing across classes, diagnostic metrics that assign partial credit to semantically reasonable mistakes with respect to the semantic hierarchy (e.g., confusing a child with an adult), and a multimodal late-fusion (MMLF) framework to fuse detections. In particular, we show that rare-class accuracy benefits substantially from MMLF of independently trained uni-modal LiDAR and RGB detectors. Because of the modular design, unlike prevailing end-to-end trained multi-modal detectors that require paired LiDAR-RGB data, MMLF enables the use of advanced unimodal detectors that are trained on large-scale uni-modal datasets with sufficient data for rare classes. Lastly, we examine three fundamental design choices in MMLF, including the RGB detector representation (2D vs. 3D), cross-modal association (3D vs. image plane), and fusion strategy. We find that 2D RGB detectors recognize rare classes more reliably than 3D RGB detectors, image-plane association is more robust to depth estimation errors, and probabilistic score-calibrated fusion consistently yields the best performance. Extensive experiments on nuScenes and Argoverse2 demonstrate substantial improvements of MMLF, establishing a new state of the art.

cs.CV↗

Towards Formal Verification of Deep Neural Networks for Object Detection

Deep neural networks (DNNs) are widely used in real-world computer vision applications, yet they remain vulnerable to errors and adversarial attacks. Formal verification offers a systematic approach to identify and mitigate these vulnerabilities, enhancing model robustness and reliability. While most existing verification methods focus on image classification models, this work extends formal verification to the more complex domain of object detection models. We propose a formulation for verifying the robustness of such models and demonstrate how state-of-the-art verification tools, originally developed for classification, can be adapted for this purpose. Through a comprehensive evaluation, we highlight the ability of formal verification to uncover vulnerabilities in object detection models, and derive formal robustness guarantees, underscoring the potential and need to further extend verification efforts in this domain. This work lays the foundation for further research into formal verification of object detection models across a broader range of computer vision applications. Our source code is publicly available online.

cs.CV↗

Patch Rebirth: Fast and Transferable Model Inversion of Vision Transformers

Model inversion is a widely adopted technique in data-free learning that reconstructs synthetic inputs from a pretrained model through iterative optimization, without access to original training data. Unfortunately, its application to state-of-the-art Vision Transformers (ViTs) poses a major computational challenge, due to their expensive self-attention mechanisms. To address this, Sparse Model Inversion (SMI) was proposed to improve efficiency by pruning and discarding seemingly unimportant patches, which were even claimed to be obstacles to knowledge transfer. However, our empirical findings suggest the opposite: even randomly selected patches can eventually acquire transferable knowledge through continued inversion. This reveals that discarding any prematurely inverted patches is inefficient, as it suppresses the extraction of class-agnostic features essential for knowledge transfer, along with class-specific features. In this paper, we propose Patch Rebirth Inversion (PRI), a novel approach that incrementally detaches the most important patches during the inversion process to construct sparse synthetic images, while allowing the remaining patches to continue evolving for future selection. This progressive strategy not only improves efficiency, but also encourages initially less informative patches to gradually accumulate more class-relevant knowledge, a phenomenon we refer to as the Re-Birth effect, thereby effectively balancing class-agnostic and class-specific knowledge. Experimental results show that PRI achieves up to 10x faster inversion than standard Dense Model Inversion (DMI) and 2x faster than SMI, while consistently outperforming SMI in accuracy and matching the performance of DMI.

cs.CV↗