Search arXiv⌕ Search

arXiv · 2609.32175

OneFixer: High-Quality and Consistent One-Step Autoregressive 3DGS Refinement for Driving Scenes

Abstract

Autoregressive video diffusion is a promising render-time fixer for 3D Gaussian Splatting (3DGS) in autonomous-driving simulation, but deployment demands high visual quality and temporal consistency at low latency. This is especially hard for one-step causal generation, where each imperfect prediction immediately becomes context for subsequent frames. Existing approaches stabilize rollouts through staged training with multiple modules and rollout-aware regularization, yet one-step quality still falls short of what deployment requires. We introduce OneFixer, a one-step autoregressive video-diffusion fixer trained in a single task-specific adaptation stage. Our key idea is a deployment-matched shared rollout: the model's own one-step predictions serve as the causal context for flow matching, exposing training to deployment-time errors, while the same rollout receives direct pixel-space perceptual supervision to preserve fine detail. Because the predictions optimized for current-frame quality are exactly those reused as future context, fidelity and autoregressive robustness are learned jointly, without bidirectional-to-causal conversion or teacher-student distillation. OneFixer further exploits cues that driving simulation readily provides, lane geometry and dynamic-agent states, to improve geometric fidelity. On Waymo and proprietary driving scenes with 900-frame rollouts, OneFixer achieves the lowest FVD, LPIPS, and DISTS among all baselines at one step, with temporal consistency matching or exceeding multi-stage DMD pipelines. Under identical backbone and conditioning, it matches a multi-stage DMD-with-Self-Forcing pipeline in under half the GPU-hours and keeps improving beyond its plateau. In closed-loop simulation with a driving policy, OneFixer reduces the collision rate by a third relative to raw 3DGS rendering. Project page: https://onefixer-web.vercel.app/

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Boseong Jeon, Junhyeop Lee, Juhan Cha, Hayoung Kim. 2026-09-26. OneFixer: High-Quality and Consistent One-Step Autoregressive 3DGS Refinement for Driving Scenes. https://arxiv.org/abs/2609.32175

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing

Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybrid while preserving the source video's motion and layout. We propose MoCA-Video, a training-free framework that steers a frozen video-diffusion denoising trajectory through concept-localized reference injection. At selected low-noise steps, MoCA-Video uses concept attention to localize the target object and injects the reference latent into the localized region, where object structure has formed but appearance remains editable. A momentum-based correction carries the injected prediction across frames to encourage coherent concept integration through the sequence. We further introduce CASS, a CLIP-based metric that measures the output's directional alignment shift toward the reference and away from the source prompt. Using the denoiser's internal attention avoids an external localization model; in our A100 FP16 setup, MoCA-Video takes 3.2 seconds per output frame, excluding preprocessing. Across the evaluated baselines, MoCA-Video achieves the highest CASS, rel-CASS, and ImageReward, while LPIPS-T and FVD expose separate temporal-coherence and video-quality trade-offs.

cs.CV↗

Matrix-game 2.0: An open-source, real-time, and streaming interactive world model

Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynamics and interactive behaviors. However, existing interactive world models depend on bidirectional attention and lengthy inference steps, severely limiting real-time performance. Consequently, they are hard to simulate real-world dynamics, where outcomes must update instantaneously based on historical context and current actions. To address this, we present Matrix-Game 2.0, an interactive world model generates long videos on-the-fly via few-step auto-regressive diffusion. Our framework consists of three key components: (1) A scalable data production pipeline for Unreal Engine and GTA5 environments to effectively produce massive amounts (about 1200 hours) of video data with diverse interaction annotations; (2) An action injection module that enables frame-level mouse and keyboard inputs as interactive conditions; (3) A few-step distillation based on the casual architecture for real-time and streaming video generation. Matrix Game 2.0 can generate high-quality minute-level videos across diverse scenes at an ultra-fast speed of 25 FPS. We open-source our model weights and codebase to advance research in interactive world modeling.

cs.CV↗

Mitigating Cross-Image Information Leakage in Multi-Image Understanding with Large Vision-Language Models

Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handling multi-image inputs. While this degradation has been observed in prior work, its nature remains poorly understood. We empirically observe visual elements from different images become entangled in the model's representations and responses. We refer to this phenomenon as cross-image information leakage. To address this issue, we propose FOCUS, a training-free and architecture-agnostic method. FOCUS masks all but one image with random noise, guiding the model to focus on the single clean image. This process is applied across the target images to obtain logits under partially masked contexts. These logits are aggregated and then refined using a noise-only reference input, which suppresses the leakage and yields more accurate outputs. FOCUS consistently improves performance on diverse multi-image benchmarks. We further show that FOCUS generalizes to video understanding, extending its applicability beyond static multi-image inputs. This demonstrates that FOCUS offers a general solution for enhancing multi-image reasoning without additional training or architectural modifications.

cs.CV↗