arXiv · 2609.25302
Multi-Agent Video Prediction: Self-Correcting Conditional Frames for Dynamic Scene Forecasting
Abstract
Transmission latency significantly degrades user quality of experience in real-time interactive perception systems. In remote driving, maintaining reliable visual feedback is critical for safe operation under dynamic network variability. Although video prediction offers a promising approach to compensate for short-term transmission delays and approximate near-zero-latency streaming, prediction-only methods remain vulnerable in highly dynamic scenes, especially when newly emerged objects appear during uplink outages. To address these challenges, we propose a multi-agent video prediction framework that combines continuous edge-side video prediction with lightweight mask-guided conditional frame reconditioning. The framework consists of three role-specialized agents: a continuous prediction agent for low-latency visual continuity, a vehicle-side trigger agent for detecting newly appeared objects, and a conditional reconditioning agent that repairs the predictor conditioning state using sparse mask guidance. This design enables semantic recovery of exogenous scene changes without requiring full-frame retransmission. We validate the proposed framework through extensive experiments on benchmark video data under realistic 5G communication traces. Results show that our method improves semantic recovery of novel objects while preserving perceptual quality and practical runtime efficiency under network-induced disruptions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Qixin Zhang, Ajay Kumar, Zhi-Li Zhang. 2026-09-21. Multi-Agent Video Prediction: Self-Correcting Conditional Frames for Dynamic Scene Forecasting. https://arxiv.org/abs/2609.25302
Cite the original work for its findings. Save a collection to share your selection of sources.