Search arXiv⌕ Search

arXiv · 2609.36565

Toward Generative Video Communication: A Dual-Stream Digital Transmission Framework

Abstract

Generative video communication has shown promise for bandwidth-constrained wireless transmission and has the potential to support personalized content delivery. In this article, we propose a dual-stream digital generative video communication (DGVC) framework that integrates a traditional digital link with a generative link. The traditional link provides source-grounded visual references, while the generative link conveys compact semantic and perceptual information for receiver-side generation. We further discuss three bandwidth-dependent operating regimes and key technologies for dual-stream coordination, synchronization, reliability, and latency control. A practical case study demonstrates the perceptual and temporal-quality benefits of DGVC under wireless fading channels. Finally, we discuss open challenges and future research directions for generative video communication.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bingyan Xie, Longyu Zhou, Tianhao Liang, Yongpeng Wu, Zehui Xiong, Wenjun Zhang, Tony Q. S. Quek. 2026-09-29. Toward Generative Video Communication: A Dual-Stream Digital Transmission Framework. https://arxiv.org/abs/2609.36565

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling

Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak textual controllability under visual-text conflict and imprecise stylistic control due to entangled temporal and timbre information in reference audio. Moreover, the lack of standardized benchmarks limits systematic evaluation. We propose ControlFoley, a unified multimodal V2A framework that enables precise control over video, text, and reference audio. We introduce a joint visual encoding paradigm that integrates CLIP with a spatio-temporal audio-visual encoder to improve alignment and textual controllability. We further propose temporal-timbre decoupling to suppress redundant temporal cues while preserving discriminative timbre features. In addition, we design a modality-robust training scheme with unified multimodal representation alignment (REPA) and random modality dropout. We also present VGGSound-TVC, a benchmark for evaluating textual controllability under varying degrees of visual-text conflict. Extensive experiments demonstrate state-of-the-art performance across multiple V2A tasks, including text-guided, text-controlled, and audio-controlled generation. ControlFoley achieves superior controllability under cross-modal conflict while maintaining strong synchronization and audio quality, and shows competitive or better performance compared to an industrial V2A system. Code, models, datasets, and demos are available at: https://github.com/xiaomi-research/controlfoley.

cs.MM↗

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning

Vision-Language Models (VLMs) are widely used for visual understanding, yet current evaluation protocols fail to assess whether these capabilities are grounded in physical reasoning. To address this gap, we introduce Retrospective Physical Process Reasoning, a new evaluation paradigm to reason backward from outcomes under explicit physical constraints. Building on the paradigm, we present RetroHolmes, the first real-world benchmark for Retrospective Physical Process Reasoning, comprising object-centric image pairs annotated with reachability labels and causal step sequences across diverse physical transitions. Using RetroHolmes, we analyze VLMs and uncover systematic failure modes, including judgment bias in reachability assessment and belief dominance over physical evidence, mirroring sycophancy behavior observed in large language models. Our quantitative analyses link these failures to reliance on linguistic priors and attention concentrated on visually invariant regions, suggesting limited physical simulation of the intermediate states connecting visual endpoints. To address these limitations, we propose Simulate-and-Verify, an analysis-by-synthesis framework that grounds reachability judgment and step reconstruction in visual simulation. Experiments show that Simulate-and-Verify improves judgment accuracy by 21.67 percentage points and reduces belief dominance by 10.39 percentage points compared with GPT-5.5, demonstrating the effectiveness of visual simulation in grounding physical reasoning.

cs.MM↗

ReVR: Dual-Path Concept Reasoning for Multimodal Fake News Detection

Vision-language models (VLMs) support multimodal fake news detection (FND) by producing explicit analyses. Recent methods further improve interpretability by organizing verification knowledge into explicit concepts. However, two questions remain: how to improve the reliability and applicability of verification concepts, and how to effectively apply reusable concepts to verify unseen news. We propose \textbf{ReVR}, a dual-path reasoning framework that constructs and applies reusable verification concepts for multimodal fake news detection. An agentic workflow grounds and consolidates candidate concepts, while statistical profiles characterize their historical behavior. During inference, a coverage-oriented path aggregates evidence from the complete concept library using a trainable encoder, while a query-focused path prompts a frozen VLM to reason over selected concepts and their observations. A learned conflict resolver selects between the two predictions when they disagree. Experiments on fake news benchmarks demonstrate the effectiveness of the method regarding detection performance and generalizability.

cs.MM↗