arXiv · 2008.13362
Sentence Guided Temporal Modulation for Dynamic Video Thumbnail Generation
Abstract
We consider the problem of sentence specified dynamic video thumbnail generation. Given an input video and a user query sentence, the goal is to generate a video thumbnail that not only provides the preview of the video content, but also semantically corresponds to the sentence. In this paper, we propose a sentence guided temporal modulation (SGTM) mechanism that utilizes the sentence embedding to modulate the normalized temporal activations of the video thumbnail generation network. Unlike the existing state-of-the-art method that uses recurrent architectures, we propose a non-recurrent framework that is simple and allows much more parallelization. Extensive experiments and analysis on a large-scale dataset demonstrate the effectiveness of our framework.
Explore related subjects
Keep this discovery
Mrigank Rochan, Mahesh Kumar Krishna Reddy, Yang Wang. 2020-08-31. Sentence Guided Temporal Modulation for Dynamic Video Thumbnail Generation. https://arxiv.org/abs/2008.13362
Cite the original work for its findings. Save a collection to share your selection of sources.