arXiv · 1812.01969
Summarizing Videos with Attention
Abstract
In this work we propose a novel method for supervised, keyshots based video summarization by applying a conceptually simple and computationally efficient soft, self-attention mechanism. Current state of the art methods leverage bi-directional recurrent networks such as BiLSTM combined with attention. These networks are complex to implement and computationally demanding compared to fully connected networks. To that end we propose a simple, self-attention based network for video summarization which performs the entire sequence to sequence transformation in a single feed forward pass and single backward pass during training. Our method sets a new state of the art results on two benchmarks TvSum and SumMe, commonly used in this domain.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiri Fajtl, Hajar Sadeghi Sokeh, Vasileios Argyriou, Dorothy Monekosso, Paolo Remagnino. 2019-02-21. Summarizing Videos with Attention. https://arxiv.org/abs/1812.01969
Cite the original work for its findings. Save a collection to share your selection of sources.