arXiv · 2006.12799
Keyframe Segmentation and Positional Encoding for Video-guided Machine Translation Challenge 2020
Abstract
Video-guided machine translation as one of multimodal neural machine translation tasks targeting on generating high-quality text translation by tangibly engaging both video and text. In this work, we presented our video-guided machine translation system in approaching the Video-guided Machine Translation Challenge 2020. This system employs keyframe-based video feature extractions along with the video feature positional encoding. In the evaluation phase, our system scored 36.60 corpus-level BLEU-4 and achieved the 1st place on the Video-guided Machine Translation Challenge 2020.
Explore related subjects
Keep this discovery
Tosho Hirasawa, Zhishen Yang, Mamoru Komachi, Naoaki Okazaki. 2020-06-23. Keyframe Segmentation and Positional Encoding for Video-guided Machine Translation Challenge 2020. https://arxiv.org/abs/2006.12799
Cite the original work for its findings. Save a collection to share your selection of sources.