Video-guided Machine Translation with Spatial Hierarchical Attention Network
Video-guided machine translation, as one type of multimodal machine translations, aims to engage video contents as auxiliary information to address the word sense ambiguity problem in machine translation. Previous studies only use features from pretrained action detection models as motion representations of the video to solve the verb sense ambiguity, leaving the noun sense ambiguity a problem. To address this problem, we propose a video-guided machine translation system by using both spatial and motion representations in videos. For spatial features, we propose a hierarchical attention network to model the spatial information from object-level to video-level. Experiments on the VATEX dataset show that our system achieves 35.86 BLEU-4 score, which is 0.51 score higher than the single model of the SOTA method.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionMachine TranslationTranslationVideo-Guided Machine TranslationSimilar Papers 제목 키워드 기반
Keyframe Segmentation and Positional Encoding for Video-guided Machine Translation Challenge 2020
Video-guided machine translation as one of multimodal neural machine translation tasks targeting on generating high-quality text translation by tangibly engaging both video and text. In this work, we presented our video-…
Machine TranslationTranslationVideo-Guided Machine TranslationRerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains…
Patch MatchingTranslationSpatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
We propose a novel spatial-temporal graph Mamba (STG-Mamba) for the music-guided dance video synthesis task, i.e., to translate the input music to a dance video. STG-Mamba consists of two translation mappings: music-to-s…
Recent Advances in Video Question Answering: A Review of Datasets and Methods
Video Question Answering (VQA) is a recent emerging challenging task in the field of Computer Vision. Several visual information retrieval techniques like Video Captioning/Description and Video-guided Machine Translation…
Information RetrievalMachine TranslationQuestion AnsweringRetrieval+6TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries
Most existing multimodal machine translation (MMT) datasets are predominantly composed of static images or short video clips, lacking extensive video data across diverse domains and topics. As a result, they fail to meet…
Domain AdaptationMachine TranslationMultimodal Machine TranslationNMT+1