paper-with-me

Papers

Video-guided Machine Translation with Spatial Hierarchical Attention Network

2021-08-01 · ACL 2021 5 · Weiqi Gu, Haiyue Song, Chenhui Chu, Sadao Kurohashi

Video-guided machine translation, as one type of multimodal machine translations, aims to engage video contents as auxiliary information to address the word sense ambiguity problem in machine translation. Previous studies only use features from pretrained action detection models as motion representations of the video to solve the verb sense ambiguity, leaving the noun sense ambiguity a problem. To address this problem, we propose a video-guided machine translation system by using both spatial and motion representations in videos. For spatial features, we propose a hierarchical attention network to model the spatial information from object-level to video-level. Experiments on the VATEX dataset show that our system achieves 35.86 BLEU-4 score, which is 0.51 score higher than the single model of the SOTA method.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionMachine TranslationTranslationVideo-Guided Machine Translation

Similar Papers 제목 키워드 기반

Keyframe Segmentation and Positional Encoding for Video-guided Machine Translation Challenge 2020

2020-06-23 · Tosho Hirasawa, Zhishen Yang, Mamoru Komachi, Naoaki Okazaki

Video-guided machine translation as one of multimodal neural machine translation tasks targeting on generating high-quality text translation by tangibly engaging both video and text. In this work, we presented our video-…

Machine TranslationTranslationVideo-Guided Machine Translation

Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation

2023-06-13 · Shuai Yang, Yifan Zhou, Ziwei Liu, Chen Change Loy

Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains…

Patch MatchingTranslation

Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis

2025-07-09 · Hao Tang, Ling Shao, Zhenyu Zhang, Luc Van Gool 외 arxiv

We propose a novel spatial-temporal graph Mamba (STG-Mamba) for the music-guided dance video synthesis task, i.e., to translate the input music to a dance video. STG-Mamba consists of two translation mappings: music-to-s…

Recent Advances in Video Question Answering: A Review of Datasets and Methods

2021-01-15 · Devshree Patel, Ratnam Parikh, Yesha Shastri

Video Question Answering (VQA) is a recent emerging challenging task in the field of Computer Vision. Several visual information retrieval techniques like Video Captioning/Description and Video-guided Machine Translation…

Information RetrievalMachine TranslationQuestion AnsweringRetrieval+6

TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries

2025-05-09 · Jinze Lv, Jian Chen, Zi Long, Xianghua Fu 외

Most existing multimodal machine translation (MMT) datasets are predominantly composed of static images or short video clips, lacking extensive video data across diverse domains and topics. As a result, they fail to meet…

Domain AdaptationMachine TranslationMultimodal Machine TranslationNMT+1