paper-with-me

홈 › Papers

BigVideo: A Large-scale Video Subtitle Translation Dataset for Multimodal Machine Translation

2023-05-23 · Liyan Kang, Luyang Huang, Ningxin Peng, Peihao Zhu, Zewei Sun, Shanbo Cheng, Mingxuan Wang, Degen Huang, Jinsong Su

We present a large-scale video subtitle translation dataset, BigVideo, to facilitate the study of multi-modality machine translation. Compared with the widely used How2 and VaTeX datasets, BigVideo is more than 10 times larger, consisting of 4.5 million sentence pairs and 9,981 hours of videos. We also introduce two deliberately designed test sets to verify the necessity of visual information: Ambiguous with the presence of ambiguous words, and Unambiguous in which the text context is self-contained for translation. To better model the common semantics shared across texts and videos, we introduce a contrastive learning method in the cross-modal encoder. Extensive experiments on the BigVideo show that: a) Visual information consistently improves the NMT model in terms of BLEU, BLEURT, and COMET on both Ambiguous and Unambiguous test sets. b) Visual information helps disambiguation, compared to the strong text baseline on terminology-targeted scores and human evaluation. Dataset and our implementations are available at https://github.com/DeepLearnXMU/BigVideo-VMT.

📄 PDF Abstract BibTeX arXiv:2305.18326

Code (1)

deeplearnxmu/bigvideo-vmt 공식 구현 pytorch

Tasks

Contrastive LearningMachine TranslationMultimodal Machine TranslationNMTSentenceTranslation

Methods 이 논문이 사용한 방법론

Test 설명 없음
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Video-guided Machine Translation with Global Video Context

2026-04-08 · Jian Chen, JinZe Lv, Zi Long, XiangHua Fu arxiv

Video-guided Multimodal Translation (VMT) has advanced significantly in recent years. However, most existing methods rely on locally aligned video segments paired one-to-one with subtitles, limiting their ability to capt…

Machine Translation

Video-Helpful Multimodal Machine Translation

2023-10-31 · Yihang Li, Shuichiro Shimizu, Chenhui Chu, Sadao Kurohashi 외

Existing multimodal machine translation (MMT) datasets consist of images and video captions or instructional video subtitles, which rarely contain linguistic ambiguity, making visual information ineffective in generating…

Machine TranslationMultimodal Machine TranslationTranslation

VISA: An Ambiguous Subtitles Dataset for Visual Scene-Aware Machine Translation

2022-01-20 · LREC 2022 6 · Yihang Li, Shuichiro Shimizu, Weiqi Gu, Chenhui Chu 외

Existing multimodal machine translation (MMT) datasets consist of images and video captions or general subtitles, which rarely contain linguistic ambiguity, making visual information not so effective to generate appropri…

Machine TranslationMultimodal Machine TranslationSentenceTranslation

HowToCaption: Prompting LLMs to Transform Video Annotations at Scale

2023-10-07 · Nina Shvetsova, Anna Kukleva, Xudong Hong, Christian Rupprecht 외

Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR) from the audio signal in the videos. H…

Automatic Speech RecognitionVideo CaptioningVideo RetrievalZero-Shot Video-Audio Retrieval+1

Generating Multilingual Parallel Corpus Using Subtitles

2018-04-11 · Farshad Jafari

Neural Machine Translation with its significant results, still has a great problem: lack or absence of parallel corpus for many languages. This article suggests a method for generating considerable amount of parallel cor…

Machine TranslationSentence