Multimodal Pivots for Image Caption Translation
We present an approach to improve statistical machine translation of image descriptions by multimodal pivots defined in visual space. The key idea is to perform image retrieval over a database of images that are captioned in the target language, and use the captions of the most similar images for crosslingual reranking of translation outputs. Our approach does not depend on the availability of large amounts of in-domain parallel data, but only relies on available large datasets of monolingually captioned images, and on state-of-the-art convolutional neural networks to compute image similarities. Our experimental evaluation shows improvements of 1 BLEU point over strong baselines.
Code (0)
등록된 구현이 없습니다.
Tasks
Image RetrievalMachine TranslationRerankingRetrievalTranslationSimilar Papers 제목 키워드 기반
From Words to Sentences: A Progressive Learning Approach for Zero-resource Machine Translation with Visual Pivots
The neural machine translation model has suffered from the lack of large-scale parallel corpora. In contrast, we humans can learn multi-lingual translations even without parallel texts by referring our languages to the e…
Machine TranslationSentenceTranslationWord TranslationNLPHut’s Participation at WAT2021
This paper provides the description of shared tasks to the WAT 2021 by our team “NLPHut”. We have participated in the English→Hindi Multimodal translation task, English→Malayalam Multimodal translation task, and Indic Mu…
Caption GenerationImage CaptioningTranslationCUNI System for the WMT17 Multimodal Translation Task
In this paper, we describe our submissions to the WMT17 Multimodal Translation Task. For Task 1 (multimodal translation), our best scoring system is a purely textual neural translation of the source image caption to the …
Image CaptioningTask 2TranslationFeature-level Incongruence Reduction for Multimodal Translation
Caption translation aims to translate image annotations (captions for short). Recently, Multimodal Neural Machine Translation (MNMT) has been explored as the essential solution. Besides of linguistic features in captions…
Machine TranslationTranslationVisually Grounded Word Embeddings and Richer Visual Features for Improving Multimodal Neural Machine Translation
In Multimodal Neural Machine Translation (MNMT), a neural model generates a translated sentence that describes an image, given the image itself and one source descriptions in English. This is considered as the multimodal…
Dense CaptioningMachine Translationobject-detectionObject Detection+3