Supervised Visual Attention for Multimodal Neural Machine Translation
This paper proposed a supervised visual attention mechanism for multimodal neural machine translation (MNMT), trained with constraints based on manual alignments between words in a sentence and their corresponding regions of an image. The proposed visual attention mechanism captures the relationship between a word and an image region more precisely than a conventional visual attention mechanism trained through MNMT in an unsupervised manner. Our experiments on English-German and German-English translation tasks using the Multi30k dataset and on English-Japanese and Japanese-English translation tasks using the Flickr30k Entities JP dataset show that a Transformer-based MNMT model can be improved by incorporating our proposed supervised visual attention mechanism and that further improvements can be achieved by combining it with a supervised cross-lingual attention mechanism (up to +1.61 BLEU, +1.7 METEOR).
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSentenceTranslationSimilar Papers 제목 키워드 기반
Supervised Visual Attention for Simultaneous Multimodal Machine Translation
Recently, there has been a surge in research in multimodal machine translation (MMT), where additional modalities such as images are used to improve translation quality of textual systems. A particular use for such multi…
Machine TranslationMultimodal Machine TranslationSentenceTranslationA Visual Attention Grounding Neural Model for Multimodal Machine Translation
We introduce a novel multimodal machine translation model that utilizes parallel visual and textual information. Our model jointly optimizes the learning of a shared visual-language embedding and a translator. The model …
Machine TranslationMultimodal Machine TranslationTranslationUnsupervised Multimodal Neural Machine Translation with Pseudo Visual Pivoting
Unsupervised machine translation (MT) has recently achieved impressive results with monolingual corpora only. However, it is still challenging to associate source-target sentences in the latent space. As people speak dif…
Machine TranslationTranslationUnsupervised Machine TranslationDouble Attention-based Multimodal Neural Machine Translation with Semantic Image Regions
Existing studies on multimodal neural machine translation (MNMT) have mainly focused on the effect of combining visual and textual modalities to improve translations. However, it has been suggested that the visual modali…
Machine TranslationTranslationDeeply Supervised Multimodal Attentional Translation Embeddings for Visual Relationship Detection
Detecting visual relationships, i.e. <Subject, Predicate, Object> triplets, is a challenging Scene Understanding task approached in the past via linguistic priors or spatial information in a single feature branch. We int…
Relationship DetectionScene UnderstandingTranslationVisual Relationship Detection