Unsupervised Multimodal Neural Machine Translation with Pseudo Visual Pivoting
Unsupervised machine translation (MT) has recently achieved impressive results with monolingual corpora only. However, it is still challenging to associate source-target sentences in the latent space. As people speak different languages biologically share similar visual systems, the potential of achieving better alignment through visual content is promising yet under-explored in unsupervised multimodal MT (MMT). In this paper, we investigate how to utilize visual content for disambiguation and promoting latent space alignment in unsupervised MMT. Our model employs multimodal back-translation and features pseudo visual pivoting in which we learn a shared multilingual visual-semantic embedding space and incorporate visually-pivoted captioning as additional weak supervision. The experimental results on the widely used Multi30K dataset show that the proposed model significantly improves over the state-of-the-art methods and generalizes well when the images are not available at the testing time.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTranslationUnsupervised Machine TranslationSimilar Papers 제목 키워드 기반
Scene Graph as Pivoting: Inference-time Image-free Unsupervised Multimodal Machine Translation with Visual Scene Hallucination
In this work, we investigate a more realistic unsupervised multimodal machine translation (UMMT) setup, inference-time image-free UMMT, where the model is trained with source-text image pairs, and tested with only source…
HallucinationMachine TranslationMultimodal Machine TranslationTranslationFiltering Back-Translated Data in Unsupervised Neural Machine Translation
Unsupervised neural machine translation (NMT) utilizes only monolingual data for training. The quality of back-translated data plays an important role in the performance of NMT systems. In back-translation, all generated…
Domain AdaptationMachine TranslationNMTSentence+1Data Augmentation with Unsupervised Machine Translation Improves the Structural Similarity of Cross-lingual Word Embeddings
Unsupervised cross-lingual word embedding (CLWE) methods learn a linear transformation matrix that maps two monolingual embedding spaces that are separately trained with monolingual corpora. This method relies on the ass…
Cross-Lingual Word EmbeddingsData AugmentationMachine TranslationTranslation+2Extract and Edit: An Alternative to Back-Translation for Unsupervised Neural Machine Translation
The overreliance on large parallel corpora significantly limits the applicability of machine translation systems to the majority of language pairs. Back-translation has been dominantly used in previous approaches for uns…
Machine TranslationSentenceTranslationUnsupervised Machine TranslationSupervised Visual Attention for Multimodal Neural Machine Translation
This paper proposed a supervised visual attention mechanism for multimodal neural machine translation (MNMT), trained with constraints based on manual alignments between words in a sentence and their corresponding region…
Machine TranslationSentenceTranslation