Image Pivoting for Learning Multilingual Multimodal Representations
In this paper we propose a model to learn multimodal multilingual representations for matching images and sentences in different languages, with the aim of advancing multilingual versions of image search and image understanding. Our model learns a common representation for images and their descriptions in two different languages (which need not be parallel) by considering the image as a pivot between two languages. We introduce a new pairwise ranking loss function which can handle both symmetric and asymmetric similarity between the two modalities. We evaluate our models on image-description ranking for German and English, and on semantic textual similarity of image descriptions in English. In both cases we achieve state-of-the-art performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Image DescriptionImage RetrievalSemantic Textual SimilaritySimilar Papers 제목 키워드 기반
Unsupervised Multimodal Neural Machine Translation with Pseudo Visual Pivoting
Unsupervised machine translation (MT) has recently achieved impressive results with monolingual corpora only. However, it is still challenging to associate source-target sentences in the latent space. As people speak dif…
Machine TranslationTranslationUnsupervised Machine TranslationComparing Multilingual NMT Models and Pivoting
Following recent advancements in multilingual machine translation at scale, our team carried out tests to compare the performance of multilingual models (M2M from Facebook and multilingual models from Helsinki-NLP) with …
Machine TranslationNMTTranslationZero-Shot Paraphrase Generation with Multilingual Language Models
Leveraging multilingual parallel texts to automatically generate paraphrases has drawn much attention as size of high-quality paraphrase corpus is limited. Round-trip translation, also known as the pivoting method, is a …
DenoisingDiversityMachine TranslationParaphrase Generation+2Assessing Multilingual Fairness in Pre-trained Multimodal Representations
Recently pre-trained multimodal models, such as CLIP, have shown exceptional capabilities towards connecting images and natural language. The textual representations in English can be desirably transferred to multilingua…
FairnessMultilingual Models for Compositional Distributed Semantics
We present a novel technique for learning semantic representations, which extends the distributional hypothesis to multilingual data and joint-space embeddings. Our models leverage parallel data and learn to strongly ali…
Cross-Lingual Document ClassificationDocument ClassificationGeneral ClassificationLearning Semantic Representations