MURAL: Multimodal, Multitask Representations Across Languages
Both image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages. We use both types of pairs in MURAL (MUltimodal, MUltitask Representations Across Languages), a dual encoder that solves two tasks: 1) image-text matching and 2) translation pair matching. By incorporating billions of translation pairs, MURAL extends ALIGN (Jia et al.)–a state-of-the-art dual encoder learned from 1.8 billion noisy image-text pairs. When using the same encoders, MURAL’s performance matches or exceeds ALIGN’s cross-modal retrieval performance on well-resourced languages across several datasets. More importantly, it considerably improves performance on under-resourced languages, showing that text-text learning can overcome a paucity of image-caption examples for these languages. On the Wikipedia Image-Text dataset, for example, MURAL-base improves zero-shot mean recall by 8.1% on average for eight under-resourced languages and by 6.8% on average when fine-tuning. We additionally show that MURAL’s text representations cluster not only with respect to genealogical connections but also based on areal linguistics, such as the Balkan Sprachbund.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Modal RetrievalImage-text matchingRetrievalText MatchingTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
MURAL: Multimodal, Multitask Retrieval Across Languages
Both image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages. We use both types of pairs in MURAL (MUltimodal, MUltitask Representations Across Langu…
Cross-Modal RetrievalImage-text matchingRetrievalSemantic Image Similarity+4M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-training
We present M3P, a Multitask Multilingual Multimodal Pre-trained model that combines multilingual pre-training and multimodal pre-training into a unified framework via multitask pre-training. Our goal is to learn universa…
Image CaptioningImage RetrievalMachine TranslationMultimodal Machine Translation+3MURAL: Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning
Multimodal Graph Neural Networks have become standard for recommendation by augmenting sparse interaction data with content features. Yet current architectures face two bottlenecks: structural rigidity, from a reliance o…
Multimodal RecommendationBuilding Bridge Across the Time: Disruption and Restoration of Murals In the Wild
In this paper, we focus on the mural-restoration task, which aims to detect damaged regions in the mural and repaint them automatically. Different from traditional image restoration tasks like in/out/blind-painting a…
DiversityImage RestorationEmbodied Multimodal Multitask Learning
Recent efforts on training visual navigation agents conditioned on language using deep reinforcement learning have been successful in learning policies for different multimodal tasks, such as semantic goal navigation and…
Deep Reinforcement LearningDisentanglementEmbodied Question AnsweringQuestion Answering+3