MURAL: Multimodal, Multitask Retrieval Across Languages
Both image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages. We use both types of pairs in MURAL (MUltimodal, MUltitask Representations Across Languages), a dual encoder that solves two tasks: 1) image-text matching and 2) translation pair matching. By incorporating billions of translation pairs, MURAL extends ALIGN (Jia et al. PMLR'21)--a state-of-the-art dual encoder learned from 1.8 billion noisy image-text pairs. When using the same encoders, MURAL's performance matches or exceeds ALIGN's cross-modal retrieval performance on well-resourced languages across several datasets. More importantly, it considerably improves performance on under-resourced languages, showing that text-text learning can overcome a paucity of image-caption examples for these languages. On the Wikipedia Image-Text dataset, for example, MURAL-base improves zero-shot mean recall by 8.1% on average for eight under-resourced languages and by 6.8% on average when fine-tuning. We additionally show that MURAL's text representations cluster not only with respect to genealogical connections but also based on areal linguistics, such as the Balkan Sprachbund.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Modal RetrievalImage-text matchingRetrievalSemantic Image SimilaritySemantic Image-Text SimilaritySemantic Textual SimilarityText MatchingTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
MURAL: Multimodal, Multitask Representations Across Languages
Both image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages. We use both types of pairs in MURAL (MUltimodal, MUltitask Representations Across Langu…
Cross-Modal RetrievalImage-text matchingRetrievalText Matching+1M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-training
We present M3P, a Multitask Multilingual Multimodal Pre-trained model that combines multilingual pre-training and multimodal pre-training into a unified framework via multitask pre-training. Our goal is to learn universa…
Image CaptioningImage RetrievalMachine TranslationMultimodal Machine Translation+3Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution
China has a long and rich history, encompassing a vast cultural heritage that includes diverse multimodal information, such as silk patterns, Dunhuang murals, and their associated historical narratives. Cross-modal retri…
Cross-Modal RetrievalImage to textImage-to-Text RetrievalRetrieval+1MURAL: Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning
Multimodal Graph Neural Networks have become standard for recommendation by augmenting sparse interaction data with content features. Yet current architectures face two bottlenecks: structural rigidity, from a reliance o…
Multimodal RecommendationA Multi-Granularity-Aware Aspect Learning Model for Multi-Aspect Dense Retrieval
Dense retrieval methods have been mostly focused on unstructured text and less attention has been drawn to structured data with various aspects, e.g., products with aspects such as category and brand. Recent work has pro…
Language ModellingRetrievalValue prediction