paper-with-me

Papers

MURAL: Multimodal, Multitask Retrieval Across Languages

2021-09-10 · Aashi Jain, Mandy Guo, Krishna Srinivasan, Ting Chen, Sneha Kudugunta, Chao Jia, Yinfei Yang, Jason Baldridge

Both image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages. We use both types of pairs in MURAL (MUltimodal, MUltitask Representations Across Languages), a dual encoder that solves two tasks: 1) image-text matching and 2) translation pair matching. By incorporating billions of translation pairs, MURAL extends ALIGN (Jia et al. PMLR'21)--a state-of-the-art dual encoder learned from 1.8 billion noisy image-text pairs. When using the same encoders, MURAL's performance matches or exceeds ALIGN's cross-modal retrieval performance on well-resourced languages across several datasets. More importantly, it considerably improves performance on under-resourced languages, showing that text-text learning can overcome a paucity of image-caption examples for these languages. On the Wikipedia Image-Text dataset, for example, MURAL-base improves zero-shot mean recall by 8.1% on average for eight under-resourced languages and by 6.8% on average when fine-tuning. We additionally show that MURAL's text representations cluster not only with respect to genealogical connections but also based on areal linguistics, such as the Balkan Sprachbund.

📄 PDF Abstract BibTeX arXiv:2109.05125

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalImage-text matchingRetrievalSemantic Image SimilaritySemantic Image-Text SimilaritySemantic Textual SimilarityText MatchingTranslation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

MURAL: Multimodal, Multitask Representations Across Languages

2021-11-01 · Findings (EMNLP) 2021 11 · Aashi Jain, Mandy Guo, Krishna Srinivasan, Ting Chen 외

Both image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages. We use both types of pairs in MURAL (MUltimodal, MUltitask Representations Across Langu…

Cross-Modal RetrievalImage-text matchingRetrievalText Matching+1

M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-training

2020-06-04 · CVPR 2021 1 · Minheng Ni, Haoyang Huang, Lin Su, Edward Cui 외

We present M3P, a Multitask Multilingual Multimodal Pre-trained model that combines multilingual pre-training and multimodal pre-training into a unified framework via multitask pre-training. Our goal is to learn universa…

Image CaptioningImage RetrievalMachine TranslationMultimodal Machine Translation+3

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution

2025-05-16 · Junyi Yuan, Jian Zhang, Fangyu Wu, Dongming Lu 외

China has a long and rich history, encompassing a vast cultural heritage that includes diverse multimodal information, such as silk patterns, Dunhuang murals, and their associated historical narratives. Cross-modal retri…

Cross-Modal RetrievalImage to textImage-to-Text RetrievalRetrieval+1

MURAL: Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning

2026-09-04 · Ahmad Mousavi, Majid Alikhani, Yeon-Chang Lee, Roberto Corizzo 외 arxiv

Multimodal Graph Neural Networks have become standard for recommendation by augmenting sparse interaction data with content features. Yet current architectures face two bottlenecks: structural rigidity, from a reliance o…

Multimodal Recommendation

A Multi-Granularity-Aware Aspect Learning Model for Multi-Aspect Dense Retrieval

2023-12-05 · Xiaojie Sun, Keping Bi, Jiafeng Guo, Sihui Yang 외

Dense retrieval methods have been mostly focused on unstructured text and less attention has been drawn to structured data with various aspects, e.g., products with aspects such as category and brand. Recent work has pro…

Language ModellingRetrievalValue prediction