paper-with-me

Papers

Generalizing Multimodal Pre-training into Multilingual via Language Acquisition

2022-05-29 · Liang Zhang, Anwen Hu, Qin Jin

English-based Vision-Language Pre-training (VLP) has achieved great success in various downstream tasks. Some efforts have been taken to generalize this success to non-English languages through Multilingual Vision-Language Pre-training (M-VLP). However, due to the large number of languages, M-VLP models often require huge computing resources and cannot be flexibly extended to new languages. In this work, we propose a \textbf{M}ulti\textbf{L}ingual \textbf{A}cquisition (MLA) framework that can easily generalize a monolingual Vision-Language Pre-training model into multilingual. Specifically, we design a lightweight language acquisition encoder based on state-of-the-art monolingual VLP models. We further propose a two-stage training strategy to optimize the language acquisition encoder, namely the Native Language Transfer stage and the Language Exposure stage. With much less multilingual training data and computing resources, our model achieves state-of-the-art performance on multilingual image-text and video-text retrieval benchmarks.

📄 PDF Abstract BibTeX arXiv:2206.11091

Code (0)

등록된 구현이 없습니다.

Tasks

Language AcquisitionRetrievalText RetrievalVideo-Text Retrieval

Similar Papers 제목 키워드 기반

M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-training

2020-06-04 · CVPR 2021 1 · Minheng Ni, Haoyang Huang, Lin Su, Edward Cui 외

We present M3P, a Multitask Multilingual Multimodal Pre-trained model that combines multilingual pre-training and multimodal pre-training into a unified framework via multitask pre-training. Our goal is to learn universa…

Image CaptioningImage RetrievalMachine TranslationMultimodal Machine Translation+3

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging

2026-05-13 · Zijing Wang, Mingyang Wang, Ercong Nie, Yongkang Liu 외 arxiv

Towards more general and human-like intelligence, large language models should seamlessly integrate both multilingual and multimodal capabilities; however, extending an existing multimodal model to many languages typical…

X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment

2024-03-18 · Dongjae Shin, HyeonSeok Lim, InHo Won, ChangSu Choi 외

The impressive development of large language models (LLMs) is expanding into the realm of large multimodal models (LMMs), which incorporate multiple types of data beyond text. However, the nature of multimodal models lea…

Diversity

CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning

2025-11-29 · Diego A. B. Moreira, Alef I. Ferreira, Jhessica Silva, Gabriel O. dos Santos 외 arxiv

As deep learning models evolve, new applications and challenges are rapidly emerging. Tasks that once relied on a single modality, such as text, images, or audio, are now enriched by seamless interactions between multimo…

Text Retrieval

Multilingual Multimodal Learning with Machine Translated Text

2022-10-24 · Chen Qiu, Dan Oneata, Emanuele Bugliarello, Stella Frank 외

Most vision-and-language pretraining research focuses on English tasks. However, the creation of multilingual multimodal evaluation datasets (e.g. Multi30K, xGQA, XVNLI, and MaRVL) poses a new challenge in finding high-q…

Zero-Shot Cross-Lingual Image-to-Text RetrievalZero-Shot Cross-Lingual Text-to-Image RetrievalZero-Shot Cross-Lingual Visual Natural Language InferenceZero-Shot Cross-Lingual Visual Question Answering+1