paper-with-me

Papers

Multilingual Multimodal Learning with Machine Translated Text

2022-10-24 · Chen Qiu, Dan Oneata, Emanuele Bugliarello, Stella Frank, Desmond Elliott

Most vision-and-language pretraining research focuses on English tasks. However, the creation of multilingual multimodal evaluation datasets (e.g. Multi30K, xGQA, XVNLI, and MaRVL) poses a new challenge in finding high-quality training data that is both multilingual and multimodal. In this paper, we investigate whether machine translating English multimodal data can be an effective proxy for the lack of readily available multilingual data. We call this framework TD-MML: Translated Data for Multilingual Multimodal Learning, and it can be applied to any multimodal dataset and model. We apply it to both pretraining and fine-tuning data with a state-of-the-art model. In order to prevent models from learning from low-quality translated text, we propose two metrics for automatically removing such translations from the resulting datasets. In experiments on five tasks across 20 languages in the IGLUE benchmark, we show that translated data can provide a useful signal for multilingual multimodal learning, both at pretraining and fine-tuning.

📄 PDF Abstract BibTeX arXiv:2210.13134

Code (1)

danoneata/td-mml 공식 구현 pytorch

Tasks

Zero-Shot Cross-Lingual Image-to-Text RetrievalZero-Shot Cross-Lingual Text-to-Image RetrievalZero-Shot Cross-Lingual Visual Natural Language InferenceZero-Shot Cross-Lingual Visual Question AnsweringZero-Shot Cross-Lingual Visual Reasoning

Similar Papers 제목 키워드 기반

Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval

2024-10-02 · Kyle Buettner, Adriana Kovashka

There is a scarcity of multilingual vision-language models that properly account for the perceptual differences that are reflected in image captions across languages and cultures. In this work, through a multimodal, mult…

Image CaptioningRetrieval

Large-scale Bilingual Language-Image Contrastive Learning

2022-03-28 · Byungsoo Ko, Geonmo Gu

This paper is a technical report to share our experience and findings building a Korean and English bilingual multimodal model. While many of the multimodal datasets focus on English and multilingual multimodal research …

Contrastive LearningProper NounRelation

Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

2024-10-21 · Xiang Yue, Yueqi Song, Akari Asai, Seungone Kim 외

Despite recent advances in multimodal large language models (MLLMs), their development has predominantly focused on English- and western-centric datasets and tasks, leaving most of the world's languages and diverse cultu…

Resource Creation and Evaluation for Multilingual Sentiment Analysis in Social Media Texts

2014-05-01 · LREC 2014 5 · Alex Balahur, ra, Marco Turchi, Ralf Steinberger 외

This paper presents an evaluation of the use of machine translation to obtain and employ data for training multilingual sentiment classifiers. We show that the use of machine translated data obtained similar results as t…

ClassificationGeneral ClassificationMachine TranslationNatural Language Inference+4

Using Machine Translation to Augment Multilingual Classification

2024-05-09 · Adam King

An all-too-present bottleneck for text classification model development is the need to annotate training data and this need is multiplied for multilingual classifiers. Fortunately, contemporary machine translation models…

ClassificationImage CaptioningMachine Translationtext-classification+2