paper-with-me

홈 › Papers

Multilingual Multimodal Pretraining for Zero-Shot Cross-Lingual Transfer of Vision-Language Models

2020-12-07 · Anonymous

This paper studies zero-shot cross-lingual transfer of vision-language models. Specifically, we focus on multilingual text-to-video search and propose a Transformer-based model that learns contextualized multilingual multimodal embeddings. Under a zero-shot setting, we empirically demonstrate that performance degrades significantly when we query the multilingual text-video model with non-English sentences. To address this problem, we introduce a multilingual multimodal pretraining strategy, and collect a new multilingual instructional video dataset (Multi-HowTo100M) for pretraining. Experiments on VTT show that our method significantly improves video search in non-English languages without additional annotations. Furthermore, when multilingual annotations are available, our method outperforms recent baselines by a large margin in multilingual text-to-video search on VTT and VATEX; as well as in multilingual text-to-image search on Multi30K. Our model and Multi-HowTo100M will be made available.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual TransferImage RetrievalText-to-video searchZero-Shot Cross-Lingual Transfer

Similar Papers 제목 키워드 기반

Match the Script, Adapt if Multilingual: Analyzing the Effect of Multilingual Pretraining on Cross-lingual Transferability

2022-03-21 · ACL 2022 5 · Yoshinari Fujinuma, Jordan Boyd-Graber, Katharina Kann

Pretrained multilingual models enable zero-shot learning even for unseen languages, and that performance can be further improved via adaptation prior to finetuning. However, it is unclear how the number of pretraining la…

Zero-Shot Learning

Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon

2024-02-03 · Fajri Koto, Tilman Beck, Zeerak Talat, Iryna Gurevych 외

Improving multilingual language models capabilities in low-resource languages is generally difficult due to the scarcity of large-scale data in those languages. In this paper, we relax the reliance on texts in low-resour…

SentenceSentiment Analysis

No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance

2024-04-04 · Vishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma 외

Web-crawled pretraining datasets underlie the impressive "zero-shot" evaluation performance of multimodal models, such as CLIP for classification/retrieval and Stable-Diffusion for image generation. However, it is unclea…

BenchmarkingImage GenerationZero-shot Generalization

AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource Languages

2021-04-18 · ACL 2022 5 · Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary 외

Pretrained multilingual models are able to perform cross-lingual transfer in a zero-shot setting, even for languages unseen during pretraining. However, prior work evaluating performance on unseen languages has largely b…

Cross-Lingual TransferNatural Language UnderstandingTranslationXLM-R+1

Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models

2021-03-16 · NAACL 2021 4 · Po-Yao Huang, Mandela Patrick, Junjie Hu, Graham Neubig 외

This paper studies zero-shot cross-lingual transfer of vision-language models. Specifically, we focus on multilingual text-to-video search and propose a Transformer-based model that learns contextualized multilingual mul…

Cross-Lingual TransferImage RetrievalText-to-video searchZero-Shot Cross-Lingual Transfer