paper-with-me

홈 › Papers

Learning to Scale Multilingual Representations for Vision-Language Tasks

2020-04-09 · ECCV 2020 8 · Andrea Burns, Donghyun Kim, Derry Wijaya, Kate Saenko, Bryan A. Plummer

Current multilingual vision-language models either require a large number of additional parameters for each supported language, or suffer performance degradation as languages are added. In this paper, we propose a Scalable Multilingual Aligned Language Representation (SMALR) that supports many languages with few model parameters without sacrificing downstream task performance. SMALR learns a fixed size language-agnostic representation for most words in a multilingual vocabulary, keeping language-specific features for just a few. We use a masked cross-language modeling loss to align features with context from other languages. Additionally, we propose a cross-lingual consistency module that ensures predictions made for a query and its machine translation are comparable. The effectiveness of SMALR is demonstrated with ten diverse languages, over twice the number supported in vision-language tasks to date. We evaluate on multilingual image-sentence retrieval and outperform prior work by 3-4% with less than 1/5th the training parameters compared to other word embedding methods.

📄 PDF Abstract BibTeX arXiv:2004.04312

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMachine TranslationRetrievalSentenceSentence RetrievalTranslation

Similar Papers 제목 키워드 기반

uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data

2025-11-17 · Dahyun Chung, Donghyun Shin, Yujin Sung, Seunggi Moon 외 arxiv

Contrastive Language-Image Pre-training (CLIP) has demonstrated strong generalization across a wide range of visual tasks by leveraging large-scale English-image pairs. However, its extension to low-resource languages re…

Efficient Multilingual Multi-modal Pre-training through Triple Contrastive Loss

2022-10-01 · COLING 2022 10 · Youhan Lee, Kyungtae Lim, Woonhyuk Baek, Byungseok Roh 외

Learning visual and textual representations in the shared space from web-scale image-text pairs improves the performance of diverse vision-and-language tasks, as well as modality-specific tasks. Many attempts in this fra…

image-classificationImage ClassificationImage-text RetrievalRetrieval+7

On Zero-shot Cross-lingual Transfer of Multilingual Neural Machine Translation

2018-10-22 · Anonymous

Transferring representations from large-scale supervised tasks to downstream tasks have shown outstanding results in Machine Learning in both Computer Vision and natural language processing (NLP). One particular example …

Cross-Lingual TransferMachine TranslationNatural Language InferenceNMT+4

Assessing Multilingual Fairness in Pre-trained Multimodal Representations

2021-06-12 · Findings (ACL) 2022 5 · Jialu Wang, Yang Liu, Xin Eric Wang

Recently pre-trained multimodal models, such as CLIP, have shown exceptional capabilities towards connecting images and natural language. The textual representations in English can be desirably transferred to multilingua…

Fairness

An Empirical Recipe for Universal Phone Recognition

2026-03-30 · Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo 외 arxiv

Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive. Highly performant English-focused models do not generalize across languages, while…