paper-with-me

홈 › Papers

IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and Languages

2022-01-27 · Emanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, Ivan Vulić

Reliable evaluation benchmarks designed for replicability and comprehensiveness have driven progress in machine learning. Due to the lack of a multilingual benchmark, however, vision-and-language research has mostly focused on English language tasks. To fill this gap, we introduce the Image-Grounded Language Understanding Evaluation benchmark. IGLUE brings together - by both aggregating pre-existing datasets and creating new ones - visual question answering, cross-modal retrieval, grounded reasoning, and grounded entailment tasks across 20 diverse languages. Our benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups. Based on the evaluation of the available state-of-the-art models, we find that translate-test transfer is superior to zero-shot transfer and that few-shot learning is hard to harness for many tasks. Moreover, downstream performance is partially explained by the amount of available unlabelled textual data for pretraining, and only weakly by the typological distance of target-source languages. We hope to encourage future research efforts in this area by releasing the benchmark to the community.

📄 PDF Abstract BibTeX arXiv:2201.11732

Code (3)

e-bug/iglue 공식 구현 pytorch
e-bug/volta 공식 구현 pytorch
shin-ee-chen/bla pytorch

Tasks

Cross-Modal RetrievalFew-Shot LearningImage-to-Text RetrievalQuestion AnsweringRetrievalTransfer LearningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models

2023-10-02 · Man Luo, Shrinidhi Kumbhar, Ming Shen, Mihir Parmar 외

Logical reasoning is fundamental for humans yet presents a substantial challenge in the domain of Artificial Intelligence. Initially, researchers used Knowledge Representation and Reasoning (KR) systems that did not scal…

Knowledge DistillationLanguage ModellingLogical Reasoning

OmniGlue: Generalizable Feature Matching with Foundation Model Guidance

2024-05-21 · CVPR 2024 1 · Hanwen Jiang, Arjun Karpur, Bingyi Cao, QiXing Huang 외

The image matching field has been witnessing a continuous emergence of novel learnable feature matching techniques, with ever-improving performance on conventional benchmarks. However, our investigation shows that despit…

ICU: Conquering Language Barriers in Vision-and-Language Modeling by Dividing the Tasks into Image Captioning and Language Understanding

2023-10-19 · Guojun Wu

Most multilingual vision-and-language (V&L) research aims to accomplish multilingual and multimodal capabilities within one model. However, the scarcity of multilingual captions for images has hindered the development. T…

Image CaptioningLanguage ModelingLanguage Modelling

Multilingual Multimodal Learning with Machine Translated Text

2022-10-24 · Chen Qiu, Dan Oneata, Emanuele Bugliarello, Stella Frank 외

Most vision-and-language pretraining research focuses on English tasks. However, the creation of multilingual multimodal evaluation datasets (e.g. Multi30K, xGQA, XVNLI, and MaRVL) poses a new challenge in finding high-q…

Zero-Shot Cross-Lingual Image-to-Text RetrievalZero-Shot Cross-Lingual Text-to-Image RetrievalZero-Shot Cross-Lingual Visual Natural Language InferenceZero-Shot Cross-Lingual Visual Question Answering+1

TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex

2026-07-24 · Yuliang Yan, Shuo Yan, Haochun Tang, Yiqin Sun 외 arxiv

Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation between an E3 ubiquitin ligase and a target protein. Despite their therapeutic potentia…