paper-with-me

IGLUE

Image-Grounded Language Understanding Evaluation

홈페이지 · 논문 31편

The Image-Grounded Language Understanding Evaluation (IGLUE) benchmark brings together—by both aggregating pre-existing datasets and creating new ones—visual question answering, cross-modal retrieval, grounded reasoning, and grounded entailment tasks across 20 diverse languages. The benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups.

ImagesTexts EnglishFrenchSpanishGermanChineseBengaliJapaneseRussianPortugueseArabicBulgarianDanishEstonianIndonesianKoreanTamilTurkishVietnameseGreekSwahili

벤치마크

Zero-Shot Cross-Lingual Transfer on MaRVL 결과 4개