paper-with-me

Papers

Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets. In this paper we present and make available the Crossmodal-3600 dataset, a geographically diverse set of 3600 images each of them annotated with human-generated reference captions in 36 languages. We select a representative set of images from across the world for this dataset, and annotate it with captions that achieve consistency in terms of style across all languages, while avoiding annotation artifacts due to direct translation. We apply this benchmark to model selection for massively multilingual image captioning models, and show superior correlation results with human evaluations when using the Crossmodal-3600 dataset as golden references for automatic metrics.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningModel SelectionTranslation

Similar Papers 제목 키워드 기반

Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset

2022-05-25 · Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu Soricut

Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets. In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically diver…

Image CaptioningImage RetrievalImage-text RetrievalImage-to-Text Retrieval+3

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

2024-12-11 · Andreas Koukounas, Georgios Mastrapas, Sedigheh Eslami, Bo wang 외

Contrastive Language-Image Pretraining (CLIP) has been widely used for crossmodal information retrieval and multimodal understanding tasks. However, CLIP models are mainly optimized for crossmodal vision-language tasks a…

Contrastive LearningCross-Modal Information RetrievalInformation RetrievalRepresentation Learning+3

Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation

2025-04-09 · Israfel Salazar, Manuel Fernández Burda, Shayekh Bin Islam, Arshia Soltani Moakhar 외

The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage. While multilingual benchmarks have expanded, bot…

Multiple-choice

Unimodal and Crossmodal Refinement Network for Multimodal Sequence Fusion

2021-11-01 · EMNLP 2021 11 · Xiaobao Guo, Adams Kong, Huan Zhou, Xianfeng Wang 외

Effective unimodal representation and complementary crossmodal representation fusion are both important in multimodal representation learning. Prior works often modulate one modal feature to another straightforwardly and…

Representation Learning

The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge Distillation

2022-06-13 · Zihui Xue, Zhengqi Gao, Sucheng Ren, Hang Zhao

Crossmodal knowledge distillation (KD) extends traditional knowledge distillation to the area of multimodal learning and demonstrates great success in various applications. To achieve knowledge transfer across modalities…

Knowledge DistillationTransfer Learning