paper-with-me

Papers

GEM: A General Evaluation Benchmark for Multimodal Tasks

2021-06-18 · Findings (ACL) 2021 8 · Lin Su, Nan Duan, Edward Cui, Lei Ji, Chenfei Wu, Huaishao Luo, Yongfei Liu, Ming Zhong, Taroon Bharti, Arun Sacheti

In this paper, we present GEM as a General Evaluation benchmark for Multimodal tasks. Different from existing datasets such as GLUE, SuperGLUE, XGLUE and XTREME that mainly focus on natural language tasks, GEM is a large-scale vision-language benchmark, which consists of GEM-I for image-language tasks and GEM-V for video-language tasks. Comparing with existing multimodal datasets such as MSCOCO and Flicker30K for image-language tasks, YouCook2 and MSR-VTT for video-language tasks, GEM is not only the largest vision-language dataset covering image-language tasks and video-language tasks at the same time, but also labeled in multiple languages. We also provide two baseline models for this benchmark. We will release the dataset, code and baseline models, aiming to advance the development of multilingual multimodal research.

📄 PDF Abstract BibTeX arXiv:2106.09889

Code (1)

microsoft/GEM 공식 구현

Similar Papers 제목 키워드 기반

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

2024-04-24 · Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li 외

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited…

UniIR: Training and Benchmarking Universal Multimodal Information Retrievers

2023-11-28 · Cong Wei, Yang Chen, Haonan Chen, Hexiang Hu 외

Existing information retrieval (IR) models often assume a homogeneous format, limiting their applicability to diverse user needs, such as searching for images with text descriptions, searching for a news article with a h…

BenchmarkingInformation RetrievalRetrievalZero-shot Generalization

Multimodal Evaluation of Russian-language Architectures

2025-11-19 · Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov, Artem Safin 외 arxiv

Multimodal large language models (MLLMs) are currently at the center of research attention, showing rapid progress in scale and capabilities, yet their intelligence, limitations, and risks remain insufficiently understoo…

MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

2023-08-04 · Weihao Yu, Zhengyuan Yang, Linjie Li, JianFeng Wang 외

We propose MM-Vet, an evaluation benchmark that examines large multimodal models (LMMs) on complicated multimodal tasks. Recent LMMs have shown various intriguing abilities, such as solving math problems written on the b…

MathMM-VetZero-Shot Visual Question Answring

MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains

2025-11-09 · Leyan Xue, Changqing Zhang, Kecheng Xue, Xiaohong Liu 외 arxiv

Although multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Current fusion methods are typically evaluated on a small selection of public …