paper-with-me

Papers

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

2024-04-24 · Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li, Han Lin, Yue Yang, Hao Zhang, Wenbo Zhang, Yuqi Lin, Shuo Liu, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu, Renrui Zhang, Haozhe Zhang, Peng Gao, Yali Wang, Yu Qiao, Ping Luo, Kaipeng Zhang, Wenqi Shao

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in tracking LVLM development. In this study, we present MMT-Bench, a comprehensive benchmark designed to assess LVLMs across massive multimodal tasks requiring expert knowledge and deliberate visual recognition, localization, reasoning, and planning. MMT-Bench comprises $31,325$ meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering $32$ core meta-tasks and $162$ subtasks in multimodal understanding. Due to its extensive task coverage, MMT-Bench enables the evaluation of LVLMs using a task map, facilitating the discovery of in- and out-of-domain tasks. Evaluation results involving $30$ LVLMs such as the proprietary GPT-4V, GeminiProVision, and open-sourced InternVL-Chat, underscore the significant challenges posed by MMT-Bench. We anticipate that MMT-Bench will inspire the community to develop next-generation multimodal foundation models aimed at achieving general-purpose multimodal intelligence.

📄 PDF Abstract BibTeX arXiv:2404.16006

Code (1)

yangyue5114/DME pytorch

Similar Papers 제목 키워드 기반

LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs

2025-04-11 · Jiarui Wang, Huiyu Duan, Yu Zhao, Juntong Wang 외

Recent breakthroughs in large multimodal models (LMMs) have significantly advanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, many generated images still suffer from issues relat…

BenchmarkingImage GenerationImage to textQuestion Answering

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

2023-06-08 · NeurIPS 2023 11 · Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia 외

Despite the existence of various benchmarks for evaluating natural language processing models, we argue that human exams are a more suitable means of evaluating general intelligence for large language models (LLMs), as t…

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges

2024-11-28 · Rao Fu, Ziyang Luo, Hongzhan Lin, Zhen Ye 외

Recent advancements in large multimodal models (LMMs) have showcased impressive code generation capabilities, primarily evaluated through image-to-code benchmarks. However, these benchmarks are limited to specific visual…

Code Generation

MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

2023-11-15 · Fuxiao Liu, Xiaoyang Wang, Wenlin Yao, Jianshu Chen 외

With the rapid development of large language models (LLMs) and their integration into large multimodal models (LMMs), there has been impressive progress in zero-shot completion of user-oriented vision-language tasks. How…

Chart Understanding

MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation

2025-02-06 · Qinhan Yu, Zhiyou Xiao, Binghui Li, Zhengren Wang 외

Recent advances in Retrieval-Augmented Generation (RAG) have significantly improved response accuracy and relevance by incorporating external knowledge into Large Language Models (LLMs). However, existing RAG methods pri…

Answer Generationmultimodal generationRAGRetrieval+1