paper-with-me

홈 › Papers

MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

2024-08-05 · Fanqing Meng, Jin Wang, Chuanhao Li, Quanfeng Lu, Hao Tian, Jiaqi Liao, Xizhou Zhu, Jifeng Dai, Yu Qiao, Ping Luo, Kaipeng Zhang, Wenqi Shao

The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their evaluation has not kept pace with their development. To fill this gap, we introduce the Multimodal Multi-image Understanding (MMIU) benchmark, a comprehensive evaluation suite designed to assess LVLMs across a wide range of multi-image tasks. MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of 24 popular LVLMs, including both open-source and proprietary models, reveals significant challenges in multi-image comprehension, particularly in tasks involving spatial understanding. Even the most advanced models, such as GPT-4o, achieve only 55.7% accuracy on MMIU. Through multi-faceted analytical experiments, we identify key performance gaps and limitations, providing valuable insights for future model and data improvements. We aim for MMIU to advance the frontier of LVLM research and development, moving us toward achieving sophisticated multimodal multi-image user interactions.

📄 PDF Abstract BibTeX arXiv:2408.02718

Code (1)

opengvlab/phygenbench pytorch

Tasks

Image ComprehensionMultiple-choice

Similar Papers 제목 키워드 기반

MMIU: Dataset for Visual Intent Understanding in Multimodal Assistants

2021-10-13 · Alkesh Patel, Joel Ruben Antony Moniz, Roman Nguyen, Nick Tzou 외

In multimodal assistant, where vision is also one of the input modalities, the identification of user intent becomes a challenging task as visual input can influence the outcome. Current digital assistants take spoken in…

intent-classificationIntent ClassificationQuestion AnsweringQuestion Generation+3

MMR: Evaluating Reading Ability of Large Multimodal Models

2024-08-26 · Jian Chen, Ruiyi Zhang, Yufan Zhou, Ryan Rossi 외

Large multimodal models (LMMs) have demonstrated impressive capabilities in understanding various types of image, including text-rich images. Most existing text-rich image benchmarks are simple extraction-based question …

Font RecognitionMMR totalOptical Character Recognition (OCR)Question Answering+2

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

2023-06-08 · NeurIPS 2023 11 · Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia 외

Despite the existence of various benchmarks for evaluating natural language processing models, we argue that human exams are a more suitable means of evaluating general intelligence for large language models (LLMs), as t…

PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

2026-07-30 · Zongyi Chen, Yu Liang, Jie Lin, Liansheng Wang arxiv

Multimodal large language models (MLLMs) are increasingly used to analyze pathology images. However, dominant multimodal benchmarks in pathology mainly score final diagnostic answers, captions, or reports. These evaluati…

Spatial Reasoning

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

2024-10-14 · Mu Cai, Reuben Tan, Jianrui Zhang, Bocheng Zou 외

Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benchmarks mostly resemble static image benc…

2kBenchmarkingQuestion AnsweringText Generation+2