paper-with-me

Papers

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

2024-09-04 · Xiang Yue, Tianyu Zheng, Yuansheng Ni, YuBo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig

This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. MMMU-Pro rigorously assesses multimodal models' true understanding and reasoning capabilities through a three-step process based on MMMU: (1) filtering out questions answerable by text-only models, (2) augmenting candidate options, and (3) introducing a vision-only input setting where questions are embedded within images. This setting challenges AI to truly "see" and "read" simultaneously, testing a fundamental human cognitive skill of seamlessly integrating visual and textual information. Results show that model performance is substantially lower on MMMU-Pro than on MMMU, ranging from 16.8% to 26.9% across models. We explore the impact of OCR prompts and Chain of Thought (CoT) reasoning, finding that OCR prompts have minimal effect while CoT generally improves performance. MMMU-Pro provides a more rigorous evaluation tool, closely mimicking real-world scenarios and offering valuable directions for future research in multimodal AI.

📄 PDF Abstract BibTeX arXiv:2409.02813

Code (2)

MMMU-Benchmark/MMMU
opendatalab/pm4bench

Tasks

Optical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context

2026-03-18 · Nahyun Lee, Guijin Son, Hyunwoo Ko, Chanyoung Kim 외 arxiv

We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,466 questions from exams natively written in Korean, covering nine dis…

CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark

2024-01-22 · Ge Zhang, Xinrun Du, Bei Chen, Yiming Liang 외

As the capabilities of large multimodal models (LMMs) continue to advance, evaluating the performance of LMMs emerges as an increasing need. Additionally, there is an even larger gap in evaluating the advanced knowledge …

JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction

2025-12-16 · Atsuyuki Miyai, Shota Onohara, Jeonghun Baek, Kiyoharu Aizawa arxiv

This paper introduces JMMMU-Pro, an image-based Japanese Multi-discipline Multimodal Understanding Benchmark, and Vibe Benchmark Construction, a scalable construction method. Following the evolution from MMMU to MMMU-Pro…

Image Generation

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

2023-11-27 · CVPR 2024 1 · Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng 외

We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected m…

Complex Query AnsweringLogical ReasoningVisual Reasoning

Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark

2025-10-15 · Kai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian 외 arxiv

Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. Existing evaluations either treat the two abilities in isolation or overl…