paper-with-me

Papers

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

2025-06-04 · Ziming Cheng, Binrui Xu, Lisheng Gong, Zuhe Song, Tianshuo Zhou, Shiqi Zhong, Siyu Ren, Mingxiang Chen, Xiangchao Meng, Yuxin Zhang, Yanlin Li, Lei Ren, Wei Chen, Zhiyuan Huang, Mingjie Zhan, Xiaojie Wang, Fangxiang Feng

With enhanced capabilities and widespread applications, Multimodal Large Language Models (MLLMs) are increasingly required to process and reason over multiple images simultaneously. However, existing MLLM benchmarks focus either on single-image visual reasoning or on multi-image understanding tasks with only final-answer evaluation, leaving the reasoning capabilities of MLLMs over multi-image inputs largely underexplored. To address this gap, we introduce the $\textbf{Multimodal Multi-image Reasoning Benchmark (MMRB)}$, the first benchmark designed to evaluate structured visual reasoning across multiple images. MMRB comprises $\textbf{92 sub-tasks}$ covering spatial, temporal, and semantic reasoning, with multi-solution, CoT-style annotations generated by GPT-4o and refined by human experts. A derivative subset is designed to evaluate multimodal reward models in multi-image scenarios. To support fast and scalable evaluation, we propose a sentence-level matching framework using open-source LLMs. Extensive baseline experiments on $\textbf{40 MLLMs}$, including 9 reasoning-specific models and 8 reward models, demonstrate that open-source MLLMs still lag significantly behind commercial MLLMs in multi-image reasoning tasks. Furthermore, current multimodal reward models are nearly incapable of handling multi-image reward ranking tasks.

📄 PDF Abstract BibTeX arXiv:2506.04280

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceVisual Reasoning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs

2026-02-20 · Ziqiao Shang, Lingyue Ge, Zi-Jian Cheng, Shi-Yu Tian 외 arxiv

Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for rigorously measuring their reasoning capa…

Multimodal Reasoning

NPHardEval4V: A Dynamic Reasoning Benchmark of Multimodal Large Language Models

2024-03-04 · Lizhou Fan, Wenyue Hua, Xiang Li, Kaijie Zhu 외

Understanding the reasoning capabilities of Multimodal Large Language Models (MLLMs) is an important area of research. In this study, we introduce a dynamic benchmark, NPHardEval4V, aimed at addressing the existing gaps …

Instruction Following

MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning

2026-03-02 · Jiachun Li, Shaoping Huang, Zhuoran Jin, Chenlong Zhang 외 arxiv

Recent progress in the reasoning capabilities of multimodal large language models (MLLMs) has empowered them to address more complex tasks such as scientific analysis and mathematical reasoning. Despite their promise, ML…

Mathematical ReasoningMultimodal Reasoning

Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models

2025-03-03 · Boyu Jia, Junzhe Zhang, Huixuan Zhang, Xiaojun Wan

In recent years, multimodal large language models (MLLMs) have achieved significant breakthroughs, enhancing understanding across text and vision. However, current MLLMs still face challenges in effectively integrating k…

MathScape: Evaluating MLLMs in multimodal Math Scenarios through a Hierarchical Benchmark

2024-08-14 · Minxuan Zhou, Hao Liang, Tianpeng Li, Zhiyu Wu 외

With the development of Multimodal Large Language Models (MLLMs), the evaluation of multimodal models in the context of mathematical problems has become a valuable research field. Multimodal visual-textual mathematical r…

MathMathematical Reasoning