paper-with-me

홈 › Papers

MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification

2025-03-16 · Zhaopan Xu, Pengfei Zhou, Jiaxin Ai, Wangbo Zhao, Kai Wang, Xiaojiang Peng, Wenqi Shao, Hongxun Yao, Kaipeng Zhang

Reasoning is an essential capacity for large language models (LLMs) to address complex tasks, where the identification of process errors is vital for improving this ability. Recently, process-level reward models (PRMs) were proposed to provide step-wise rewards that facilitate reinforcement learning and data production during training and guide LLMs toward correct steps during inference, thereby improving reasoning accuracy. However, existing benchmarks of PRMs are text-based and focus on error detection, neglecting other scenarios like reasoning search. To address this gap, we introduce MPBench, a comprehensive, multi-task, multimodal benchmark designed to systematically assess the effectiveness of PRMs in diverse scenarios. MPBench employs three evaluation paradigms, each targeting a specific role of PRMs in the reasoning process: (1) Step Correctness, which assesses the correctness of each intermediate reasoning step; (2) Answer Aggregation, which aggregates multiple solutions and selects the best one; and (3) Reasoning Process Search, which guides the search for optimal reasoning steps during inference. Through these paradigms, MPBench makes comprehensive evaluations and provides insights into the development of multimodal PRMs.

📄 PDF Abstract BibTeX arXiv:2503.12505

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs

2024-07-23 · Jihyung Kil, Zheda Mai, Justin Lee, Zihe Wang 외

The ability to compare objects, scenes, or situations is crucial for effective decision-making and problem-solving in everyday life. For instance, comparing the freshness of apples enables better choices during grocery s…

Attribute

T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation

2023-07-12 · NeurIPS 2023 11 · Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li 외

Despite the stunning ability to generate high-quality images by recent text-to-image models, current approaches often struggle to effectively compose objects with different attributes and relationships into a complex and…

AttributeImage GenerationText to Image GenerationText-to-Image Generation

CompBench: Benchmarking Complex Instruction-guided Image Editing

2025-05-18 · Bohan Jia, Wenxuan Huang, Yuntian Tang, Junbo Qiao 외

While real-world applications increasingly demand intricate scene manipulation, existing instruction-guided image editing benchmarks often oversimplify task complexity and lack comprehensive, fine-grained instructions. T…

BenchmarkingInstruction Following

T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

2024-07-19 · CVPR 2025 1 · Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu 외

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also …

AttributeLanguage ModelingLanguage ModellingLarge Language Model+3

MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning

2025-09-26 · Yapeng Mi, Yanpeng Zhao, Hengli Li, Chenxi Li 외 arxiv

Reasoning-augmented machine learning systems have shown improved performance in various domains, including image generation. However, existing reasoning-based methods for image generation either restrict reasoning to a s…

Image Generation