paper-with-me

홈 › Papers

PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

2025-01-06 · Mingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou, Yu Cheng

Process-level Reward Models (PRMs) are crucial for complex reasoning and decision-making tasks, where each intermediate step plays an important role in the reasoning process. Since language models are prone to various types of errors during the reasoning process, PRMs are required to possess nuanced capabilities for detecting various implicit error types in real-world scenarios. However, current benchmarks primarily focus on step correctness, failing to evaluate PRMs' performance systematically. To address this gap, we introduce PRMBench, a process-level benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs. PRMBench comprises 6,216 carefully designed problems and 83,456 step-level labels, evaluating models across multiple dimensions, including simplicity, soundness, and sensitivity. In our experiments on 15 models, spanning both open-source PRMs and closed-source large language models prompted as critic models, we uncover significant weaknesses in current PRMs. These findings underscore the challenges inherent in process-level evaluation and highlight key directions for future research. We hope PRMBench can be a robust bench for advancing research on PRM evaluation and development.

📄 PDF Abstract BibTeX arXiv:2501.03124

Code (1)

ssmisya/PRMBench 공식 구현 pytorch

Tasks

Decision Making

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents

2026-01-18 · Dawei Li, Yuguang Yao, Zhen Tan, Huan Liu 외 arxiv

Reward-guided search methods have demonstrated strong potential in enhancing tool-using agents by effectively guiding sampling and exploration over complex action spaces. As a core design, those search methods utilize pr…

MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning

2026-04-19 · Lingyan Wu, Xiang Zheng, Weiqi Zhai, Wei Wang 외 arxiv

Process-Level Reward Models (PRMs) are essential for guiding complex reasoning in large language models, yet existing PRM benchmarks cover only general domains such as mathematics, failing to address medical reasoning --…

Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns

2025-05-29 · Xiang Li, Haiyang Yu, Xinghua Zhang, Ziyang Huang 외

Process Reward Models (PRMs) are crucial in complex reasoning and problem-solving tasks (e.g., LLM agents with long-horizon decision-making) by verifying the correctness of each intermediate reasoning step. In real-world…

Benchmarking

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

2025-05-26 · Tej Deep Pala, Panshul Sharma, Amir Zadeh, Chuan Li 외

Large Language Models (LLMs) are prone to hallucination, especially during multi-hop and reasoning-intensive tasks such as mathematical problem solving. While Outcome Reward Models verify only final answers, Process Rewa…

HallucinationMathMathematical Problem-SolvingMathematical Reasoning+1

Efficient Process Reward Model Training via Active Learning

2025-04-14 · Keyu Duan, Zichen Liu, Xin Mao, Tianyu Pang 외

Process Reward Models (PRMs) provide step-level supervision to large language models (LLMs), but scaling up training data annotation remains challenging for both humans and LLMs. To address this limitation, we propose an…

Active LearningMath