paper-with-me

Papers

GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models

2026-05-02 · Zhouhao Sun, Xuan Zhang, Xiao Ding, Bibo Cai, Li Du, Kai Xiong, Xinran Dai, Fei Zhang, weidi tang, Zhiyuan Kan, Yang Zhao, Bing Qin, Ting Liu arxiv

Currently, process reward models (PRMs) have exhibited remarkable potential for test-time scaling. Since large language models (LLMs) regularly generate flawed intermediate reasoning steps when tackling a broad spectrum of reasoning and decision-making tasks, PRMs are required to possess capabilities for detecting process-level errors in real-world scenarios. However, existing benchmarks primarily focus on mathematical reasoning, thereby failing to comprehensively evaluate the error detection ability of PRMs across diverse reasoning scenarios. To mitigate this gap, we introduce GR-Ben, a process-level benchmark specifically designed for assessing PRM's performance across two primary reasoning domains (science and logic) and nine subdomains. We conduct extensive experiments on a diverse set of 22 models, encompassing both PRMs and LLMs, and derive two key findings: (1) In domains beyond mathematical reasoning, the error-detection ability of existing PRMs and LLMs is found to be markedly weaker by comparison.(2) In general, PRMs are less adept at identifying knowledge-based errors, whereas LLMs exhibit poorer performance in detecting computational errors. We hope GR-Ben can foster future researches on PRMs for general domains, thereby enhancing the reasoning capabilities of LLMs.

📄 PDF Abstract BibTeX arXiv:2605.01203

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling

2024-12-19 · Zihan Liu, Yang Chen, Mohammad Shoeybi, Bryan Catanzaro 외

In this paper, we introduce AceMath, a suite of frontier math models that excel in solving complex math problems, along with highly effective reward models capable of evaluating generated solutions and reliably identifyi…

Math

Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

2025-02-20 · Michihiro Yasunaga, Luke Zettlemoyer, Marjan Ghazvininejad

Reward models play an essential role in training vision-language models (VLMs) by assessing output quality to enable aligning with human preferences. Despite their importance, the research community lacks comprehensive o…

Question AnsweringVisual Question Answering

A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

2025-10-09 · Congmin Zheng, Jiachen Zhu, Zhuoying Ou, Yuxiang Chen 외 arxiv

Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward models (ORMs) that judge only final answers. Process Reward Models(PRMs) addres…

Reinforcement LearningMultimodal Reasoning

Process Reward Agents for Steering Knowledge-Intensive Reasoning

2026-04-10 · Jiwoong Sohn, Tomasz Sternal, Kenneth Styppa, Torsten Hoefler 외 arxiv

Reasoning in knowledge-intensive domains remains challenging as intermediate steps are often not locally verifiable: unlike math or code, evaluating step correctness may require synthesizing clues across large external k…

RewardBench: Evaluating Reward Models for Language Modeling

2024-03-20 · Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda 외

Reward models (RMs) are at the crux of successfully using RLHF to align pretrained models to human preferences, yet there has been relatively little study that focuses on evaluation of those models. Evaluating reward mod…

Instruction FollowingLanguage ModelingLanguage Modelling