paper-with-me

홈 › Papers

ProcessBench: Identifying Process Errors in Mathematical Reasoning

2024-12-09 · Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, Junyang Lin

As language models regularly make mistakes when solving math problems, automated identification of errors in the reasoning process becomes increasingly significant for their scalable oversight. In this paper, we introduce ProcessBench for measuring the ability to identify erroneous steps in mathematical reasoning. It consists of 3,400 test cases, primarily focused on competition- and Olympiad-level math problems. Each test case contains a step-by-step solution with error location annotated by human experts. Models are required to identify the earliest step that contains an error, or conclude that all steps are correct. We conduct extensive evaluation on ProcessBench, involving two types of models: process reward models (PRMs) and critic models, where for the latter we prompt general language models to critique each solution step by step. We draw two main observations: (1) Existing PRMs typically fail to generalize to more challenging math problems beyond GSM8K and MATH. They underperform both critic models (i.e., prompted general language models) and our own trained PRM that is straightforwardly fine-tuned on the PRM800K dataset. (2) The best open-source model, QwQ-32B-Preview, has demonstrated the critique capability competitive with the proprietary model GPT-4o, despite that it still lags behind the reasoning-specialized o1-mini. We hope ProcessBench can foster future research in reasoning process assessment, paving the way toward scalable oversight of language models.

📄 PDF Abstract BibTeX arXiv:2412.06559

Code (1)

qwenlm/processbench 공식 구현 pytorch

Tasks

GSM8KMathMathematical Reasoning

Similar Papers 제목 키워드 기반

AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

2026-03-15 · Shengda Fan, Xuyan Ye, Yupeng Huo, Zhi-Yuan Chen 외 arxiv

While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reasoning where errors are often rectifiable via backtracking, tool-use failur…

Mathematical Reasoning

Temporal Consistency for LLM Reasoning Process Error Identification

2025-03-18 · Jiacheng Guo, Yue Wu, Jiahao Qiu, Kaixuan Huang 외

Verification is crucial for effective mathematical reasoning. We present a new temporal consistency method where verifiers iteratively refine their judgments based on the previous assessment. Unlike one-round verificatio…

Mathematical Reasoning

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

2025-04-27 · Jiaqi Chen, Bang Zhang, Ruotian Ma, Peisong Wang 외

Evaluating the step-by-step reliability of large language model (LLM) reasoning, such as Chain-of-Thought, remains challenging due to the difficulty and cost of obtaining high-quality step-level supervision. In this pape…

Large Language ModelMathematical Reasoning

Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical Reasoning

2025-08-03 · Jiuzhou Han, Wray Buntine, Ehsan Shareghi arxiv

Large language models have demonstrated remarkable capabilities in complex mathematical reasoning tasks, but they inevitably generate errors throughout multi-step solutions. Process-level Reward Models (PRMs) have shown …

Mathematical Reasoning

ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth

2025-12-02 · Salman Rahman, Sruthi Gorantla, Arpit Gupta, Swastik Roy 외 arxiv

Training process reward models (PRMs) requires step-level correctness labels, obtained either through expensive human annotation or by relying on ground-truth answers, limiting the ability to scale process-level supervis…

Reinforcement LearningMathematical Reasoning