paper-with-me

홈 › Papers

Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search

2025-01-02 · Shuangtao Li, Shuaihao Dong, Kexin Luan, Xinhan Di, Chaofan Ding

Large language models (LLMs) have demonstrated their remarkable capacity across a variety of tasks. However, reasoning remains a challenge for LLMs. To improve LLMs' reasoning ability, process supervision has proven to be better than outcome supervision. In this work, we study using Monte Carlo Tree Search (MCTS) to generate process supervision data with LLMs themselves for training them. We sample reasoning steps with an LLM and assign each step a score that captures its "relative correctness," and the LLM is then trained by minimizing weighted log-likelihood of generating the reasoning steps. This generate-then-train process is repeated iteratively until convergence.Our experimental results demonstrate that the proposed methods considerably improve the performance of LLMs on two mathematical reasoning datasets. Furthermore, models trained on one dataset also exhibit improved performance on the other, showing the transferability of the enhanced reasoning ability.

📄 PDF Abstract BibTeX arXiv:2501.01478

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Improve Mathematical Reasoning in Language Models by Automated Process Supervision

2024-06-05 · Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale 외

Complex multi-step reasoning tasks, such as solving mathematical problems or generating code, remain a significant hurdle for even the most advanced large language models (LLMs). Verifying LLM outputs with an Outcome Rew…

GSM8KMathMathematical Reasoning

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision

2025-05-19 · Lingxiao Du, Fanqing Meng, Zongkai Liu, Zhixiang Zhou 외

While Multimodal Large Language Models (MLLMs) have achieved impressive progress in vision-language understanding, they still struggle with complex multi-step reasoning, often producing logically inconsistent or partiall…

MathMathematical ReasoningMultimodal Reasoning

GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning

2026-04-22 · Jingyi Wang, Lei Zhu, Tengjin Weng, Song-Li Wu 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Language Models (LLMs) by leveraging direct outcome verification instead of learned reward models. Building on this p…

Reinforcement Learning

An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning

2025-03-04 · Wei Sun, Qianlong Du, Fuwei Cui, Jiajun Zhang

Enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) is of great scientific and practical significance. Researchers typically employ process-supervised reward models (PRMs) to guide the reaso…

Mathematical Reasoning

From Static to Dynamic: Adaptive Monte Carlo Search for Mathematical Process Supervision

2025-09-29 · Jie Ma, Shihao Qi, Rui Xing, Ziang Yin 외 arxiv

The quality of process data plays a key role in training a Process Reward Model (PRM), which can enhance the complex mathematical reasoning capability of large language models. Existing methods estimate the quality of re…

Mathematical Reasoning