paper-with-me

홈 › Papers

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment

2024-11-18 · Jiawei Li, Xinyue Liang, Junlong Zhang, Yizhe Yang, Chong Feng, Yang Gao

Process supervision enhances the performance of large language models in reasoning tasks by providing feedback at each step of chain-of-thought reasoning. However, due to the lack of effective process supervision methods, even advanced large language models are prone to logical errors and redundant reasoning. We claim that the effectiveness of process supervision significantly depends on both the accuracy and the length of reasoning chains. Moreover, we identify that these factors exhibit a nonlinear relationship with the overall reward score of the reasoning process. Inspired by these insights, we propose a novel process supervision paradigm, PSPO*, which systematically outlines the workflow from reward model training to policy optimization, and highlights the importance of nonlinear rewards in process supervision. Based on PSPO*, we develop the PSPO-WRS, which considers the number of reasoning steps in determining reward scores and utilizes an adjusted Weibull distribution for nonlinear reward shaping. Experimental results on six mathematical reasoning datasets demonstrate that PSPO-WRS consistently outperforms current mainstream models.

📄 PDF Abstract BibTeX arXiv:2411.11681

Code (1)

direct-bit/pspo 공식 구현 pytorch

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Offline Policy Optimization with Posterior Sampling

2026-05-08 · Hongqiang Lin, Dongxu Zhang, Yiding Sun, Mingzhe Li 외 arxiv

A fundamental challenge in model-based offline reinforcement learning (RL) lies in the trade-off between generalization and robustness against exploitation errors in out-of-distribution (OOD) regions. While OOD samples m…

Reinforcement LearningBayesian Inference

It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL

2025-09-25 · Madeleine Dwyer, Adam Sobey, Adriane Chapman arxiv

Training large language models (LLMs) with reinforcement learning (RL) methods such as PPO and GRPO commonly relies on ratio clipping to stabilise updates. While effective at preventing instability, clipping discards inf…

Reinforcement LearningMathematical Reasoning

The SkipSponge Attack: Sponge Weight Poisoning of Deep Neural Networks

2024-02-09 · Jona te Lintelo, Stefanos Koffas, Stjepan Picek

Sponge attacks aim to increase the energy consumption and computation time of neural networks. In this work, we present a novel sponge attack called SkipSponge. SkipSponge is the first sponge attack that is performed dir…

image-classificationImage Classification

PaperScout: An Autonomous Agent for Academic Paper Search with Process-Aware Sequence-Level Policy Optimization

2026-01-15 · Tingyue Pan, Jie Ouyang, Mingyue Cheng, Qingchuan Li 외 arxiv

Academic paper search is a fundamental task in scientific research, yet most existing approaches rely on rigid, predefined workflows that struggle with complex, conditional queries. To address this limitation, we propose…

Reinforcement Learning

Benchmark-Ready 3D Anatomical Shape Classification

2025-11-03 · Tomáš Krsička, Tibor Kubík arxiv

Progress in anatomical 3D shape classification is limited by the complexity of mesh data and the lack of standardized benchmarks, highlighting the need for robust learning methods and reproducible evaluation. We introduc…

3D Shape Classification