paper-with-me

홈 › Papers

Boosting LLM Reasoning via Human-Inspired Reward Shaping

2026-02-04 · Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma, Gao Huang arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for enhancing reasoning in Large Language Models (LLMs). However, existing reward formulations typically treat exploration and consolidation as a monolithic process, resulting in entangled stage-wise learning dynamics. This contradicts the natural learning behavior of human learners. In human learning, individuals adopt distinct behavioral patterns toward mastered versus unfamiliar problems. When confronting unmastered challenges, humans prioritize broad exploration to seek viable solutions. By contrast, for well-mastered problems, they focus instead on reasoning condensation and knowledge abstraction to distill concise underlying principles. Motivated by this gap, we introduce T2T(Thickening-to-Thinning), a dynamic reward framework inspired by human learning processes. Specifically, it implements a dual-phase mechanism: (1) On incorrect attempts, T2T incentivizes "thickening" to broaden the search space and explore novel solution paths; (2) Upon achieving correctness, it shifts to "thinning", imposing length penalties to discourage redundancy, thereby fostering model confidence and crystallizing reasoning capabilities. Extensive experiments on mathematical benchmarks (MATH-500, AIME, AMC) across 5 mainstream LLMs demonstrate that T2T significantly outperforms standard GRPO and recent baselines, achieving superior performance.

📄 PDF Abstract BibTeX arXiv:2602.04265

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

A Narration-based Reward Shaping Approach using Grounded Natural Language Commands

2019-10-31 · Nicholas Waytowich, Sean L. Barton, Vernon Lawhern, Garrett Warnell

While deep reinforcement learning techniques have led to agents that are successfully able to learn to perform a number of tasks that had been previously unlearnable, these techniques are still susceptible to the longsta…

Deep Reinforcement LearningReinforcement LearningStarcraftStarcraft II

Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping

2020-11-05 · NeurIPS 2020 12 · Yujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 외

Reward shaping is an effective technique for incorporating domain knowledge into reinforcement learning (RL). Existing approaches such as potential-based reward shaping normally make full use of a given shaping reward fu…

MuJoCoReinforcement Learning (RL)

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment

2024-11-18 · Jiawei Li, Xinyue Liang, Junlong Zhang, Yizhe Yang 외

Process supervision enhances the performance of large language models in reasoning tasks by providing feedback at each step of chain-of-thought reasoning. However, due to the lack of effective process supervision methods…

Mathematical Reasoning

Learning from Expert Factors: Trajectory-level Reward Shaping for Formulaic Alpha Mining

2025-07-27 · Junjie Zhao, Chengxi Zhang, Chenkai Wang, Peng Yang arxiv

Reinforcement learning (RL) has successfully automated the complex process of mining formulaic alpha factors, for creating interpretable and profitable investment strategies. However, existing methods are hampered by the…

Computational EfficiencyReinforcement Learning

Magnetic Field-Based Reward Shaping for Goal-Conditioned Reinforcement Learning

2023-07-16 · Hongyu Ding, Yuanze Tang, Qing Wu, Bo wang 외

Goal-conditioned reinforcement learning (RL) is an interesting extension of the traditional RL framework, where the dynamic environment and reward sparsity can cause conventional learning algorithms to fail. Reward shapi…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)