paper-with-me

홈 › Papers

ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs

2025-09-22 · Bonan Zhang, Zhongqi Chen, Bowen Song, Qinya Li, Fan Wu, Guihai Chen arxiv

Reinforcement learning (RL) has become a standard paradigm for refining large language models (LLMs) beyond pre-training and instruction tuning. A prominent line of work is RL with verifiable rewards (RLVR), which leverages automatically verifiable outcomes (e.g., correctness or executability) to generate reward signals. While efficient, this framework faces two key limitations: First, its binary feedback is too sparse to capture the quality of the reasoning process. Second, its coarse-grained rewards potentially lead to vanishing gradients. Inspired by observations from human learning, we introduce a RL technique that integrates verifiable outcomes with the model's own confidence estimates. This joint design enriches the reward signal, providing finer-grained feedback and implicitly supervising the reasoning process. Experimental results demonstrate that our proposed method enhances RL performance across multiple datasets and reduces token consumption during inference, while incurring negligible additional training cost. Moreover, it can be used as a plug-in module to enhance other state-of-the-art RL methods.

📄 PDF Abstract BibTeX arXiv:2509.17730

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards

2026-06-27 · Yupeng Chang, Yuan Wu, Yi Chang arxiv

Critic-free reinforcement learning with verifiable rewards (RLVR), exemplified by Group Relative Policy Optimization (GRPO), avoids training a value function (critic) and reduces memory and compute overhead relative to c…

Reinforcement Learning

Generalisation of RLHF under Reward Shift and Clipped KL Regularisation

2026-02-25 · Kenton Tang, Yuzhu Chen, Fengxiang He arxiv

Alignment and adaptation in large language models heavily rely on reinforcement learning from human feedback (RLHF); yet, theoretical understanding of its generalisability remains premature, especially when the learned r…

Reinforcement Learning

Process Supervision of Confidence Margin for Calibrated LLM Reasoning

2026-04-25 · Liaoyaqi Wang, Chunsheng Zuo, William Jurayj, Benjamin Van Durme 외 arxiv

Scaling test-time computation with reinforcement learning (RL) has emerged as a reliable path to improve large language models (LLM) reasoning ability. Yet, outcome-based reward often incentivizes models to be overconfid…

Reinforcement Learning

Learning values across many orders of magnitude

2016-02-24 · NeurIPS 2016 12 · Hado van Hasselt, Arthur Guez, Matteo Hessel, Volodymyr Mnih 외

Most learning algorithms are not invariant to the scale of the function that is being approximated. We propose to adaptively normalize the targets used in learning. This is useful in value-based reinforcement learning, w…

Atari Gamesreinforcement-learningReinforcement LearningReinforcement Learning (RL)

PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR

2026-08-11 · Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes arxiv

Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories. Recent allocators reduce this cost by assigning budgets to prompts, rollouts, or tokens a…

Reinforcement Learning