paper-with-me

홈 › Papers

Rethinking Entropy Regularization in Large Reasoning Models

2025-09-29 · Yuxian Jiang, Yafu Li, Guanxu Chen, Dongrui Liu, Yu Cheng, Jing Shao arxiv

Reinforcement learning with verifiable rewards (RLVR) has shown great promise in enhancing the reasoning abilities of large reasoning models (LRMs). However, it suffers from a critical issue: entropy collapse and premature convergence. Naive entropy regularization, a common approach for encouraging exploration in the traditional RL literature, fails to address this problem in the context of LRM. Our analysis reveals that this failure stems from the vast action space and long trajectories in LRMs, which easily trigger a global entropy explosion as the model indiscriminately explores all possible actions and states. To address this, we propose SIREN (SelectIve entRopy rEgularizatioN), a method that confines exploration to a meaningful subset of actions and states. SIREN achieves this through a two-step entropy masking mechanism, consisting of a top-p mask and a peak-entropy mask. In addition, regularization is transformed into a self-anchored form to stabilize training. Across five mathematical benchmarks, SIREN attains superior average performance over previous entropy-related RLVR approaches, exemplified by a +6.6 maj@k improvement on AIME24/25 with Qwen2.5-Math-7B. Further analysis confirms that SIREN promotes greater response diversity and maintains entropy at an appropriate level, which helps to preserve the validation pass@k throughout training. This effectively mitigates the premature convergence problem common in RLVR for LRM.

📄 PDF Abstract BibTeX arXiv:2509.25133

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning

2025-10-13 · Xiaoyun Zhang, Xiaojian Yuan, Di Huang, Wang You 외 arxiv

Reasoning ability has become a defining capability of Large Language Models (LLMs), with Reinforcement Learning with Verifiable Rewards (RLVR) emerging as a key paradigm to enhance it. However, RLVR training often suffer…

Reinforcement LearningMathematical Reasoning

A Comparative Theoretical Analysis of Entropy Control Methods in Reinforcement Learning

2026-04-02 · Ming Lei, Christophe Baehr arxiv

Reinforcement learning (RL) has become a key approach for enhancing reasoning in large language models (LLMs), yet scalable training is often hindered by the rapid collapse of policy entropy, which leads to premature con…

Reinforcement Learning

DSDR: Dual-Scale Diversity Regularization for Exploration in LLM Reasoning

2026-02-23 · Zhongwei Wan, Yun Shen, Zhihao Dou, Donghao Zhou 외 arxiv

Reinforcement learning with verifiers (RLVR) is a central paradigm for improving large language model (LLM) reasoning, yet existing methods often suffer from limited exploration. Policies tend to collapse onto a few reas…

Reinforcement Learning

PAEC: Position-Aware Entropy Calibration for LLM Reasoning in RLVR

2026-06-07 · Shumeng Yang, Yisu Liu, Jiayi Zheng, Zhaohui Yang 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning but often suffers from rapid policy-entropy collapse, where the policy prematurely concentrates on narrow high-probability rea…

Reinforcement LearningMathematical Reasoning

Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood

2026-04-14 · Xingyu Lin, Yilin Wen, Du Su, Jinchang Hou 외 arxiv

Group Relative Policy Optimization (GRPO) has significantly advanced the reasoning ability of large language models (LLMs), particularly in their mathemat ical reasoning performance. However, GRPO and related entropy reg…

Mathematical Reasoning