First Return, Entropy-Eliciting Explore
Reinforcement Learning from Verifiable Rewards (RLVR) improves the reasoning abilities of Large Language Models (LLMs) but it struggles with unstable exploration. We propose FR3E (First Return, Entropy-Eliciting Explore), a structured exploration framework that identifies high-uncertainty decision points in reasoning trajectories and performs targeted rollouts to construct semantically grounded intermediate feedback. Our method provides targeted guidance without relying on dense supervision. Empirical results on mathematical reasoning benchmarks(AIME24) show that FR3E promotes more stable training, produces longer and more coherent responses, and increases the proportion of fully correct trajectories. These results highlight the framework's effectiveness in improving LLM reasoning through more robust and structured exploration.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningMathematical ReasoningSimilar Papers 제목 키워드 기반
Identifying and Estimating Perceived Returns to Binary Investments
I describe a method for estimating agents' perceived returns to investments that relies on cross-sectional data containing binary choices and prices, where prices may be imperfectly known to agents. This method identifie…
Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces
Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a penalty on the divergence from a reference policy. Despite its success,…
Reinforcement LearningJust Ask Them Twice: Choice Probabilities and Identification of Ex ante returns and Willingness-To-Pay
One of the exciting developments in the stated preference literature is the use of probabilistic stated preference experiments to estimate semi-parametric population distributions of ex ante returns and willingness-to-pa…
AttributecounterfactualSurveyAction Redundancy in Reinforcement Learning
Maximum Entropy (MaxEnt) reinforcement learning is a powerful learning paradigm which seeks to maximize return under entropy regularization. However, action entropy does not necessarily coincide with state entropy, e.g.,…
MuJoCoreinforcement-learningReinforcement LearningReinforcement Learning (RL)Leveraging Sample Entropy for Enhanced Volatility Measurement and Prediction in International Oil Price Returns
This paper explores the application of Sample Entropy (SampEn) as a sophisticated tool for quantifying and predicting volatility in international oil price returns. SampEn, known for its ability to capture underlying pat…
Decision MakingregressionTime SeriesTime Series Regression