paper-with-me

Papers

First Return, Entropy-Eliciting Explore

2025-07-09 · Tianyu Zheng, Tianshun Xing, Qingshui Gu, Taoran Liang, Xingwei Qu, Xin Zhou, Yizhi Li, Zhoufutu Wen, Chenghua Lin, Wenhao Huang, Qian Liu, Ge Zhang, Zejun Ma arxiv

Reinforcement Learning from Verifiable Rewards (RLVR) improves the reasoning abilities of Large Language Models (LLMs) but it struggles with unstable exploration. We propose FR3E (First Return, Entropy-Eliciting Explore), a structured exploration framework that identifies high-uncertainty decision points in reasoning trajectories and performs targeted rollouts to construct semantically grounded intermediate feedback. Our method provides targeted guidance without relying on dense supervision. Empirical results on mathematical reasoning benchmarks(AIME24) show that FR3E promotes more stable training, produces longer and more coherent responses, and increases the proportion of fully correct trajectories. These results highlight the framework's effectiveness in improving LLM reasoning through more robust and structured exploration.

📄 PDF Abstract BibTeX arXiv:2507.07017

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

Identifying and Estimating Perceived Returns to Binary Investments

2021-01-26 · Clint Harris

I describe a method for estimating agents' perceived returns to investments that relies on cross-sectional data containing binary choices and prices, where prices may be imperfectly known to agents. This method identifie…

Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces

2026-04-15 · Pranav Mahajan, Ben Seymour arxiv

Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a penalty on the divergence from a reference policy. Despite its success,…

Reinforcement Learning

Just Ask Them Twice: Choice Probabilities and Identification of Ex ante returns and Willingness-To-Pay

2023-03-06 · Romuald Meango, Esther Mirjam Girsberger

One of the exciting developments in the stated preference literature is the use of probabilistic stated preference experiments to estimate semi-parametric population distributions of ex ante returns and willingness-to-pa…

AttributecounterfactualSurvey

Action Redundancy in Reinforcement Learning

2021-02-22 · Nir Baram, Guy Tennenholtz, Shie Mannor

Maximum Entropy (MaxEnt) reinforcement learning is a powerful learning paradigm which seeks to maximize return under entropy regularization. However, action entropy does not necessarily coincide with state entropy, e.g.,…

MuJoCoreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Leveraging Sample Entropy for Enhanced Volatility Measurement and Prediction in International Oil Price Returns

2023-12-20 · Radhika Prosad Datta

This paper explores the application of Sample Entropy (SampEn) as a sophisticated tool for quantifying and predicting volatility in international oil price returns. SampEn, known for its ability to capture underlying pat…

Decision MakingregressionTime SeriesTime Series Regression