paper-with-me

홈 › Papers

Model-free Policy Learning with Reward Gradients

2021-03-09 · Qingfeng Lan, Samuele Tosatto, Homayoon Farrahi, A. Rupam Mahmood

Despite the increasing popularity of policy gradient methods, they are yet to be widely utilized in sample-scarce applications, such as robotics. The sample efficiency could be improved by making best usage of available information. As a key component in reinforcement learning, the reward function is usually devised carefully to guide the agent. Hence, the reward function is usually known, allowing access to not only scalar reward signals but also reward gradients. To benefit from reward gradients, previous works require the knowledge of environment dynamics, which are hard to obtain. In this work, we develop the \textit{Reward Policy Gradient} estimator, a novel approach that integrates reward gradients without learning a model. Bypassing the model dynamics allows our estimator to achieve a better bias-variance trade-off, which results in a higher sample efficiency, as shown in the empirical analysis. Our method also boosts the performance of Proximal Policy Optimization on different MuJoCo control tasks.

📄 PDF Abstract BibTeX arXiv:2103.05147

Code (1)

qlan3/Explorer 공식 구현 pytorch

Tasks

Continuous ControlmodelMuJoCoPolicy Gradient Methods

Similar Papers 제목 키워드 기반

Efficient Baseline-free Sampling in Parameter Exploring Policy Gradients: Super Symmetric PGPE

2013-12-13 · Frank Sehnke

Policy Gradient methods that explore directly in parameter space are among the most effective and robust direct policy search methods and have drawn a lot of attention lately. The basic method from this field, Policy Gra…

Policy Gradient Methods

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

2026-08-04 · Yongshi Ye, Liang Zhang, Yidong Chen, Xiaodong Shi 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-leve…

Reinforcement Learning

Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward

2026-05-11 · Xuexiang Wen, Hang Yu, Linchao Zhu, Gaoang Wang arxiv

While Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a promising post-training paradigm for Large Language Models (LLMs), its dependency on the gold label or domain-specific verifiers limit…

Reinforcement LearningMathematical Reasoning

Model-Based Policy Gradients with Parameter-Based Exploration by Least-Squares Conditional Density Estimation

2013-07-19 · Syogo Mori, Voot Tangkaratt, Tingting Zhao, Jun Morimoto 외

The goal of reinforcement learning (RL) is to let an agent learn an optimal control policy in an unknown environment so that future expected rewards are maximized. The model-free RL approach directly learns the policy ba…

Density EstimationReinforcement LearningReinforcement Learning (RL)

SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints

2025-12-29 · Dominik Wagner, Ankit Kanwar, Luke Ong arxiv

In safety-critical domains, reinforcement learning (RL) agents must often satisfy strict, zero-cost safety constraints while accomplishing tasks. Existing model-free methods frequently either fail to achieve near-zero sa…

Reinforcement Learning