paper-with-me

홈 › Papers

Exploration by Maximizing Rényi Entropy for Reward-Free RL Framework

2020-06-11 · Chuheng Zhang, Yuanying Cai, Longbo Huang, Jian Li

Exploration is essential for reinforcement learning (RL). To face the challenges of exploration, we consider a reward-free RL framework that completely separates exploration from exploitation and brings new challenges for exploration algorithms. In the exploration phase, the agent learns an exploratory policy by interacting with a reward-free environment and collects a dataset of transitions by executing the policy. In the planning phase, the agent computes a good policy for any reward function based on the dataset without further interacting with the environment. This framework is suitable for the meta RL setting where there are many reward functions of interest. In the exploration phase, we propose to maximize the Renyi entropy over the state-action space and justify this objective theoretically. The success of using Renyi entropy as the objective results from its encouragement to explore the hard-to-reach state-actions. We further deduce a policy gradient formulation for this objective and design a practical exploration algorithm that can deal with complex environments. In the planning phase, we solve for good policies given arbitrary reward functions using a batch RL algorithm. Empirically, we show that our exploration algorithm is effective and sample efficient, and results in superior policies for arbitrary reward functions in the planning phase.

📄 PDF Abstract BibTeX arXiv:2006.06193

Code (0)

등록된 구현이 없습니다.

Tasks

Q-LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Entropy Regularization 설명 없음
PPO Proximal Policy Optimization, or PPO, is a policy gradient method for reinforcement learning. The motivation was to have an algorithm with the data efficiency and reliable…
Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…

Similar Papers 제목 키워드 기반

MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization

2024-12-16 · Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel 외

Reinforcement learning (RL) algorithms aim to balance exploiting the current best strategy with exploring new options that could lead to higher rewards. Most common RL algorithms use undirected exploration, i.e., select …

Multi-Armed BanditsReinforcement Learning (RL)

k-Means Maximum Entropy Exploration

2022-05-31 · Alexander Nedergaard, Matthew Cook

Exploration in high-dimensional, continuous spaces with sparse rewards is an open problem in reinforcement learning. Artificial curiosity algorithms address this by creating rewards that lead to exploration. Given a rein…

Density Estimationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Skew-Explore: Learn faster in continuous spaces with sparse rewards

2019-09-25 · Xi Chen, Yuan Gao, Ali Ghadirzadeh, Marten Bjorkman 외

In many reinforcement learning settings, rewards which are extrinsically available to the learning agent are too sparse to train a suitable policy. Beside reward shaping which requires human expertise, utilizing better e…

Off-Policy Actor-Critic in an Ensemble: Achieving Maximum General Entropy and Effective Environment Exploration in Deep Reinforcement Learning

2019-02-14 · Gang Chen, Yiming Peng

We propose a new policy iteration theory as an important extension of soft policy iteration and Soft Actor-Critic (SAC), one of the most efficient model free algorithms for deep reinforcement learning. Supported by the n…

Deep Reinforcement LearningReinforcement Learning

Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning

2026-03-26 · Jiajun Hu, Nuria Armengol Urpi, Jin Cheng, Stelian Coros arxiv

Zero-shot reinforcement learning (RL) algorithms aim to learn a family of policies from a reward-free dataset, and recover optimal policies for any reward function directly at test time. Naturally, the quality of the pre…

Reinforcement Learning