paper-with-me

홈 › Papers

Generative Planning for Temporally Coordinated Exploration in Reinforcement Learning

2022-01-24 · ICLR 2022 4 · Haichao Zhang, Wei Xu, Haonan Yu

Standard model-free reinforcement learning algorithms optimize a policy that generates the action to be taken in the current time step in order to maximize expected future return. While flexible, it faces difficulties arising from the inefficient exploration due to its single step nature. In this work, we present Generative Planning method (GPM), which can generate actions not only for the current step, but also for a number of future steps (thus termed as generative planning). This brings several benefits to GPM. Firstly, since GPM is trained by maximizing value, the plans generated from it can be regarded as intentional action sequences for reaching high value regions. GPM can therefore leverage its generated multi-step plans for temporally coordinated exploration towards high value regions, which is potentially more effective than a sequence of actions generated by perturbing each action at single step level, whose consistent movement decays exponentially with the number of exploration steps. Secondly, starting from a crude initial plan generator, GPM can refine it to be adaptive to the task, which, in return, benefits future explorations. This is potentially more effective than commonly used action-repeat strategy, which is non-adaptive in its form of plans. Additionally, since the multi-step plan can be interpreted as the intent of the agent from now to a span of time period into the future, it offers a more informative and intuitive signal for interpretation. Experiments are conducted on several benchmark environments and the results demonstrated its effectiveness compared with several baseline methods.

📄 PDF Abstract BibTeX arXiv:2201.09765

Code (1)

Haichao-Zhang/generative-planning 공식 구현 pytorch

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Subwords as Skills: Tokenization for Sparse-Reward Reinforcement Learning

2023-09-08 · David Yunis, Justin Jung, Falcon Dai, Matthew Walter

Exploration in sparse-reward reinforcement learning is difficult due to the requirement of long, coordinated sequences of actions in order to achieve any reward. Moreover, in continuous action spaces there are an infinit…

reinforcement-learningReinforcement Learning

Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings

2018-06-07 · ICML 2018 7 · John D. Co-Reyes, Yuxuan Liu, Abhishek Gupta, Benjamin Eysenbach 외

In this work, we take a representation learning perspective on hierarchical reinforcement learning, where the problem of learning lower layers in a hierarchy is transformed into the problem of learning trajectory-level g…

Hierarchical Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

Learning to Plan via Deep Optimistic Value Exploration

2020-06-08 · L4DC 2020 6 · Tim Seyde, Wilko Schwarting, Sertac Karaman, Daniela Rus

Deep exploration requires coordinated long-term planning. We present a model-based reinforcement learning algorithm that guides policy learning through a value function that exhibits optimism in the face of uncertainty. …

BenchmarkingModel-based Reinforcement LearningReinforcement Learning (RL)

Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control

2018-11-05 · ICLR 2019 5 · Kendall Lowrey, Aravind Rajeswaran, Sham Kakade, Emanuel Todorov 외

We propose a plan online and learn offline (POLO) framework for the setting where an agent, with an internal model, needs to continually act and learn in the world. Our work builds on the synergistic relationship between…

Flexible and Efficient Long-Range Planning Through Curious Exploration

2020-04-22 · ICML 2020 1 · Aidan Curtis, Minjian Xin, Dilip Arumugam, Kevin Feigelis 외

Identifying algorithms that flexibly and efficiently discover temporally-extended multi-phase plans is an essential step for the advancement of robotics and model-based reinforcement learning. The core problem of long-ra…

Deep Reinforcement LearningImitation LearningModel-based Reinforcement LearningMotion Planning+4