paper-with-me

홈 › Papers

Hierarchical Reinforcement Learning via Advantage-Weighted Information Maximization

2019-01-05 · ICLR 2019 5 · Takayuki Osa, Voot Tangkaratt, Masashi Sugiyama

Real-world tasks are often highly structured. Hierarchical reinforcement learning (HRL) has attracted research interest as an approach for leveraging the hierarchical structure of a given task in reinforcement learning (RL). However, identifying the hierarchical policy structure that enhances the performance of RL is not a trivial task. In this paper, we propose an HRL method that learns a latent variable of a hierarchical policy using mutual information maximization. Our approach can be interpreted as a way to learn a discrete and latent representation of the state-action space. To learn option policies that correspond to modes of the advantage function, we introduce advantage-weighted importance sampling. In our HRL method, the gating policy learns to select option policies based on an option-value function, and these option policies are optimized based on the deterministic policy gradient method. This framework is derived by leveraging the analogy between a monolithic policy in standard RL and a hierarchical policy in HRL by using a deterministic option policy. Experimental results indicate that our HRL approach can learn a diversity of options and that it can enhance the performance of RL in continuous control tasks.

📄 PDF Abstract BibTeX arXiv:1901.01365

Code (1)

TakaOsa/adInfoHRL 공식 구현 tf

Tasks

continuous-controlContinuous ControlDiversityHierarchical Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Off-policy Maximum Entropy Reinforcement Learning : Soft Actor-Critic with Advantage Weighted Mixture Policy(SAC-AWMP)

2020-02-07 · Zhimin Hou, Kuangen Zhang, Yi Wan, Dongyu Li 외

The optimal policy of a reinforcement learning problem is often discontinuous and non-smooth. I.e., for two states with similar representations, their optimal policies can be significantly different. In this case, repres…

continuous-controlContinuous ControlMixture-of-ExpertsReinforcement Learning

FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification

2026-06-29 · Zheming Fu, Ruizhe He, Wei Shang, Xiaoxiao Ma 외 arxiv

Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing density-approximated policy gradient methods rely on stochastic SDE sa…

Reinforcement Learning

Reward is not Necessary: How to Create a Modular & Compositional Self-Preserving Agent for Life-Long Learning

2022-11-20 · Thomas J. Ringstrom

Reinforcement Learning views the maximization of rewards and avoidance of punishments as central to explaining goal-directed behavior. However, over a life, organisms will need to learn about many different aspects of th…

Hierarchical Importance Weighted Autoencoders

2019-05-13 · Chin-wei Huang, Kris Sankaran, Eeshan Dhekane, Alexandre Lacoste 외

Importance weighted variational inference (Burda et al., 2015) uses multiple i.i.d. samples to have a tighter variational lower bound. We believe a joint proposal has the potential of reducing the number of redundant sam…

Variational Inference

Hierarchical Reinforcement Learning in Multi-Goal Spatial Navigation with Autonomous Mobile Robots

2025-04-26 · Brendon Johnson, Alfredo Weitzenfeld

Hierarchical reinforcement learning (HRL) is hypothesized to be able to take advantage of the inherent hierarchy in robot learning tasks with sparse reward schemes, in contrast to more traditional reinforcement learning …

Hierarchical Reinforcement Learningreinforcement-learningReinforcement Learning